Smoking behavior detection method and device

By combining the residual network model and the feature pyramid network, image feature extraction and semantic mapping processing are performed, which solves the problem of low accuracy in smoking behavior detection and achieves higher detection accuracy.

CN114140878BActive Publication Date: 2025-09-12SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111443689.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-09-12
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

The accuracy of smoking behavior detection in the existing technology is low.

Method used

A residual network model is used to extract the first feature of the input image, and then the features are fused through feature pyramid network and semantic mapping processing to detect smoking behavior.

Benefits of technology

Improved the accuracy of smoking behavior detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140878B_ABST
    Figure CN114140878B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technology and provides a method and device for detecting smoking behavior. The method comprises: obtaining an input image and extracting a first feature of the input image using a residual network model; inputting the first feature into a feature pyramid network to output a second feature; performing a first semantic mapping process on the second feature to obtain a third feature; performing a second semantic mapping process on the second feature to obtain a fourth feature; performing a feature fusion process on the second, third, and fourth features to obtain a fused feature; and detecting smoking behavior based on the fused feature. The above-described technical approach addresses the low accuracy of smoking behavior detection in existing technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a method and device for detecting smoking behavior. Background Art

[0002] Smoking is harmful to health. For the sake of public health, smoking is prohibited in many places. To achieve intelligent smoking detection, traditional methods rely on cigarette detection. Detecting a cigarette indicates that someone is holding a cigarette and that smoking has occurred or is about to occur. However, because cigarette butts are small and their features are weak, these detection methods struggle to achieve high accuracy in smoking detection based solely on cigarette detection.

[0003] In the process of realizing the concept of the present disclosure, the inventors discovered that there are at least the following technical problems in the related art: the problem of low accuracy in detecting smoking behavior. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a method and apparatus for detecting smoking behavior to solve the problem of low accuracy in detecting smoking behavior in the prior art.

[0005] A first aspect of an embodiment of the present disclosure provides a method for detecting smoking behavior, comprising: obtaining an input image and extracting a first feature of the input image through a residual network model; inputting the first feature into a feature pyramid network to output a second feature; performing a first semantic mapping process on the second feature to obtain a third feature, performing a second semantic mapping process on the second feature to obtain a fourth feature, performing a feature fusion process on the second feature, the third feature, and the fourth feature to obtain a fused feature; and detecting smoking behavior based on the fused feature.

[0006] According to a second aspect of an embodiment of the present disclosure, a device for detecting smoking behavior is provided, comprising: a feature extraction module configured to obtain an input image and extract a first feature of the input image through a residual network model; a pyramid network module configured to input the first feature into a feature pyramid network and output a second feature; a semantic mapping module configured to perform a first semantic mapping process on the second feature to obtain a third feature, and perform a second semantic mapping process on the second feature to obtain a fourth feature; a feature fusion module configured to perform a feature fusion process on the second feature, the third feature, and the fourth feature to obtain a fused feature; and a detection module configured to detect smoking behavior based on the fused feature.

[0007] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0008] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0009] Compared with the prior art, the embodiments of the present disclosure have the following beneficial effects: because the embodiments of the present disclosure extract the first feature of the input image through the residual network model; input the first feature into the feature pyramid network to output the second feature; perform a first semantic mapping process on the second feature to obtain a third feature, perform a second semantic mapping process on the second feature to obtain a fourth feature, perform a feature fusion process on the second feature, the third feature and the fourth feature to obtain a fused feature; and detect smoking behavior based on the fused feature. Therefore, the above technical means can solve the problem of low accuracy in detecting smoking behavior in the prior art, thereby improving the accuracy of detecting smoking behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 is a schematic diagram of an application scenario of an embodiment of the present disclosure;

[0012] Figure 2 is a flow chart of a method for detecting smoking behavior provided by an embodiment of the present disclosure;

[0013] Figure 3 1 is a schematic structural diagram of a smoking behavior detection device provided by an embodiment of the present disclosure;

[0014] Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present disclosure with unnecessary detail.

[0016] A method and device for detecting smoking behavior according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0017] Figure 1 FIG2 is a schematic diagram of an application scenario of an embodiment of the present disclosure. The application scenario may include terminal devices 1, 2, and 3, a server 4, and a network 5.

[0018] The terminal devices 1, 2 and 3 can be hardware or software. When the terminal devices 1, 2 and 3 are hardware, they can be various electronic devices with display screens and supporting communication with the server 4, including but not limited to smart phones, tablet computers, laptop computers and desktop computers; when the terminal devices 1, 2 and 3 are software, they can be installed in the above electronic devices. The terminal devices 1, 2 and 3 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiments of the present disclosure are not limited to this. Furthermore, various applications can be installed on the terminal devices 1, 2 and 3, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0019] Server 4 can be a server that provides various services, for example, a backend server that receives requests sent by terminal devices that establish communication connections with it. The backend server can receive and analyze the requests sent by the terminal devices, and generate processing results. Server 4 can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center, all of which are not limited in the present embodiment.

[0020] It should be noted that the server 4 can be either hardware or software. When the server 4 is hardware, it can be various electronic devices that provide various services to the terminal devices 1, 2, and 3. When the server 4 is software, it can be multiple software programs or software modules that provide various services to the terminal devices 1, 2, and 3, or it can be a single software program or software module that provides various services to the terminal devices 1, 2, and 3. This is not limited in the present embodiment.

[0021] The network 5 can be a wired network connected by coaxial cable, twisted pair and optical fiber, or it can be a wireless network that can interconnect various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), infrared, etc. The embodiments of the present disclosure are not limited to this.

[0022] Users can establish a communication connection with the server 4 via the network 5 through the terminal devices 1, 2, and 3 to receive or send information, etc. It should be noted that the specific types, quantities, and combinations of the terminal devices 1, 2, and 3, the server 4, and the network 5 can be adjusted according to the actual needs of the application scenario, and the embodiments of the present disclosure are not limited thereto.

[0023] Figure 2 It is a flowchart of a method for detecting smoking behavior provided by an embodiment of the present disclosure. Figure 2 The detection method of smoking behavior can be Figure 1 The terminal device or server executes. Figure 2 As shown, the smoking behavior detection method includes:

[0024] S201, obtaining an input image, and extracting a first feature of the input image through a residual network model;

[0025] S202, inputting the first feature into a feature pyramid network and outputting a second feature;

[0026] S203, performing a first semantic mapping process on the second feature to obtain a third feature, and performing a second semantic mapping process on the second feature to obtain a fourth feature;

[0027] S204, performing feature fusion processing on the second feature, the third feature, and the fourth feature to obtain a fused feature;

[0028] S205: Perform smoking behavior detection based on the fusion features.

[0029] Resnet50 can be used as the residual network model. The feature pyramid network FPN (feature pyramid networks) can fuse the underlying semantic information and high-level semantic information of the first feature to obtain the second feature. The first semantic mapping process is to map the second feature to the feature of the previous stage of the second feature in the residual network model, that is, the third feature. For example, the residual network model has four stages, and the second feature is the final output of the residual network model, that is, the output of the fourth stage of the residual network model, then the third feature is the output of the third stage of the residual network model. The first semantic mapping process is to map the second feature to the feature of the previous stage of the previous stage of the second feature in the residual network model, that is, the third feature. For example, the residual network model has four stages, and the second feature is the final output of the residual network model, that is, the output of the fourth stage of the residual network model, then the fourth feature is the output of the second stage of the residual network model.

[0030] According to the technical solution provided by the embodiment of the present disclosure, because the embodiment of the present disclosure extracts the first feature of the input image through the residual network model; inputs the first feature into the feature pyramid network to output the second feature; performs a first semantic mapping process on the second feature to obtain a third feature, performs a second semantic mapping process on the second feature to obtain a fourth feature, performs feature fusion processing on the second feature, the third feature and the fourth feature to obtain a fused feature; and performs smoking behavior detection based on the fused feature. Therefore, the above-mentioned technical means can solve the problem of low accuracy in smoking behavior detection in the existing technology, thereby improving the accuracy of smoking behavior detection.

[0031] In step S204, feature fusion processing is performed on the second feature, the third feature, and the fourth feature to obtain a fused feature, including: performing pooling processing on the second feature, the third feature, and the fourth feature respectively to obtain a fifth feature, a sixth feature, and a seventh feature; and performing the feature fusion processing on the fifth feature, the sixth feature, and the seventh feature to obtain the fused feature.

[0032] Before performing the first semantic mapping process on the second feature and performing the second semantic mapping process on the second feature, a non-maximum suppression (NMS) process may be further performed on the second feature.

[0033] For example: the first feature is input into the feature pyramid network, and after the second feature is output, the second feature is subjected to non-maximum suppression (NMS) processing to obtain three human frames (the reason for predicting the human body is that the semantic information of the human body is advanced and rich. First detecting whether it is a human body, and then detecting the subsequent face / hand / cigarette butt can effectively reduce subsequent false detections). The three human frames corresponding to the second feature are subjected to the first semantic mapping processing to obtain the third feature. The third feature is pooled to obtain the fifth feature with a dimension of (24, 8, 1024), where 24 represents the height of the fifth feature, 8 represents the width of the fifth feature, and 1024 is the number of channels. The three human frames corresponding to the second feature are subjected to the second semantic mapping processing to obtain the fourth feature. The fourth feature is pooled to obtain the fifth feature with a dimension of (48, 16, 512), where 48 represents the height of the fifth feature, 16 represents the width of the fifth feature, and 512 is the number of channels. The second feature is pooled to obtain a fifth feature with a dimension of (12, 4, 2048). 12 represents the height of the fifth feature, 4 represents the width of the fifth feature, and 2048 is the number of channels.

[0034] The reason why the second feature is subjected to the first semantic mapping process to obtain the third feature, and the second feature is subjected to the second semantic mapping process to obtain the fourth feature, is because the second feature is the final output of the residual network model, such as the detected human body, the third feature is the feature of the previous stage of the second feature in the residual network model, such as the detected face / hand, and the fourth feature is the feature of the previous stage of the previous stage of the second feature in the residual network model, such as the detected cigarette butt. Here, we draw on the idea of ​​faster-rcnn and innovate. Faster-rcnn maps back to the corresponding layer, while the embodiment of the present disclosure predicts the human body through the second feature and maps back to the third feature to predict the face / hand. The third and fourth features have richer detail information than the second feature, and are very friendly to the detection and positioning of small objects.

[0035] In step S204, the feature fusion processing is performed on the fifth feature, the sixth feature, and the seventh feature to obtain the fused feature, including: inputting the sixth feature and the seventh feature into a preset network group respectively, and outputting the eighth feature and the ninth feature, wherein the preset network group is composed of a convolution layer, a normalization layer, and an activation layer connected in sequence; and the feature fusion processing is performed on the fifth feature, the eighth feature, and the ninth feature to obtain the fused feature.

[0036] The preset network group can be understood as a component, which is composed of a convolution layer, a normalization layer, and an activation layer connected in sequence.

[0037] In step S204, the feature fusion processing is performed on the fifth feature, the eighth feature, and the ninth feature to obtain the fused feature, including: performing a first interactive calculation on the fifth feature and the eighth feature to obtain the tenth feature; and performing a second interactive calculation on the ninth feature and the tenth feature to obtain the fused feature.

[0038] The fifth feature is information about the human body, the eighth feature is information about the face and hand, and the ninth feature is information about the cigarette butt. This disclosed embodiment first performs interactive calculations on the face / hand and the human body, and then performs interactive calculations on the cigarette butt with the above calculation results. This completes cross-scale information fusion and solves the problem of errors that can easily occur when determining whether smoking has occurred based solely on detecting cigarette butts or cigarettes.

[0039] In an optional embodiment, a first interactive calculation is performed on the fifth feature and the eighth feature to obtain a tenth feature, including: flattening the eighth feature to obtain an eleventh feature, and multiplying the transpose of the eleventh feature by a first parameter matrix to obtain a first matrix; passing the fifth feature through a first preset convolution layer to obtain a convolution result, and flattening the convolution result to obtain a second matrix; and multiplying the first matrix by the transpose of the second matrix and then by the second matrix to obtain the tenth feature.

[0040] It should be noted that, because feature information can be represented by vectors or matrices, the features in the present disclosure, including the first feature, the second feature, etc., can be understood to exist in the form of matrices or vectors.

[0041] For example, the eighth feature is flattened into the eleventh feature with a dimension of (24x8, 1024). The eleventh feature is transposed and multiplied by the first parameter matrix Q1 to obtain a first matrix with a dimension of (12x4, 1024). The dimension of the first parameter matrix Q1 is (24x8, 12x4). The flattening process can be understood as a dimensionality reduction process. The eighth feature is a matrix, so the eleventh feature obtained by flattening the eighth feature is a vector. For example, 24x8 is used to represent the size of the eleventh feature after flattening, and 1024 is the number of channels. The first preset convolution layer is a convolution layer with a convolution kernel of 1x1 and 1024 channels. The fifth feature is passed through the first preset convolution layer to obtain a first convolution result, and the first convolution result is flattened into a second matrix with a dimension of (12x4, 1024). The first matrix is ​​multiplied by the transpose of the second matrix and then multiplied by the second matrix to obtain the tenth feature with a dimension of (12x4, 1024). The tenth feature is obtained through the above calculation. In smoking detection, the tenth feature integrates the specific relationship between the face, hand and body.

[0042] In an optional embodiment, a second interactive calculation is performed on the ninth feature and the tenth feature to obtain the fused feature, including: flattening the eighth feature to obtain the eleventh feature, and multiplying the transpose of the eleventh feature by the first parameter matrix to obtain a first matrix; passing the fifth feature through a first preset convolution layer to obtain a first convolution result, and performing the flattening process on the first convolution result to obtain a second matrix; multiplying the first matrix by the transpose of the second matrix and then by the second matrix to obtain the tenth feature; performing the flattening process on the ninth feature to obtain the twelfth feature, and multiplying the transpose of the twelfth feature by the second parameter matrix to obtain a third matrix; passing the tenth feature through a second preset convolution layer to obtain a second convolution result, and performing the flattening process on the second convolution result to obtain a fourth matrix; multiplying the third matrix by the transpose of the fourth matrix and then by the fourth matrix to obtain the fused feature.

[0043] Then, the ninth and tenth features are calculated, and the ninth feature is flattened into a twelfth feature of dimension (48x16, 512). The transpose of the twelfth feature is multiplied by the second parameter matrix Q2, which has a dimension of (48x16, 12x4), to obtain a third matrix of dimension (12x4, 512). The second preset convolution layer is a convolution layer with a convolution kernel of 1x1 and 512 channels. The tenth feature is passed through the second preset convolution layer to obtain a second convolution result, and the second convolution result is flattened into a fourth matrix of dimension (12x4, 1024). The third matrix is ​​multiplied by the transpose of the fourth matrix and then multiplied by the fourth matrix to obtain the fused feature, which integrates the specific relationship between cigarette butts, faces, hands, and human bodies.

[0044] In step S205, smoking behavior detection is performed based on the fused features, including: inputting the fused features into a fully connected layer network and outputting a detection result; wherein the detection result includes whether the target object has smoking behavior, whether the target object does not have smoking behavior, whether the target area has smoking behavior, and whether the target area does not have smoking behavior. The fully connected layer network includes multiple fully connected layers, and the fully connected layer network has been trained to learn and store the correspondence between the fused features and the smoking behavior.

[0045] In an optional embodiment, the input image is input into a neural network, and a detection result is output; wherein the detection result includes the presence of smoking behavior of the target object, the absence of smoking behavior of the target object, the presence of smoking behavior in the target area, and the absence of smoking behavior in the target area. The neural network has been trained to learn and save the correspondence between the input image and the smoking behavior.

[0046] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0047] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0048] Figure 3 FIG. 1 is a schematic diagram of a smoking behavior detection device provided by an embodiment of the present disclosure. Figure 3 As shown, the smoking behavior detection device includes:

[0049] The feature extraction module 301 is configured to obtain an input image and extract a first feature of the input image through a residual network model;

[0050] A pyramid network module 302 is configured to input the first feature into a feature pyramid network and output a second feature;

[0051] The semantic mapping module 303 is configured to perform a first semantic mapping process on the second feature to obtain a third feature, and perform a second semantic mapping process on the second feature to obtain a fourth feature;

[0052] A feature fusion module 304 is configured to perform feature fusion processing on the second feature, the third feature, and the fourth feature to obtain a fused feature;

[0053] The detection module 305 is configured to perform smoking behavior detection based on the fused features.

[0054] Resnet50 can be used as the residual network model. The feature pyramid network FPN (feature pyramid networks) can fuse the underlying semantic information and high-level semantic information of the first feature to obtain the second feature. The first semantic mapping process is to map the second feature to the feature of the previous stage of the second feature in the residual network model, that is, the third feature. For example, the residual network model has four stages, and the second feature is the final output of the residual network model, that is, the output of the fourth stage of the residual network model, then the third feature is the output of the third stage of the residual network model. The first semantic mapping process is to map the second feature to the feature of the previous stage of the previous stage of the second feature in the residual network model, that is, the third feature. For example, the residual network model has four stages, and the second feature is the final output of the residual network model, that is, the output of the fourth stage of the residual network model, then the fourth feature is the output of the second stage of the residual network model.

[0055] According to the technical solution provided by the embodiment of the present disclosure, because the embodiment of the present disclosure extracts the first feature of the input image through the residual network model; inputs the first feature into the feature pyramid network to output the second feature; performs a first semantic mapping process on the second feature to obtain a third feature, performs a second semantic mapping process on the second feature to obtain a fourth feature, performs feature fusion processing on the second feature, the third feature and the fourth feature to obtain a fused feature; and performs smoking behavior detection based on the fused feature. Therefore, the above-mentioned technical means can solve the problem of low accuracy in smoking behavior detection in the existing technology, thereby improving the accuracy of smoking behavior detection.

[0056] Optionally, the feature fusion module 304 is further configured to perform pooling processing on the second feature, the third feature and the fourth feature respectively to obtain the fifth feature, the sixth feature and the seventh feature; and perform the feature fusion processing on the fifth feature, the sixth feature and the seventh feature to obtain the fused feature.

[0057] Before performing the first semantic mapping process on the second feature and performing the second semantic mapping process on the second feature, a non-maximum suppression (NMS) process may be further performed on the second feature.

[0058] For example: the first feature is input into the feature pyramid network, and after the second feature is output, the second feature is subjected to non-maximum suppression (NMS) processing to obtain three human frames (the reason for predicting the human body is that the semantic information of the human body is advanced and rich. First detecting whether it is a human body, and then detecting the subsequent face / hand / cigarette butt can effectively reduce subsequent false detections). The three human frames corresponding to the second feature are subjected to the first semantic mapping processing to obtain the third feature. The third feature is pooled to obtain the fifth feature with a dimension of (24, 8, 1024), where 24 represents the height of the fifth feature, 8 represents the width of the fifth feature, and 1024 is the number of channels. The three human frames corresponding to the second feature are subjected to the second semantic mapping processing to obtain the fourth feature. The fourth feature is pooled to obtain the fifth feature with a dimension of (48, 16, 512), where 48 represents the height of the fifth feature, 16 represents the width of the fifth feature, and 512 is the number of channels. The second feature is pooled to obtain a fifth feature with a dimension of (12, 4, 2048). 12 represents the height of the fifth feature, 4 represents the width of the fifth feature, and 2048 is the number of channels.

[0059] The reason why the second feature is subjected to the first semantic mapping process to obtain the third feature, and the second feature is subjected to the second semantic mapping process to obtain the fourth feature, is because the second feature is the final output of the residual network model, such as the detected human body, the third feature is the feature of the previous stage of the second feature in the residual network model, such as the detected face / hand, and the fourth feature is the feature of the previous stage of the previous stage of the second feature in the residual network model, such as the detected cigarette butt. Here, we draw on the idea of ​​faster-rcnn and innovate. Faster-rcnn maps back to the corresponding layer, while the embodiment of the present disclosure predicts the human body through the second feature and maps back to the third feature to predict the face / hand. The third and fourth features have richer detail information than the second feature, and are very friendly to the detection and positioning of small objects.

[0060] Optionally, the feature fusion module 304 is further configured to input the sixth feature and the seventh feature into a preset network group respectively, and output the eighth feature and the ninth feature, wherein the preset network group is composed of a convolution layer, a normalization layer and an activation layer connected in sequence; and perform the feature fusion processing on the fifth feature, the eighth feature and the ninth feature to obtain the fused feature.

[0061] The preset network group can be understood as a component, which is composed of a convolution layer, a normalization layer, and an activation layer connected in sequence.

[0062] Optionally, the feature fusion module 304 is further configured to perform a first interactive calculation on the fifth feature and the eighth feature to obtain a tenth feature; and perform a second interactive calculation on the ninth feature and the tenth feature to obtain the fused feature.

[0063] The fifth feature is information about the human body, the eighth feature is information about the face and hand, and the ninth feature is information about the cigarette butt. This disclosed embodiment first performs interactive calculations on the face / hand and the human body, and then performs interactive calculations on the cigarette butt with the above calculation results. This completes cross-scale information fusion and solves the problem of errors that can easily occur when determining whether smoking has occurred based solely on detecting cigarette butts or cigarettes.

[0064] Optionally, the feature fusion module 304 is also configured to flatten the eighth feature to obtain an eleventh feature, and multiply the transpose of the eleventh feature by the first parameter matrix to obtain a first matrix; pass the fifth feature through a first preset convolution layer to obtain a convolution result, and flatten the convolution result to obtain a second matrix; multiply the first matrix by the transpose of the second matrix and then by the second matrix to obtain the tenth feature.

[0065] It should be noted that, because feature information can be represented by vectors or matrices, the features in the present disclosure, including the first feature, the second feature, etc., can be understood to exist in the form of matrices or vectors.

[0066] For example, the eighth feature is flattened into the eleventh feature with a dimension of (24x8, 1024). The eleventh feature is transposed and multiplied by the first parameter matrix Q1 to obtain a first matrix with a dimension of (12x4, 1024). The dimension of the first parameter matrix Q1 is (24x8, 12x4). The flattening process can be understood as a dimensionality reduction process. The eighth feature is a matrix, so the eleventh feature obtained by flattening the eighth feature is a vector. For example, 24x8 is used to represent the size of the eleventh feature after flattening, and 1024 is the number of channels. The first preset convolution layer is a convolution layer with a convolution kernel of 1x1 and 1024 channels. The fifth feature is passed through the first preset convolution layer to obtain a first convolution result, and the first convolution result is flattened into a second matrix with a dimension of (12x4, 1024). The first matrix is ​​multiplied by the transpose of the second matrix and then multiplied by the second matrix to obtain the tenth feature with a dimension of (12x4, 1024). The tenth feature is obtained through the above calculation. In smoking detection, the tenth feature integrates the specific relationship between the face, hand and body.

[0067] Optionally, the feature fusion module 304 is also configured to flatten the eighth feature to obtain the eleventh feature, and multiply the transpose of the eleventh feature by the first parameter matrix to obtain the first matrix; pass the fifth feature through the first preset convolution layer to obtain the first convolution result, and perform the flattening process on the first convolution result to obtain the second matrix; multiply the first matrix by the transpose of the second matrix and then multiply by the second matrix to obtain the tenth feature; perform the flattening process on the ninth feature to obtain the twelfth feature, and multiply the transpose of the twelfth feature by the second parameter matrix to obtain the third matrix; pass the tenth feature through the second preset convolution layer to obtain the second convolution result, and perform the flattening process on the second convolution result to obtain the fourth matrix; multiply the third matrix by the transpose of the fourth matrix and then multiply by the fourth matrix to obtain the fused feature.

[0068] Then, the ninth and tenth features are calculated, and the ninth feature is flattened into a twelfth feature of dimension (48x16, 512). The transpose of the twelfth feature is multiplied by the second parameter matrix Q2, which has a dimension of (48x16, 12x4), to obtain a third matrix of dimension (12x4, 512). The second preset convolution layer is a convolution layer with a convolution kernel of 1x1 and 512 channels. The tenth feature is passed through the second preset convolution layer to obtain a second convolution result, and the second convolution result is flattened into a fourth matrix of dimension (12x4, 1024). The third matrix is ​​multiplied by the transpose of the fourth matrix and then multiplied by the fourth matrix to obtain the fused feature, which integrates the specific relationship between cigarette butts, faces, hands, and human bodies.

[0069] Optionally, the detection module 305 is further configured to input the fused features into a fully connected layer network and output a detection result; wherein the detection result includes the presence of smoking behavior in the target object, the absence of smoking behavior in the target object, the presence of smoking behavior in the target area, and the absence of smoking behavior in the target area. The fully connected layer network includes multiple fully connected layers, and the fully connected layer network has been trained to learn and save the correspondence between the fused features and the smoking behavior.

[0070] Optionally, the detection module 305 is further configured to input the input image into a neural network and output a detection result; wherein the detection result includes whether the target object has smoking behavior, whether the target object does not have smoking behavior, whether the target area has smoking behavior, and whether the target area does not have smoking behavior. The neural network has been trained to learn and save the correspondence between the input image and the smoking behavior.

[0071] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.

[0072] Figure 4 FIG. 4 is a schematic diagram of an electronic device 4 provided in an embodiment of the present disclosure. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable by the processor 401. When the processor 401 executes the computer program 403, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-described device embodiments are implemented.

[0073] For example, computer program 403 may be divided into one or more modules / units, which are stored in memory 402 and executed by processor 401 to implement the present disclosure. One or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of computer program 403 in electronic device 4.

[0074] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include but is not limited to a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 4 and does not constitute a limitation of the electronic device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0075] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0076] Memory 402 can be an internal storage unit of electronic device 4, such as a hard drive or memory of electronic device 4. Memory 402 can also be an external storage device of electronic device 4, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on electronic device 4. Furthermore, memory 402 can include both an internal storage unit of electronic device 4 and an external storage device. Memory 402 is used to store computer programs and other programs and data required by the electronic device. Memory 402 can also be used to temporarily store data that has been output or is about to be output.

[0077] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0078] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0079] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0080] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0081] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0082] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0083] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0084] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the scope of protection of the present disclosure.

Claims

1. A method for detecting smoking behavior, characterized in that: include: Obtain an input image, and extract a first feature of the input image through a residual network model; Input the first feature into a feature pyramid network and output a second feature; Performing a first semantic mapping process on the second feature to obtain a third feature, and performing a second semantic mapping process on the second feature to obtain a fourth feature; performing feature fusion processing on the second feature, the third feature, and the fourth feature to obtain a fused feature; performing smoking behavior detection according to the fused features; The performing feature fusion processing on the second feature, the third feature, and the fourth feature to obtain a fused feature includes: Performing pooling processing on the second feature, the third feature, and the fourth feature respectively to obtain a fifth feature, a sixth feature, and a seventh feature; Inputting the sixth feature and the seventh feature into a preset network group respectively, and outputting an eighth feature and a ninth feature, wherein the preset network group is composed of a convolution layer, a normalization layer, and an activation layer connected in sequence; performing a first interactive calculation on the fifth feature and the eighth feature to obtain a tenth feature; performing a second interactive calculation on the ninth feature and the tenth feature to obtain the fused feature; The performing a first interactive calculation on the fifth feature and the eighth feature to obtain a tenth feature includes: Flattening the eighth feature to obtain an eleventh feature, and multiplying the transpose of the eleventh feature by the first parameter matrix to obtain a first matrix; Passing the fifth feature through a first preset convolution layer to obtain a convolution result, and performing the flattening process on the convolution result to obtain a second matrix; The tenth feature is obtained by multiplying the first matrix by the transpose of the second matrix and then multiplying the first matrix by the second matrix.

2. The method according to claim 1, characterized in that The performing a second interactive calculation on the ninth feature and the tenth feature to obtain the fused feature includes: Flattening the eighth feature to obtain an eleventh feature, and multiplying the transpose of the eleventh feature by the first parameter matrix to obtain a first matrix; Passing the fifth feature through a first preset convolution layer to obtain a first convolution result, and performing the flattening process on the first convolution result to obtain a second matrix; Multiplying the first matrix by the transpose of the second matrix and then by the second matrix to obtain a tenth feature; Performing the flattening process on the ninth feature to obtain a twelfth feature, and multiplying the transpose of the twelfth feature by the second parameter matrix to obtain a third matrix; Passing the tenth feature through a second preset convolution layer to obtain a second convolution result, and performing the flattening process on the second convolution result to obtain a fourth matrix; The fusion feature is obtained by multiplying the third matrix by the transpose of the fourth matrix and then multiplying the third matrix by the fourth matrix.

3. The method according to claim 1, characterized in that The detecting of smoking behavior according to the fusion feature includes: Input the fusion features into the fully connected layer network and output the detection results; The detection results include the presence of smoking behavior in the target object, the absence of smoking behavior in the target object, the presence of smoking behavior in the target area, and the absence of smoking behavior in the target area. The fully connected layer network includes multiple fully connected layers, and the fully connected layer network has been trained to learn and save the correspondence between the fusion features and the smoking behavior.

4. A smoking behavior detection device, characterized in that: include: a feature extraction module configured to obtain an input image and extract a first feature of the input image through a residual network model; a pyramid network module, configured to input the first feature into a feature pyramid network and output a second feature; a semantic mapping module configured to perform a first semantic mapping process on the second feature to obtain a third feature, and perform a second semantic mapping process on the second feature to obtain a fourth feature; a feature fusion module configured to perform feature fusion processing on the second feature, the third feature, and the fourth feature to obtain a fused feature; a detection module, configured to detect smoking behavior based on the fused features; The feature fusion module is further configured to: perform pooling processing on the second feature, the third feature, and the fourth feature respectively to obtain a fifth feature, a sixth feature, and a seventh feature; The sixth feature and the seventh feature are respectively input into a preset network group, and an eighth feature and a ninth feature are output, wherein the preset network group is composed of a convolution layer, a normalization layer and an activation layer connected in sequence; a first interactive calculation is performed on the fifth feature and the eighth feature to obtain a tenth feature; a second interactive calculation is performed on the ninth feature and the tenth feature to obtain the fusion feature; wherein, the first interactive calculation is performed on the fifth feature and the eighth feature to obtain the tenth feature, including: flattening the eighth feature to obtain an eleventh feature, and multiplying the transpose of the eleventh feature by a first parameter matrix to obtain a first matrix; passing the fifth feature through a first preset convolutional layer to obtain a convolution result, and performing the flattening process on the convolution result to obtain a second matrix; multiplying the first matrix by the transpose of the second matrix and then by the second matrix to obtain the tenth feature.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • River sewage draining exit detection method and system based on high-definition image

    CN112308040A