Method, apparatus, and electronic device for detecting smoking behavior based on video
Deep learning-based video analysis improves smoking detection accuracy by encoding pose and temporal features, addressing the challenges of pose variation and environmental interference in existing systems.
Patent Information
- Application Number
- CN202210270938.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-18
AI Technical Summary
In the prior art, the accuracy of detecting smoking behavior is low, especially due to the difference in smoking postures of people and posture confusion.
Using video-based deep learning method, the single-frame feature vector and frame sequence feature vector are extracted, and the posture feature information and posture feature change information are combined to identify smoking behavior.
It improves the accuracy of detecting smoking behavior, reduces the impact of non-smoking behavior postures on recognition results, and can accurately identify smoking behaviors.
Smart Images

Figure CN114612840B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of deep learning, and for example, relates to a method, an apparatus, and an electronic device for detecting smoking behavior based on video. Background Art
[0002] Currently, some public places have attempted to install a no-smoking alarm system. When the system detects someone smoking, it will issue a warning and can even capture the actual situation of the smoking scene. For example, the detection of smoking behavior can be achieved by detecting the smoke generated by smoking through a smoke sensor, or by using image recognition technology to identify the cigarette characteristics in the image. However, the detection result of detecting the smoke generated by smoking using a smoke sensor is affected by the environmental ventilation condition and the space volume, and the accuracy of detecting smoking behavior is relatively low; the target characteristics of the cigarette are relatively small, and moreover, it is also affected by adverse factors such as the diversity of cigarette shapes, the diversity of light changes, and the diversity of backgrounds, resulting in a relatively low accuracy of identifying the cigarette characteristics in the image.
[0003] In some existing technologies, the human body posture in the image is detected, and the smoking posture is determined as the smoking behavior, which improves the accuracy of detecting smoking behavior to a certain extent.
[0004] In the process of implementing the embodiments of this application, it is found that there are at least the following problems in the related technologies:
[0005] There are certain differences in the smoking postures of people, and moreover, the smoking posture is easily confused with other postures, such as the posture of making a phone call, the posture of eating, etc. Therefore, the recognition accuracy of identifying smoking behavior by detecting the human body posture is relatively low. Summary of the Invention
[0006] To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. The summary is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments, but rather serves as a preamble to the subsequent detailed description.
[0007] The embodiments of this application provide a method, an apparatus, and an electronic device for detecting smoking behavior based on video, so as to improve the detection accuracy of detecting smoking behavior.
[0008] In some embodiments, the method for detecting smoking behavior based on video includes: obtaining a video segment to be recognized; performing encoding processing on each video frame in the video segment to be recognized to obtain a single-frame feature vector for each video frame; performing encoding processing on all the single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector, where the position encoding corresponding to the single-frame feature vector is used to represent the sequence order of the video frame corresponding to each single-frame feature vector in the video segment to be recognized; and performing classification processing on the frame sequence feature vector to obtain a detection result of smoking behavior.
[0009] Optionally, performing encoding processing on each video frame in the video segment to be recognized to obtain a single-frame feature vector for each video frame includes: performing convolution processing on each video frame using a first-layer convolutional network to obtain the output of the first-layer convolutional network; performing max-pooling processing on the output of the first-layer convolutional network to obtain a pooling result; performing convolution processing on the pooling result through an intermediate-layer convolutional network, and determining the combination of the convolution processing and the output of the previous-layer convolutional network as the output of the intermediate-layer convolutional network; performing average-pooling processing on the output of the last-layer convolutional network, and processing the average-pooling result using a fully-connected layer to obtain the single-frame feature vector.
[0010] Optionally, performing encoding processing on all the single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector includes: obtaining a vector to be encoded according to all the single-frame feature vectors and their corresponding position encodings; processing each sub-vector to be encoded in the vector to be encoded in the following manner: performing encoding processing on the sub-vector to be encoded according to the adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded and all the sub-vectors in the vector to be encoded to obtain a sub-feature vector of the sub-vector to be encoded, where the frame sequence feature vector includes the sub-feature vectors of all the sub-vectors to be encoded.
[0011] Optionally, performing encoding processing on the sub-vector to be encoded according to the adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded and all the sub-vectors in the vector to be encoded to obtain a sub-feature vector of the sub-vector to be encoded includes:
[0012] In the case where the sub-vector to be encoded is an unlabeled sub-vector, obtaining a set number of adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded; determining a first weight corresponding to the sub-vector to be encoded according to the set number of sub-vectors and the sub-vector to be encoded; and performing encoding processing on the product of the sub-vector to be encoded and the first weight to obtain a sub-feature vector of the sub-vector to be encoded.
[0013] When the sub-vector to be encoded is a marked sub-vector, obtain a plurality of sub-vectors in the vector to be encoded that are in the same row and / or the same column as the sub-vector to be encoded; determine a second weight corresponding to the sub-vector to be encoded according to the plurality of sub-vectors and the sub-vector to be encoded; perform encoding processing on the product of the sub-vector to be encoded and the second weight to obtain a sub-feature vector of the sub-vector to be encoded.
[0014] Optionally, performing classification processing on the frame sequence feature vector to obtain a detection result of a smoking behavior, including: performing feature representation integration processing on the frame sequence feature vector to obtain a feature representation integration vector; wherein, the dimension of the feature representation integration vector is lower than that of the frame sequence feature vector; performing normalization processing and classification processing on the feature representation integration vector to obtain the detection result of the smoking behavior.
[0015] Optionally, obtaining a video segment to be recognized, including: performing segment sampling on an original video to obtain an original video segment; performing normalization processing on the original video segment; sampling in the normalized video to obtain the video segment to be recognized.
[0016] Optionally, the method for detecting a smoking behavior based on a video further includes: obtaining an intermediate video frame in the video segment to be recognized; performing human body detection on the intermediate video frame to obtain a human body detection result.
[0017] In some embodiments, a device for detecting a smoking behavior based on a video includes: a first obtaining module, a first encoding module, a second encoding module, and a classification module; the first obtaining module is configured to obtain a video segment to be recognized; the first encoding module is configured to perform encoding processing on each video frame in the video segment to be recognized to obtain a single-frame feature vector of each video frame; the second encoding module is configured to perform encoding processing on all the single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector; wherein, the position encoding corresponding to the single-frame feature vector is used to represent the order of each video frame corresponding to the single-frame feature vector in the video segment to be recognized; the classification module is configured to perform classification processing on the frame sequence feature vector to obtain a detection result of a smoking behavior.
[0018] In some embodiments, an electronic device includes a processor and a memory storing program instructions, and the processor is configured to execute the method for detecting a smoking behavior based on a video provided in the foregoing embodiments when executing the program instructions.
[0019] In some embodiments, a storage medium stores program instructions, and the program instructions execute the method for detecting a smoking behavior based on a video provided in the foregoing embodiments when running.
[0020] The method, device, and electronic device for detecting smoking behavior based on video provided by the embodiments of the present application can achieve the following technical effects:
[0021] The embodiments of the present application adopt deep learning technology and use computer vision to identify smoking behavior. In this technical solution, each video frame in the video segment to be recognized is encoded to obtain a single-frame feature vector, which contains pose feature information. Then, the single feature vector and its corresponding position encoding are encoded to obtain a frame sequence feature vector, which contains the change information of the pose feature. In this way, in the process of identifying smoking behavior by combining the pose feature information and the change information of the pose feature, the influence of the single-frame pose feature information on the recognition result is reduced, and the influence of the pose information of non-smoking behavior on the recognition result can be reduced. Since this recognition process combines the change information of the pose feature, the smoking behavior can be recognized more accurately.
[0022] The above general description and the following description are only exemplary and explanatory, and are not used to limit the present application. Description of the Drawings
[0023] One or more embodiments are exemplarily illustrated by the corresponding drawings. These exemplary illustrations and the drawings do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are regarded as similar elements, and among them:
[0024] Figure 1 is a schematic flowchart of a method for detecting smoking behavior based on video provided by the embodiments of the present application;
[0025] Figure 2 is a schematic flowchart of a method for detecting smoking behavior based on video provided by the embodiments of the present application;
[0026] Figure 3 is a schematic flowchart of a method for detecting smoking behavior based on video provided by the embodiments of the present application;
[0027] Figure 4 is a schematic diagram of a device for detecting smoking behavior based on video provided by the embodiments of the present application;
[0028] Figure 5 is a schematic diagram of an electronic device for detecting smoking behavior based on video provided by the embodiments of the present application. Detailed Embodiments
[0029] In order to understand the features and technical content of the embodiments of the present application in more detail, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration purposes only and are not used to limit the embodiments of the present application. In the following technical description, for the sake of explanation, multiple details are provided to provide a full understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices can be shown in a simplified manner to simplify the drawings.
[0030] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0031] Unless otherwise specified, the term "plurality" means more than two.
[0032] In the embodiments of the present application, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.
[0033] The term "and / or" is a description of the associated relationship of objects and indicates that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B.
[0034] The smoking postures of people have certain differences, and moreover, the smoking postures are easily confused with other postures, such as the posture of making a phone call, the posture of eating, etc. In the embodiments of the present application, by combining the posture information included in the single-frame feature vector of each video frame in the video and the order of posture changes included in the frame sequence feature vector, the influence of other behaviors on the detection result of the smoking behavior is reduced, and the accuracy of detecting the smoking behavior is improved.
[0035] Figure 1 It is a schematic flow chart of a method for detecting smoking behavior based on video provided by the embodiments of the present application.
[0036] Combined with Figure 1 As shown, the method for detecting smoking behavior based on video includes:
[0037] S101. Obtain a video segment to be recognized.
[0038] The video segment to be recognized here refers to the video segment used for detecting smoking behavior.
[0039] The video segment to be recognized can be obtained in the following way: perform segment sampling on the original video to obtain the original video segment; perform normalization processing on the original video segment; sample in the normalized video to obtain the video segment to be recognized.
[0040] The original video here can be a video obtained by a camera device, and the video frame rate in the original video is related to the performance indicators of the camera device.
[0041] Segment sampling can be performed on the original video through a sliding window. Sampling in the original video to obtain the original video segment may include: collecting a segment of the first set number of video frames in the original video to obtain the original video segment of the first set number of video frames; or, collecting a segment of the set duration in the original video to obtain the original video segment of the set duration.
[0042] The above-mentioned first set number of video frames can be 50-70 video frames. For example, the first set number of video frames is 50 video frames, 52 video frames, 54 video frames, 56 video frames, 58 video frames, 60 video frames, 62 video frames, 64 video frames, 66 video frames, 68 video frames or 70 video frames.
[0043] In some application scenarios, the size of the sliding window can be set to 64 video frames, and the sliding interval is 8 video frames. In this way, the process of starting sampling from the first video frame in the original video is as follows: collect the 1st to 64th video frames in the original video to obtain an original video segment; move the sliding window 8 video frames, collect the 9th to 72nd video frames in the original video to obtain an original video segment; move the sliding window 8 video frames again, collect the 17th to 81st video frames in the original video to obtain an original video segment. The sampling process listed here is only used to exemplarily illustrate the process of obtaining the video to be recognized, and those skilled in the art can also set sliding windows and sliding intervals of other sizes according to the actual situation.
[0044] The original video segment can be normalized in the following way: calculate the average value and variance of all video frames in the original video segment; crop each video frame in the original video segment to obtain a video frame of the set size (such as 224*224); subtract the average value from the video frame of the set size and then divide by the variance to obtain the normalized video frame.
[0045] Sampling can be performed in the normalized video in the following way to obtain the video segment to be recognized: sample one video frame every second set number of video frames. The second set number of video frames can be 1-3 video frames. For example, the second set number of video frames can be 1 video frame, 2 video frames or 3 video frames. In this way, it is beneficial to reduce the video to be recognized to reduce the calculation amount and improve the speed of detecting smoking behavior.
[0046] Through the above technical solutions, the video clip to be recognized can be obtained from the original video.
[0047] S102. Perform encoding processing on each video frame in the video clip to be recognized to obtain the single-frame feature vector of each video frame.
[0048] The purpose of this encoding processing is to extract the features of each video frame in the video clip to be recognized. The single-frame feature vector of each video frame contains the pose features in this video frame. The following method can be used to perform encoding processing on each video frame in the video clip to be recognized to obtain the single-frame feature vector of each video frame: Use the first convolutional network to perform convolution processing on each video frame to obtain the output of the first convolutional network; perform max pooling processing on the output of the first convolutional network to obtain the pooling processing result; perform convolution processing on the pooling processing result through the intermediate convolutional network, and determine the combination of the convolution processing and the output of the previous convolutional network as the output of the intermediate convolutional network; perform average pooling processing on the output of the last convolutional network, and use the fully connected layer to process the average pooling processing result to obtain the single-frame feature vector.
[0049] For example, a residual network (ResNet) can be used to perform encoding processing on each video frame in the video clip to be recognized to obtain the single-frame feature vector of each video frame.
[0050] Of course, a recurrent neural network (RNN), a long short-term memory network (LSTM), etc. can also be used to perform encoding processing on each video frame in the video clip to be recognized to obtain the single-frame feature vector of each video frame. The present application does not make specific limitations in this regard.
[0051] S103. Perform encoding processing on all single-frame feature vectors and their corresponding position encodings to obtain the frame sequence feature vector.
[0052] Among them, the position encoding corresponding to the single-frame feature vector is used to represent the sequence order of the video frames corresponding to each single-frame feature vector in the video clip to be recognized. In this way, the frame sequence feature vector contains the change order of multiple single-frame feature vectors.
[0053] Optionally, encoding processing is performed on all single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector, including: obtaining a vector to be encoded based on all single-frame feature vectors and their corresponding position encodings; processing each sub-vector to be encoded in the vector to be encoded in the following manner: encoding the sub-vector to be encoded based on the adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded and all the sub-vectors in the vector to be encoded to obtain a sub-feature vector of the sub-vector to be encoded; wherein, the frame sequence feature vector includes the sub-feature vectors of all sub-vectors to be encoded.
[0054] Further, obtaining a vector to be encoded based on a single-frame feature vector and its corresponding position encoding may include: determining the sum of all video frame feature vectors and a position encoding vector as the vector to be encoded; wherein, the sub-vectors in all video frame feature vectors are single-frame feature vectors, and the sub-vectors in the position encoding vector are used to represent the position encodings corresponding to the single-frame feature vectors.
[0055] Encoding the sub-vector to be encoded based on the adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded, then the sub-vector encoding corresponding to the sub-vector to be encoded contains the information of the adjacent sub-vectors of the sub-vector to be encoded; encoding the sub-vector to be encoded based on all the sub-vectors in the vector to be encoded, then the sub-vector encoding corresponding to the sub-vector to be encoded contains the information of all the sub-vectors in the vector to be encoded. According to this encoding method, more accurate frame sequence feature vectors can be extracted from the vector to be encoded, which is beneficial to obtaining more accurate detection results of smoking behaviors.
[0056] The following further explains obtaining the sub-feature vector of the sub-vector to be encoded:
[0057] In the case where the sub-vector to be encoded is an unlabeled sub-vector, obtaining a set number of adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded; determining a first weight corresponding to the sub-vector to be encoded based on the set number of sub-vectors and the sub-vector to be encoded; encoding the product of the sub-vector to be encoded and the first weight to obtain a sub-feature vector of the sub-vector to be encoded;
[0058] In the case where the sub-vector to be encoded is a labeled sub-vector, obtaining a plurality of sub-vectors in the same row and / or the same column as the sub-vector to be encoded in the vector to be encoded; determining a second weight corresponding to the sub-vector to be encoded based on the plurality of sub-vectors and the sub-vector to be encoded; encoding the product of the sub-vector to be encoded and the second weight to obtain a sub-feature vector of the sub-vector to be encoded.
[0059] Wherein, the labeled sub-vector refers to the sub-vector that needs to be focused on in the vector to be encoded. Those skilled in the art can determine the labeled positions according to the segments that need to be focused on in the video segment to be recognized, and the sub-vectors at the labeled positions in the vector to be recognized are the labeled sub-vectors.
[0060] In the process of obtaining the sub-feature vector of the sub-vector to be encoded, by considering the features represented by a set number of sub-vectors adjacent to the sub-vector to be encoded, or by considering the features represented by multiple sub-vectors that are in the same row and / or the same column as the sub-vector to be encoded among the adjacent ones to be encoded, a more accurate sub-feature vector of the sub-vector to be encoded can be obtained, which is conducive to obtaining a more accurate detection result of the smoking behavior.
[0061] The following details the process of obtaining the sub-feature vector of the sub-vector to be encoded:
[0062] Obtaining a set number of sub-vectors adjacent to the sub-vector to be encoded in the vector to be encoded may include: obtaining a set number of consecutive sub-vectors adjacent to the sub-vector to be encoded in the vector to be encoded, or obtaining a set number of spaced sub-vectors adjacent to the sub-vector to be encoded in the vector to be encoded.
[0063] The above-mentioned consecutive sub-vectors refer to the sub-vectors with consecutive positions in the vector to be encoded. The set number of consecutive sub-vectors can be the sub-vectors whose positions are before the sub-vector to be encoded, or the sub-vectors whose positions are after the sub-vector to be encoded, or the sub-vectors whose positions are on both sides of the sub-vector to be encoded. For example, half of the set number of sub-vectors are before the sub-vector to be encoded, and the other half are after the sub-vector to be encoded.
[0064] For example: the vector to be encoded is V, and the sub-vector to be encoded is v ij , i is the row number of the sub-vector to be encoded v ij in the vector to be encoded V, j is the column number of the sub-vector to be encoded v ij in the vector to be encoded V, the set number is m, then the set number of sub-vectors can be v ij-m , v ij-m+1 , v ij-m+2 , …, v ij-1 ; or, the set number of sub-vectors can be v ij+1 , v ij+2 , …, v ij+m ; or, the set number of sub-vectors can be v ij-m1 , …, v ij-2 , v ij-1 , v ij+1 , v ij+2 , …, v ij+m2 , where m2 - m1 = m. Here, only the set number of consecutive sub-vectors adjacent to the sub-vector to be encoded is exemplarily illustrated, which does not constitute a specific limitation on the embodiments of the present application. Those skilled in the art can select appropriate consecutive sub-vectors with a set number adjacent to the sub-vector to be encoded according to the actual situation.
[0065] The above-mentioned spacer vectors refer to the sub-vectors with position intervals in the vector to be encoded. A set number of spacer vectors can be the sub-vectors before the sub-vector to be encoded, or the sub-vectors after the sub-vector to be encoded, or the sub-vectors on both sides of the sub-vector to be encoded. For example, half of the set number of sub-vectors are before the sub-vector to be encoded, and the other half of the set number of sub-vectors are after the sub-vector to be encoded.
[0066] For example, the vector to be encoded is V, and the sub-vector to be encoded is v ij , where i is the row number of the sub-vector v to be encoded in the vector V to be encoded ij , and j is the column number of the sub-vector v to be encoded in the vector V to be encoded. If the set number is n, the set number of sub-vectors can be v ij , v ij-2n , V ij-n+2 , …, v ij-n+4 ; or, the set number of sub-vectors can be v ij-2 , v ij+2 , …, v ij+4 ; or, the set number of sub-vectors can be v ij+2n , …, v ij-n1 , v ij-4 , v ij-2 , v ij+2 , v ij+4 , …, v ij+n2 , where n2 - n1 = 2n. Here, only the set number of spacer vectors adjacent to the sub-vector to be encoded are exemplarily described, which does not constitute a specific limitation on the embodiments of the present application. Those skilled in the art can select the appropriate set number of spacer vectors adjacent to the sub-vector to be encoded according to the actual situation.
[0067] Further, determining the first weight corresponding to the sub-vector to be encoded according to the set number of sub-vectors and the sub-vector to be encoded includes: obtaining the first matching degree between each of the set number of sub-vectors and the first query vector, and the second matching degree between the sub-vector to be encoded and the first query vector; dividing the second matching degree by the sum of all the first matching degrees and the second matching degree to obtain the first weight. Wherein, the first query vector corresponds to the sub-vector to be encoded. For example, the sub-vector to be encoded can be determined as the first query vector.
[0068] Specifically, the first weight can be obtained in the following manner:
[0069]
[0070] Where a1 is the first weight, s(·) is the matching degree evaluation model, v ij is the vector to be encoded, q1 is the first query vector, v ilis the l-th of a set number of sub-vectors, and m is the set number.
[0071] Further, determining the second weight corresponding to the vector to be encoded according to a plurality of sub-vectors and the vector to be encoded includes: obtaining the third matching degree between each sub-vector in the plurality of sub-vectors and the second query vector, and the fourth matching degree between the vector to be encoded and the second query vector; using the fourth matching degree divided by the sum of all the third matching degrees and the fourth matching degree to obtain the second weight. Wherein, the second query vector corresponds to the sub-vector to be encoded. For example, the sub-vector to be encoded can be determined as the second query vector.
[0072] Specifically, the second weight can be obtained in the following manner:
[0073]
[0074] Wherein, a2 is the second weight, s(·) is the matching degree evaluation model, v ij is the vector to be encoded, q1 is the first query vector, v il is the l-th of a plurality of sub-vectors, and p is the number of the plurality of sub-vectors.
[0075] The matching degree evaluation model s(·) will be further described in detail below:
[0076] s(·) can be a Feedforward Neural Network (FNN). In this way, inputting the vector to be encoded v ij and the first query vector q1 into the FNN, the FNN can output the matching degree between the vector to be encoded v ij and the first query vector q1; inputting the vector to be encoded v ij and the second query vector q2 into the FNN, the FNN can output the matching degree between the vector to be encoded v ij and the second query vector q2.
[0077] Or, s(v, q) = h T tanh(Wv + Uq);
[0078] Or, s(v, q) = v T q;
[0079] Or, s(v, q) = v T Wq
[0080] Wherein, W, U, and h are all learnable parameter matrices or vectors.
[0081] The first weight or the second weight can be obtained through the above technical solution, and then the product of the sub-vector to be encoded and the first weight can be encoded to obtain the sub-feature vector of the sub-vector to be encoded; or, the product of the sub-vector to be encoded and the second weight can be encoded to obtain the sub-feature vector of the sub-vector to be encoded, and then the frame sequence feature vector can be obtained.
[0082] S104. Classify the frame sequence feature vector to obtain the detection result of the smoking behavior.
[0083] For example, a classifier can be used to classify the frame sequence feature vector to obtain the detection result of the smoking behavior.
[0084] Alternatively, classifying the frame sequence feature vector to obtain the detection result of the smoking behavior may include: performing feature representation integration processing on the frame sequence feature vector to obtain a feature representation integration vector; wherein, the dimension of the feature representation integration vector is lower than that of the frame sequence feature vector; performing normalization processing and classification processing on the feature representation integration vector to obtain the detection result of the smoking behavior.
[0085] For example, the frame sequence feature vector can be subjected to feature representation processing through one or more fully-connected neural networks to obtain a feature representation integration vector.
[0086] Specifically, when the number of layers of the fully-connected neural network is one, the frame sequence feature vector is input into the fully-connected neural network, and the output of the fully-connected neural network is determined as the feature representation integration vector;
[0087] When the number of layers of the fully-connected neural network is multiple, the frame sequence feature vector is input into the first layer of the fully-connected neural network; for the intermediate layer fully-connected neural network other than the first layer, the output of the previous layer of the fully-connected neural network is determined as the input of the intermediate layer fully-connected neural network; the output of the last layer of the fully-connected neural network is determined as the feature representation integration vector.
[0088] By performing feature representation integration processing on the frame sequence feature representation vector, a feature representation integration vector with a lower dimension (lower than the dimension of the frame sequence feature representation vector) is obtained, which is convenient for the classifier to accurately classify the feature representation integration vector.
[0089] The classifier can output the detection result and the probability distribution of the detection result, and the detection result with the highest probability can be determined as the detection result of the smoking behavior.
[0090] The embodiments of the present application adopt deep learning technology and use computer vision to identify smoking behavior. In this technical solution, in the method for detecting smoking behavior based on video provided by the embodiments of the present application, each video frame in the video segment to be recognized is encoded to obtain a single-frame feature vector, which contains pose feature information. Then, the single feature vector and its corresponding position encoding are encoded to obtain a frame sequence feature vector, which contains the change information of the pose feature. In this way, in the process of identifying smoking behavior by combining the pose feature information and the change information of the pose feature, the influence of single-frame pose feature information on the recognition result can be reduced, and the influence of the pose information of non-smoking behavior on the recognition result can be reduced. Since this recognition process combines the change information of the pose feature, the smoking behavior can be recognized more accurately.
[0091] Figure 2 It is a schematic flowchart of a method for detecting smoking behavior based on video provided by the embodiments of the present application.
[0092] Combined with Figure 2 As shown, the method for detecting smoking behavior based on video includes:
[0093] S201. Obtain the video segment to be recognized.
[0094] S202. Encode each video frame in the video segment to be recognized to obtain a single-frame feature vector for each video frame.
[0095] S203. Encode all the single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector.
[0096] Among them, the position encoding corresponding to the single-frame feature vector is used to represent the order of the video frames corresponding to each single-frame feature vector.
[0097] S204. Classify the frame sequence feature vector to obtain the detection result of the smoking behavior.
[0098] S205. Obtain the intermediate video frame in the video segment to be recognized.
[0099] For example, when the number of video frames in the video segment to be recognized is 32 video frames, the 16th video frame or the 17th video frame is determined as the intermediate video frame.
[0100] S206. Perform human detection on the intermediate video frame to obtain the human detection result.
[0101] By using the above-mentioned method for smoking behavior, not only the detection result of the smoking behavior can be obtained, but also the human detection result can be obtained. When warning about the smoking behavior, the smoking behavior and the smoking person can be displayed more clearly.
[0102] In addition, after obtaining the detection result of the smoking behavior, regardless of whether the detection result indicates the existence of the smoking behavior, the steps of obtaining the intermediate video frames of the video segment to be recognized and performing human body detection on the intermediate video frames to obtain the human body detection result are continued. If the detection result indicates the existence of the smoking behavior, the human body detection result is displayed for warning. Alternatively, after obtaining the detection result of the smoking behavior, it is first determined whether the detection result indicates the existence of the smoking behavior. If there is no smoking behavior, the subsequent steps are not executed. If there is a smoking behavior, the subsequent steps are continued: obtaining the intermediate video frames of the video segment to be recognized, performing human body detection on the intermediate video frames to obtain the human body detection result, and displaying the human body detection result to realize the warning of the smoking behavior.
[0103] Specifically, the model of the object detection dynamic training algorithm (Dynamic RCNN) can be used to perform human body detection on the intermediate video frames to obtain the human body detection result.
[0104] In some specific applications, to obtain the model parameters of the above method for detecting smoking behavior based on video, a large number of video segments can be collected and labeled by a camera device, such as 10,000 video segments, 15,000 video segments, or 20,000 video segments; and these video segments are manually labeled, for example, the video segments with smoking behavior are labeled, or, while labeling the video segments with smoking behavior, the smoking personnel are labeled. Using these labeled video segments as training samples, the model of the method for detecting smoking behavior based on video is trained. During the training process, stochastic gradient descent can be used to optimize the model parameters, with a learning rate of 0.1, a momentum of 0.9, and a weight decay coefficient of 0.0001, and the training stops after 30 epochs.
[0105] Figure 3 It is a schematic flowchart of a method for detecting smoking behavior based on video provided by an embodiment of the present application, which is used to exemplarily illustrate the method for detecting smoking behavior based on video in combination with an application scenario.
[0106] Combined with Figure 3 As shown, the method for detecting smoking behavior based on video includes:
[0107] S301. Obtain the original video.
[0108] The video stream collected by the camera device can be used as the original video.
[0109] S302. Use a sliding window to perform segment sampling on the original video to obtain the original video segment.
[0110] The sliding window can be moved once every 8 video frames, or once every 0.5 s.
[0111] S303. Preprocess the original video clip.
[0112] The preprocessing here may include normalizing the original video clip and sampling in the original video clip to obtain the video clip to be recognized (preprocessing result).
[0113] S304. Recognize the preprocessing result to obtain the detection result of the smoking behavior.
[0114] S305. Perform human body detection to obtain the human body detection result.
[0115] S306. When the detection result of the smoking behavior indicates the existence of the smoking behavior, display the detection result of the smoking behavior and the human body detection result.
[0116] For example, the human body detection result can be displayed through a human body bounding box.
[0117] Figure 4 It is a schematic diagram of a device for detecting smoking behavior based on video provided by an embodiment of the present application.
[0118] Combined with Figure 4 As shown, the device for detecting smoking behavior based on video includes a first acquisition module 41, a first encoding module 42, a second encoding module 43, and a classification module 44; the first acquisition module 41 is used to acquire the video clip to be recognized; the first encoding module 42 is used to perform encoding processing on each video frame in the video clip to be recognized to obtain a single-frame feature vector of each video frame; the second encoding module 43 is used to perform encoding processing on all single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector; wherein, the position encoding corresponding to the single-frame feature vector is used to represent the sequence order of the video frame corresponding to each single-frame feature vector in the video clip to be recognized; the classification module 44 is used to perform classification processing on the frame sequence feature vector to obtain the detection result of the smoking behavior.
[0119] Optionally, the first encoding module 42 includes a first convolutional unit, a first pooling unit, a second convolutional unit, and a second pooling unit; the first convolutional unit is used to perform convolutional processing on each video frame by using a first-layer convolutional network to obtain the output of the first-layer convolutional network; the first pooling unit performs max pooling processing on the output of the first-layer convolutional network to obtain a pooling processing result; the second convolutional unit is used to perform convolutional processing on the pooling processing result through an intermediate-layer convolutional network and determine the combination of the convolutional processing and the output of the previous-layer convolutional network as the output of the intermediate-layer convolutional network; the second pooling unit is used to perform average pooling processing on the output of the last-layer convolutional network and process the average pooling processing result by using a fully connected layer to obtain a single-frame feature vector.
[0120] Optionally, the second encoding module 43 includes a first obtaining unit and an encoding unit; the first obtaining unit is configured to obtain a vector to be encoded according to all single-frame feature vectors and their corresponding position encodings; the encoding unit is configured to process each sub-vector to be encoded in the vector to be encoded in the following manner: perform encoding processing on the sub-vector to be encoded according to the adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded and all sub-vectors in the vector to be encoded, to obtain a sub-feature vector of the sub-vector to be encoded; wherein, the frame sequence feature vector includes sub-feature vectors of all sub-vectors to be encoded.
[0121] Optionally, performing encoding processing on the sub-vector to be encoded according to the adjacent sub-vectors of the sub-vector to be encoded in the vector to be encoded and all sub-vectors in the vector to be encoded, to obtain a sub-feature vector of the sub-vector to be encoded, includes: in the case where the sub-vector to be encoded is an unmarked sub-vector, obtaining a set number of sub-vectors adjacent to the sub-vector to be encoded in the vector to be encoded; determining a first weight corresponding to the sub-vector to be encoded according to the set number of sub-vectors and the sub-vector to be encoded; performing encoding processing on the product of the sub-vector to be encoded and the first weight, to obtain a sub-feature vector of the sub-vector to be encoded; in the case where the sub-vector to be encoded is a marked sub-vector, obtaining a plurality of sub-vectors in the same row and / or the same column as the sub-vector to be encoded in the vector to be encoded; determining a second weight corresponding to the sub-vector to be encoded according to the plurality of sub-vectors and the sub-vector to be encoded; performing encoding processing on the product of the sub-vector to be encoded and the second weight, to obtain a sub-feature vector of the sub-vector to be encoded.
[0122] Optionally, the classification module 44 includes a second obtaining unit and a classification unit; the second obtaining unit is configured to perform feature representation integration processing on the frame sequence feature vector to obtain a feature representation integration vector; wherein, the dimension of the feature representation integration vector is lower than that of the frame sequence feature vector; the classification unit is configured to perform normalization processing and classification processing on the feature representation integration vector to obtain a detection result of the smoking behavior.
[0123] Optionally, the first obtaining module 41 includes a third obtaining unit, a normalization processing unit, and a fourth obtaining unit; the third obtaining unit is configured to perform segment sampling on the original video to obtain an original video segment; the normalization processing unit is configured to perform normalization processing on the original video segment; the fourth obtaining unit is configured to sample in the normalized video to obtain a video segment to be recognized.
[0124] Optionally, the device for detecting smoking behavior based on video further includes a second obtaining module and a third obtaining module; the second obtaining module is configured to obtain an intermediate video frame in the video segment to be recognized; the third obtaining module is configured to perform human body detection on the intermediate video frame to obtain a human body detection result.
[0125] In some embodiments, the device for detecting smoking behavior includes a processor and a memory storing program instructions. The processor is configured to execute the method for detecting smoking behavior based on video provided in the foregoing embodiments when executing the program instructions.
[0126] Figure 5 is a schematic diagram of an electronic device for detecting smoking behavior based on video provided by an embodiment of the present application. Combining Figure 5 As shown, the electronic device for detecting smoking behavior based on video includes:
[0127] A processor 51 and a memory 52, and may further include a communication interface 53 and a bus 54. Among them, the processor 51, the communication interface 53, and the memory 52 can complete mutual communication through the bus 54. The communication interface 53 can be used for information transmission. The processor 51 can call the logical instructions in the memory 52 to execute the method for detecting smoking behavior based on video provided in the foregoing embodiments.
[0128] In addition, when the logical instructions in the above-mentioned memory 52 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0129] The memory 52, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of the present application. The processor 51 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 52, that is, implements the methods in the above method embodiments.
[0130] The memory 52 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 52 may include a high-speed random access memory and may also include a non-volatile memory.
[0131] An embodiment of the present application provides a storage medium storing program instructions, and the program instructions execute the method for detecting smoking behavior based on video provided in the foregoing embodiments when running.
[0132] An embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the method for detecting smoking behavior based on video provided in the foregoing embodiments.
[0133] The above computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.
[0134] The technical solution of the embodiment of the present application may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method in the embodiment of the present application. The foregoing storage medium may be a non-transient storage medium, including: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or may also be a transient storage medium.
[0135] The above description and the drawings fully illustrate the embodiments of the present application so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments only represent possible variations. Unless explicitly required, separate components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in the present application are only for describing the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Additionally, when used in the present application, the term "comprise" and its variants "comprises" and / or "comprising", etc. refer to the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, or device including the element. In this document, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method parts disclosed in the embodiments, the relevant parts may refer to the description of the method parts.
[0136] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner can depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of this application. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0137] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. In addition, in the embodiments of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0138] The flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the systems, methods, and computer program products according to the embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block can also occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which can depend on the functions involved. Each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for detecting smoking behavior based on video, characterized in that, Including: Obtain the video segment to be recognized; Perform encoding processing on each video frame in the video segment to be recognized to obtain a single-frame feature vector for each video frame; Perform encoding processing on all the single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector; wherein, the position encoding corresponding to the single-frame feature vector is used to represent the sequence order of each video frame corresponding to the single-frame feature vector in the video segment to be recognized; Perform classification processing on the frame sequence feature vector to obtain the detection result of the smoking behavior; Performing encoding processing on all the single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector, including: Obtain the vector to be encoded according to all the single-frame feature vectors and their corresponding position encodings; Process each sub-vector to be encoded in the vector to be encoded in the following manner: In the case where the sub-vector to be encoded is an unlabeled sub-vector, obtain a set number of sub-vectors adjacent to the sub-vector to be encoded in the vector to be encoded; determine the first weight corresponding to the sub-vector to be encoded according to the set number of sub-vectors and the sub-vector to be encoded; perform encoding processing on the product of the sub-vector to be encoded and the first weight to obtain the sub-feature vector of the sub-vector to be encoded; In the case where the sub-vector to be encoded is a labeled sub-vector, obtain a plurality of sub-vectors in the same row and / or the same column as the sub-vector to be encoded in the vector to be encoded; determine the second weight corresponding to the sub-vector to be encoded according to the plurality of sub-vectors and the sub-vector to be encoded; perform encoding processing on the product of the sub-vector to be encoded and the second weight to obtain the sub-feature vector of the sub-vector to be encoded; Wherein, the frame sequence feature vector includes the sub-feature vectors of all the sub-vectors to be encoded.
2. The method according to claim 1, wherein Performing encoding processing on each video frame in the video segment to be recognized to obtain a single-frame feature vector for each video frame, including: Perform convolution processing on each video frame using the first-layer convolution network to obtain the output of the first-layer convolution network; Perform max pooling processing on the output of the first-layer convolution network to obtain the result of the pooling processing; Perform convolution processing on the result of the pooling processing through the intermediate-layer convolution network, and determine the output of the intermediate-layer convolution network by combining the convolution processing result and the output of the previous-layer convolution network; Perform average pooling processing on the output of the last-layer convolution network, and process the result of the average pooling processing using the fully connected layer to obtain the single-frame feature vector.
3. The method according to claim 1, wherein Performing classification processing on the frame sequence feature vector to obtain the detection result of the smoking behavior, including: Perform feature representation integration processing on the frame sequence feature vector to obtain a feature representation integration vector; wherein, the dimension of the feature representation integration vector is lower than that of the frame sequence feature vector; Perform normalization processing and classification processing on the feature representation integration vector to obtain the detection result of the smoking behavior.
4. The method according to any one of claims 1 to 3, characterized in that Obtain the video segment to be recognized, including: Perform segment sampling on the original video to obtain the original video segment; Perform normalization processing on the original video segment; Perform sampling on the normalized video to obtain the video segment to be recognized.
5. The method according to any one of claims 1 to 3, characterized in that, Further comprising: Obtaining an intermediate video frame in the video segment to be recognized; Performing human body detection on the intermediate video frame to obtain a human body detection result.
6. A video-based device for detecting smoking behavior, characterized in that, Comprising: A first obtaining module, configured to obtain a video segment to be recognized; A first encoding module, configured to perform encoding processing on each video frame in the video segment to be recognized to obtain a single-frame feature vector of each video frame; A second encoding module, configured to perform encoding processing on all the single-frame feature vectors and their corresponding position encodings to obtain a frame sequence feature vector; wherein, the position encoding corresponding to the single-frame feature vector is used to represent the sequence of the video frames corresponding to each single-frame feature vector in the video segment to be recognized; A classification module, configured to perform classification processing on the frame sequence feature vector to obtain a detection result of a smoking behavior; The second encoding module includes a first obtaining unit and an encoding unit; the first obtaining unit is configured to obtain a vector to be encoded according to all the single-frame feature vectors and their corresponding position encodings; the encoding unit is configured to process each sub-vector to be encoded in the vector to be encoded in the following manner: in the case where the sub-vector to be encoded is an unmarked sub-vector, obtaining a set number of sub-vectors adjacent to the sub-vector to be encoded in the vector to be encoded; determining a first weight corresponding to the sub-vector to be encoded according to the set number of sub-vectors and the sub-vector to be encoded; performing encoding processing on the product of the sub-vector to be encoded and the first weight to obtain a sub-feature vector of the sub-vector to be encoded; in the case where the sub-vector to be encoded is a marked sub-vector, obtaining a plurality of sub-vectors in the same row and / or the same column as the sub-vector to be encoded in the vector to be encoded; determining a second weight corresponding to the sub-vector to be encoded according to the plurality of sub-vectors and the sub-vector to be encoded; performing encoding processing on the product of the sub-vector to be encoded and the second weight to obtain a sub-feature vector of the sub-vector to be encoded; wherein, the frame sequence feature vector includes sub-feature vectors of all the sub-vectors to be encoded.
7. An electronic device, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the method for detecting smoking behavior based on video according to any one of claims 1 to 5 when executing the program instructions.
8. A storage medium stores program instructions, characterized in that, The program instructions execute the method for detecting smoking behavior based on video according to any one of claims 1 to 5 when running.
Citation Information
Patent Citations
Video processing method and device
CN111222493A
Information identification method and device, storage medium and electronic equipment
CN113947694A