A method for recognizing phone call or smoking behaviors in special scenarios

Through face detection and image correction in special scenarios, combined with deep learning and multi-frame decision-making, the problems of low accuracy and high false alarm rate for phone calls and smoking behavior recognition are solved, and high-precision behavior recognition is achieved.

CN115731592BActive Publication Date: 2025-07-22HAINA CLOUD IOT TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210565525.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-07-22
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

The prior art has low accuracy in identifying calls and smoking behaviors in special scenarios, high false alarm rates and missed alarm rates, and insufficient data volume leads to poor model effects.

Method used

Face detection is used to determine the ROI area, image correction is used to generate images to be classified, multi-classification models are trained using deep learning, combined with feature similarity calculation and multi-frame target matching decisions, and output the final behavior results.

Benefits of technology

In the case of limited data volume, the recognition accuracy is significantly improved, the false positive rate is reduced, and the accuracy of the algorithm is further improved through the error classification feature library and multi-frame decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731592B_ABST
    Figure CN115731592B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for recognizing calling or smoking behaviors in special scenarios. The specific steps include: S1, collecting real-time monitoring images; S2, performing face detection to determine potential ROI regions; S3, correcting the images to generate images to be classified; S4, classifying the behaviors, and extracting the features of the images to be classified; S5, comparing the features and calculating the similarity; S6, comprehensively calculating the final result based on multi-frame comparison information and classification information; S7, outputting the result of whether calling or smoking according to the behavior corresponding to the final category. The present invention can greatly improve the algorithm accuracy under the condition of limited data volume, and further improve the algorithm accuracy by constructing a feature library for misclassified cases, and finally make a decision through the method of multi-frame target tracking and matching, reducing the false alarm rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and specifically, relates to a method for recognizing the behaviors of making a phone call or smoking in a special scenario. Background Art

[0002] Special scenarios such as gas stations, flour mills, and coal mines have a very high fire prevention requirement level. Therefore, behaviors such as smoking and making a phone call need to be closely monitored and prohibited to prevent fires caused by the above behaviors, thereby causing huge losses of life and property. The traditional recognition method is to use the entire image to detect the action behaviors of smoking and making a phone call to identify dangerous behaviors and give early warnings. However, since objects such as phones and cigarettes are relatively small in themselves, and the field of view angle of the monitoring camera is generally relatively large, resulting in a small proportion of target pixels, the false alarm rate and missed detection rate of smoking and making a phone call recognition are relatively high. For example, actions such as holding the cheek and touching the mouth are easily recognized as smoking and making a phone call behaviors, and are easily affected by light and background interference, so the application scenarios are limited. On the other hand, due to the lack of data on smoking and making a phone call behaviors, the data diversity and quantity generally used for training the detection model are relatively small, further resulting in poor model inference effects and low accuracy.

[0003] Chinese Patent with application number CN202110154950.X discloses a method for recognizing smoking and making a phone call in a specific scenario, which uses two cameras, a fixed camera and a PTZ camera. The fixed camera roughly locates through a human body detection algorithm and sends the coordinate points to the PTZ camera. The PTZ camera automatically zooms in according to the coordinates sent by the fixed camera, enlarges the human target area and inputs it into the classifier for recognition. The disadvantages of this method are: 1. It requires two devices to be combined for image acquisition, with high hardware requirements; 2. Selecting the human target with the largest target area for detection is likely to miss the behaviors of making a phone call and smoking; 3. Even if the target area image is enlarged, the resolution still needs to be reduced when entering the classifier, and the effective pixel proportion of the phone and smoking in the image is not increased, resulting in low classification accuracy.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a method for recognizing the behaviors of making a phone call or smoking in a special scenario

[0006] To solve the above technical problem, the basic concept of the technical solution adopted by the present invention is:

[0007] The present invention provides a method for recognizing the behaviors of making a phone call or smoking in a special scenario, and the specific steps include,

[0008] S1. Collect real-time monitoring images;

[0009] S2. Perform face monitoring to determine the potential ROI area;

[0010] S3. Correct the image to generate the image to be classified;

[0011] S4. Divide the behaviors into several categories, collect the training data of each category for deep learning training to obtain a multi-classification model, input the image to be classified into the classification model, and output the confidence value corresponding to each category and the features of the image to be classified;

[0012] S5. Calculate the similarity between the features of the image to be classified and each element in the feature set, and save the value with the highest similarity calculation in each category.

[0013] S6. Multiply the confidence value of the corresponding category by the similarity value, and take the category with the maximum product value as the final category of the current frame. If two-thirds of the images in a continuous number of frames have the same final category, output the final category;

[0014] S7. Output the result of whether a person is making a call or smoking according to the behavior corresponding to the final category.

[0015] Further, in step S2, the centerface network is used for face detection. Specifically, first scale the video frame output in step S1 to the image size supported by the centerface network, and then obtain the positions of all faces in the image and five key points of each face through convolution, pooling or activation. The five key points are the left eye center point, the right eye center point, the nose tip point, the left mouth corner point, and the right mouth corner point.

[0016] Further, the specific steps for generating the image to be classified by image correction in step S3 are as follows

[0017] S31. Take the set of key points to be corrected P = {p1, p2, p3, p4, p5}, where p1, p2, p3, p4, and p5 are the two-dimensional coordinates (xi, yi) of the key points to be corrected, and i belongs to [1, 5]. Let the set of standard key points S = {s1, s2, s3, s4, s5}, where s1, s2, s3, s4, and s5 are the two-dimensional coordinates (xi, yi) of the standard key points, and i belongs to [1, 5];

[0018] S32. Calculate the standard deviations of set P and set S respectively;

[0019] S33. Divide the elements of set P and set S by their corresponding standard deviations respectively to obtain new sets P' and S';

[0020] S34. Multiply the two new sets P' and S' to obtain a matrix M, perform SVD decomposition on matrix M, combine to obtain a rotation matrix and a parallel matrix, and intercept an image with a size of 224×224 pixels centered on the converted nose tip point as the image to be classified.

[0021] Further, the specific number of categories into which the behaviors are divided in step S4 is 8, namely Category 1: holding a cigarette with the hand, Category 2: holding a cigarette with the mouth, Category 3: holding a cigarette with the hand while smoking with the mouth, Category 4: holding a phone in the hand, Category 5: holding a phone next to the ear, Category 6: touching the mouth with the hand, Category 7: touching the ear with the hand, and Category 8: others.

[0022] Further, in step S4, the resnet50 network is used for deep learning training. The image to be classified is input into the resnet50 network, and while outputting the confidence levels of each classification, 256-dimensional features after dimensionality reduction are output.

[0023] Further, when the category with the maximum confidence level output by the training data does not match the true category in step S4, the 2048-dimensional deep features before the fully connected layer of the network are extracted, and using the PCA dimensionality reduction method, the 2048-dimensional features are reduced to 256-dimensional features and saved as the feature database for misclassification.

[0024] Further, the similarity calculation method in step 5 is as follows. The feature of the image to be classified is denoted as f, the feature set in the feature database is denoted as F, F = {f’1, f’2, ……, f’n}, where f, is an element in the feature set F, and n is the number of features in the feature set. The similarity between f and f,1 to f,n is calculated one by one, and the calculation formula is

[0025] sim = f × f’ / (norm(f) × norm(f’))

[0026] where the calculation formula of the norm() function is

[0027]

[0028] where V is an element in the feature f.

[0029] Further, the specific method of "multiplying the confidence level value corresponding to the category by the similarity value and taking the category with the maximum product as the final category of the current frame" in step S6 is

[0030] Calculated according to the following formula

[0031] max(w × t i + (1 - w) × T i )

[0032] where the max() function is the function for taking the maximum value, t i is the confidence level corresponding to the category, T i is the similarity to the corresponding category closest in the feature set, and w is the weighted weight, with a range of (0, 1).

[0033] Further, the weighted weight w is taken as 0.7.

[0034] Further, in step S6, the matching method for "consecutive frames of images" is to perform object matching by calculating the IOU value. Specifically, record the rectangular coordinate frames of all face regions in the current frame. When the next frame is input, multiple rectangular frames will also be obtained. Then, use the rectangular frames in the current frame to traverse the rectangular frames in the previous frame one by one. When traversing, calculate the overlapping area and the combined area of the two rectangular frames, and at the same time calculate the IOU value by dividing the overlapping area by the combined area. After traversing once, if the IOU value is greater than the set threshold, the matching is successful; otherwise, the matching fails. The face regions with failed matching are used as new objects and are prepared to perform the same calculation with the next frame of data.

[0035] After adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art.

[0036] 1. In the case of limited training data for smoking and making phone calls, by using image correction, behavior fine classification, feature extraction and comparison in the present invention, the model training can easily converge, and through algorithm strategies, the recognition accuracy of the algorithm can be greatly improved.

[0037] 2. During the testing and use process, extract the features of the images with deduced wrong categories and form a feature library for wrong classification. Subsequently, when making a decision, a certain weight can be given to the feature similarity of the feature library to assist in making a decision on the final category and improve the algorithm accuracy.

[0038] 3. Through multi-frame IOU object matching and comprehensive multi-frame information, finally make a decision on smoking and making phone call behaviors, which can greatly reduce the false alarm rate.

[0039] The following further describes in detail the specific implementation manners of the present invention with reference to the accompanying drawings. Description of the Drawings

[0040] The accompanying drawings, as a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not constitute an improper limitation to the present invention. Obviously, the accompanying drawings in the following description are only some embodiments. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0041] Figure 1 is a flow schematic diagram of a method for identifying phone call or smoking behaviors in a special scenario provided by the present invention.

[0042] It should be noted that these drawings and the textual descriptions are not intended to limit the scope of the concept of the present invention in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0043] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0044] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0045] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0046] As Figure 1 shown, the present invention provides a method for identifying calling or smoking behaviors in special scenarios. The specific steps include

[0047] S1. The camera collects real-time monitoring images and outputs them in the form of video frames;

[0048] S2. Perform face monitoring on the video frames to determine potential ROI regions;

[0049] S3. Correct the image to generate an image to be classified;

[0050] S4. Divide the behaviors into several categories, collect training data for each category for deep learning training to obtain a multi-classification model, input the image to be classified into the classification model, and output the confidence value corresponding to each category and the features of the image to be classified;

[0051] S5. Calculate the similarity between the features of the image to be classified and each element in the feature set, and save the value with the highest similarity calculation in each category.

[0052] S6. Multiply the confidence value of the corresponding category by the similarity value, and the category with the maximum product value is the final category of the current frame. If two-thirds of the images in several consecutive frames have the same final category, output the final category;

[0053] S7. Output the result of whether to make a call or smoke according to the behavior corresponding to the final category.

[0054] The above steps S1 to S7 are elaborated in detail below.

[0055] Step S1:

[0056] Use a fixed camera to collect the monitoring images of the scene in real time and input them into Step 2 in the form of video frames.

[0057] Step S2:

[0058] First, scale the input video frame to the image size supported by the network input. In this embodiment, the centerface network is used for face detection, and the image size supported by the centerface is 640×480 pixels. Then, through operations such as convolution, pooling, and activation, the centerface network obtains the positions of all faces in the entire image and five key points of each face. The five key points are the left eye pupil point, the right eye pupil point, the nose tip point, the left mouth corner point, and the right mouth corner point.

[0059] Step S3:

[0060] Correct each face in the image. The purpose of correction is to better perform behavior classification. After correction, a face region with an image size of 224×224 is intercepted to generate an image to be classified.

[0061] The method for correcting and intercepting each face is as follows:

[0062] S31. Take the set of key points of the face to be corrected as P = {p1, p2, p3, p4, p5}, where p1, p2, p3, p4, and p5 represent the left eye pupil point, the right eye pupil point, the nose tip point, the left mouth corner point, and the right mouth corner point respectively. Each point is a two-dimensional coordinate, that is, (xi, yi), and i belongs to [1, 5].

[0063] Another fixed set of standard key points is set as S = {s1, s2, s3, s4, s5}. Similarly, s1, s2, s3, s4, and s5 represent the left eye pupil point, the right eye pupil point, the nose tip point, the left mouth corner point, and the right mouth corner point respectively. Each point is a two-dimensional coordinate, that is, (xi, yi), and i belongs to [1, 5].

[0064] The purpose is to obtain the transformation matrix from the points to be corrected to the standard points, and then obtain the corrected face of the face to be corrected through the transformation matrix.

[0065] S32. Calculate the standard deviations of set P and set S respectively.

[0066] The mathematical definition of the standard deviation is the square root of the arithmetic mean of the squared deviations of the standard values of each unit in the population from their mean. Therefore, before calculating the standard deviation, it is necessary to calculate the means of two sets P and S, and the calculation formula is:

[0067] mean_px = (p1.x + p2.x + p3.x + p4.x + p5.x) / 5

[0068] mean_py = (p1.y + p2.y + p3.y + ≤p4.y + p5.y) / 5

[0069] mean_sx = (s1.x + s2.x + s3.x + s4.x + s5.x) / 5

[0070] mean_sy = (s1.x + s2.x + s3.x + s4.x + s5.x) / 5

[0071] After obtaining the means of sets P and S, subtract the corresponding means from the elements in the sets to obtain new sets P and S. Then, use the following formula to calculate the standard deviations of the new sets P and S:

[0072] std_px = sqrt((p1.x * p1.x + p2.x * p2.x + p3.x * p3.x + p4.x * p4.x + p5.x * p5.x) / 5)

[0073] std_py = sqrt((p1.y * p1.y + p2.y * p2.y + p3.y * p3.y + p4.y * p4.y + p5.y * p5.y) / 5)

[0074] std_sx = sqrt((s1.x * s1.x + s2.x * s2.x + s3.x * s3.x + s4.x * s4.x + s5.x * s5.x) / 5)

[0075] std_sy = sqrt((s1.y * s1.y + s2.y * s2.y + s3.y * s3.y + s4.y * s4.y + s5.y * s5.y) / 5)

[0076] Where the function sqrt() represents the square root function;

[0077] S33. Then calculate the normalized point sets, that is, the elements (x i , y i ) of sets P and S are divided by their corresponding standard deviations respectively to obtain new sets, denoted as P' and S'.

[0078] S34. Multiply two sets P' and S' to obtain matrix M. Perform SVD decomposition on matrix M. After combination, finally obtain the rotation matrix and the translation matrix, and take the center of the converted human face nose tip as the center to intercept an image of 224×224 pixels as the image to be classified.

[0079] Step S4:

[0080] Classify the behaviors into 8 categories, namely Category 1: holding a cigarette with hand, Category 2: holding a cigarette with mouth, Category 3: smoking with mouth while holding a cigarette with hand, Category 4: holding a phone, Category 5: holding a phone next to the ear, Category 6: touching the mouth with hand, Category 7: touching the ear with hand, and Category 8: others. Among the above 8 categories, Categories 1, 2, and 3 are all smoking behaviors, Categories 4 and 5 are phone call behaviors, and Categories 6 and 7 are normal behaviors.

[0081] Collect the data of the above 8 categories and preprocess them in the manner of Steps S2 and S3 to generate training data with each image size of 224×224.

[0082] Select the resnet50 network and perform deep learning training according to the training data to obtain a multi-classification model.

[0083] Before using the model, first conduct tests on the training data and actual scenarios. During the test process, input the image into the network and output the confidence levels belonging to the eight categories. If there is a difference between the category corresponding to the maximum confidence level and the true category, directly extract the 2048-dimensional deep features before the fully connected layer of the network, and use the PCA dimensionality reduction method to reduce the 2048-dimensional features to 256-dimensional features and save them.

[0084] After the test, in addition to the classification model, a feature database of misclassified images is also obtained. The database contains the features of the misclassified images and their true corresponding categories. In actual use, misclassified data can also be added manually to expand the feature database.

[0085] When actually using the model, input an image into the resnet50 network, which will output the confidence levels belonging to the eight categories. At the same time, it will output the features reduced to 256 dimensions.

[0086] Step S5:

[0087] Denote the 256-dimensional features of the image to be classified as f, and denote the feature set in the feature database as F, F = {f'1, f'2, ……, f'n}, where f' is an element in the feature set F, and n is the number of features in the feature set. Calculate the similarity between f and f'1 to f'n one by one, and the calculation formula is

[0088] sim = f × f' / (norm(f) × norm(f'))

[0089] The calculation formula of the norm() function is

[0090]

[0091] where V is an element in the feature f.

[0092] According to the above method, calculate the similarity values between f and each element in the set F respectively. At the same time, during the calculation process, retain the numerical values with the highest calculated similarity in each of the eight categories for future use.

[0093] Step S6:

[0094] First, obtain the classification result of the current face region and the similarity calculation result according to Step S4 and Step S5. The eight classification results and their corresponding confidence levels are denoted as: [(c1,t1),(c2,t2),(c3,t3), (c4,t4), (c5,t5), (c6,t6), (c7,t7), (c8,t8)], where c1 represents category 1, t1 represents the confidence level of category 1, and so on. Based on the feature similarity calculation, also obtain the eight highest similarities, which are respectively denoted as: [(c1,T1),(c2,T2),(c3,T3), (c4,T4), (c5,T5), (c6,T6), (c7,T7), (c8,T8)]. Similarly, c1 represents category 1, and T1 represents the similarity to the closest category 1 in the database. Based on the above results, multiply the confidence level and similarity of the corresponding category, and take the category corresponding to the largest product among the eight categories as the category of the current frame.

[0095] The calculation formula is as follows:

[0096] max(w×t i +(1 - w)×T i )

[0097] where the max() function is the function to take the maximum value, t i is the confidence level of the corresponding category, T i is the similarity to the closest corresponding category in the feature set, and w is the weighting weight, with a range of (0, 1).

[0098] Preferably, in this embodiment, the weighting weight w takes a value of 0.7.

[0099] Secondly, multi-frame target matching is performed. Since there may be multiple face regions to be classified in the surveillance video, when multi-frame decision-making is required, target matching for the corresponding face regions needs to be carried out first. Since the surveillance camera is fixed, the method of calculating IOU is used for target matching, that is, the rectangular coordinate frames of all face regions in the current frame are recorded. When the next frame is input, multiple rectangular frames will also be obtained. Then, the rectangular frames in the current frame are traversed respectively with those in the previous frame. During the traversal, the overlapping area and the combined area of the two rectangular frames are calculated, and the IOU value is calculated by dividing the overlapping area by the combined area. After one round of traversal, if the IOU value is greater than the set threshold, the matching is successful; otherwise, the matching fails. The face regions with failed matching are regarded as new targets and are prepared to perform the same calculation with the next frame of data.

[0100] Finally, for the face regions that are continuously matched for several frames, if two-thirds of the categories in several frames are the same category, then the category C is output; otherwise, the invalid category is output. In this embodiment, the number of continuously matched frames is 3 frames, that is, if 2 out of 3 frames are of the same category, then the category is output.

[0101] Step S7:

[0102] The final result is output according to the output category C, that is, if the output category C belongs to category one, two, or three, then the smoking behavior is output; if category C belongs to category four or five, then the calling behavior is output; if the output category C belongs to category six, seven, or eight, then the normal behavior is output.

[0103] The recognition method for calling or smoking behavior in a special scenario provided by the present invention can greatly improve the algorithm accuracy in the case of limited data volume, and further improve the algorithm accuracy by constructing a feature library for misclassified cases, and finally make a decision through the method of multi-frame target tracking and matching, reducing the false alarm rate.

[0104] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art of this patent, without departing from the scope of the technical solution of the present invention, can make some changes or modifications to the above-mentioned disclosed technical content to obtain equivalent embodiments with equivalent changes. The implementation schemes in the above embodiments can also be further combined or replaced. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A method for identifying phone call or smoking behaviors in special scenarios, characterized in that: The specific steps include: S1. Collect real-time monitoring images; S2. Perform face detection to determine potential ROI regions; S3. Correct the image to generate an image to be classified; S4. Divide the behaviors into several categories, collect training data for each category for deep learning training to obtain a multi-classification model, input the image to be classified into the classification model, and output the confidence values corresponding to each category and the features of the image to be classified; S5. Calculate the similarity between the features of the image to be classified and each element in the feature set, and save the value with the highest similarity calculation in each category; S6. Multiply the confidence value of the corresponding category by the similarity value, and take the category with the maximum product value as the final category of the current frame. If two-thirds of the images in a continuous number of frames have the same final category, output the final category; S7. According to the behavior corresponding to the final category, output the result of whether making a call or smoking.

2. The method for identifying a calling or smoking behavior in a special scenario according to claim 1, wherein: In step S2, the centerface network is used for face detection. Specifically, first, the video frame output in step S1 is scaled to the image size supported by the centerface network, and then the positions of all faces in the image and five key points of each face are obtained through convolution, pooling, or activation. The five key points are the left eye pupil point, the right eye pupil point, the nose tip point, the left mouth corner point, and the right mouth corner point.

3. The method for identifying a calling or smoking behavior in a special scenario according to claim 1, wherein: The specific steps for image correction to generate an image to be classified in step S3 are as follows: S31. Take the set of key points to be corrected P = {p1, p2, p3, p4, p5}, where p1, p2, p3, p4, and p5 are the two-dimensional coordinates (xi, yi) of the key points to be corrected, and i belongs to [1, 5]. Let the set of standard key points S = {s1, s2, s3, s4, s5}, where s1, s2, s3, s4, and s5 are the two-dimensional coordinates (xi, yi) of the standard key points, and i belongs to [1, 5]; S32. Calculate the standard deviations of set P and set S respectively; S33. Divide the elements of set P and set S by their corresponding standard deviations to obtain new sets P' and S'; S34. Multiply the two new sets P' and S' to obtain matrix M, perform SVD decomposition on matrix M, combine to obtain a rotation matrix and a parallel matrix, and intercept an image with a size of 224×224 pixels centered on the converted nose tip point as the image to be classified.

4. A method for identifying calling or smoking behaviors in a special scenario according to claim 1, characterized in that: In step S4, the specific number of categories for dividing the behaviors into several categories is 8, namely Category 1: Holding a cigarette between fingers, Category 2: Holding a cigarette in the mouth, Category 3: Smoking with a cigarette in the mouth and holding a cigarette between fingers, Category 4: Holding a phone in hand, Category 5: Holding a phone to the ear, Category 6: Touching the mouth with hand, Category 7: Touching the ear with hand, Category 8: Others.

5. A method for identifying phone call or smoking behaviors in a special scenario according to claim 1, characterized in that: In step S4, the resnet50 network is used for deep learning training. The image to be classified is input into the resnet50 network, and while outputting the confidence of each classification, a 256-dimensional feature after dimensionality reduction is also output.

6. The method for identifying a calling or smoking behavior in a special scenario according to claim 1, wherein: When the maximum confidence class output by the training data in step S4 does not match the true class, extract the 2048-dimensional deep features before the fully connected layer of the network. Use the PCA dimensionality reduction method to reduce the 2048-dimensional features to 256-dimensional features, and save them as the feature database for misclassification.

7. A method for identifying phone call or smoking behaviors in a special scenario according to claim 1, characterized in that: In step 5, the similarity calculation method is as follows: Denote the feature of the image to be classified as f, and the set of features in the feature database as F, F = {f’1, f’2, ……, f’n}, where f’ is an element in the feature set F, and n is the number of features in the feature set. Calculate the similarity between f and f’1 to f’n one by one. The calculation formula is sim = f × f′ / (norm(f) × norm(f’)) The calculation formula of the norm() function is where V is an element in the feature f.

8. A method for identifying calling or smoking behaviors in a special scenario according to claim 1, characterized in that: In step S6, the specific method of "multiplying the confidence value of the corresponding class by the similarity value and taking the class with the maximum product as the final class of the current frame" is Calculate according to the following formula max(w×t i +(1 - w)×T i ) where the max() function is the maximum value function, and t i is the confidence of the corresponding category, and T i is the similarity to the corresponding category that is closest in the feature set, and w is the weighting weight, with a range of (0, 1).

9. A method for identifying calling or smoking behaviors in a special scenario according to claim 8, characterized in that: The weighted weight w takes 0.

7.

10. A method for identifying calling or smoking behaviors in a special scenario according to claim 1, characterized in that: In step S6, the matching method for "consecutive frames of images" is to use the IOU value calculation for object matching. Specifically, record the rectangular coordinate frames of all face regions in the current frame. When the next frame is input, multiple rectangular frames will also be obtained. Then, use the rectangular frames of the current frame to traverse the rectangular frames of the previous frame one by one. During the traversal, calculate the overlapping area and the combined area of the two rectangular frames, and calculate the IOU value by dividing the overlapping area by the combined area. After one round of traversal, if the IOU value is greater than the set threshold, the matching is successful; otherwise, the matching fails. The face regions with failed matching are used as new targets and are ready to perform the same calculation with the next frame of data.

Citation Information

Patent Citations

  • Method for identifying smoking and calling in specific scene

    CN112836643A

  • Deep-learning and cloud service-based face identification attendance system and method

    CN106204780A

  • Liveness detection

    US20210334570A1