Method and device for detecting mobile phone playing behavior of personnel in bright kitchen and bright stove scene
By detecting human key points and target objects in single-frame images under open kitchen scenarios, and combining deep learning and small object detection algorithms, the system can identify mobile phone use behavior, solving the problems of insufficient accuracy and timeliness in existing technologies, and achieving efficient mobile phone use behavior recognition.
Patent Information
- Application Number
- CN202211034251.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-26
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-08-26
AI Technical Summary
Existing intelligent monitoring methods struggle to simultaneously guarantee the accuracy and timeliness of identifying people using mobile phones in open kitchen scenarios. Single-frame image analysis is susceptible to background interference, leading to false alarms and missed alarms, while continuous multi-frame image analysis consumes significant computing resources and incurs latency.
We use single-frame images to detect key points of the human body and target objects. Combining human skeletal joint information and target object information, we determine whether the behavior is playing with a mobile phone by considering distance, similarity, and deviation. We use a bottom-up algorithm and a small object detection algorithm to run in parallel, and combine deep learning algorithms for pose estimation.
It improved system operating efficiency, enhanced the accuracy and timeliness of recognizing mobile phone use, especially in overhead and side-view shots, and reduced computing power consumption and latency.
Smart Images

Figure CN115424205B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine vision, and in particular to a personnel mobile phone playing behavior detection method and device in a bright kitchen and bright stove scene. BACKGROUND
[0002] With the rapid development of information technology, people's dependence on mobile phones is also becoming more and more serious. Due to playing mobile phones, negative shirking, and even causing accidents are not uncommon. For example, during the cooking process of a chef, the chef may play mobile phones in violation of regulations, forget to turn off the gas stove and other fire sources, and cause fire accidents and other safety accidents. In addition, playing mobile phones on duty may also lead to reduced work efficiency and affect food safety and hygiene. Therefore, in the bright kitchen and bright stove scene, shirking and safety risks caused by playing mobile phones need to be highly valued. If personnel are monitored by means of manpower supervision, the cost of manpower will be increased, and therefore, intelligent monitoring of target behaviors such as personnel playing mobile phones is needed.
[0003] At present, intelligent monitoring of target behaviors has the following disadvantages: if single-frame images are used for analysis, due to objective factors such as target background interference, it is difficult to extract the mobile phone playing behavior features, which may easily lead to false positives and false negatives; if continuous multiple frames of images are used for analysis, a large amount of computing resources will be consumed, and the analysis result will have a time delay.
[0004] Therefore, the existing intelligent monitoring method cannot guarantee both the accuracy and timeliness of target behavior recognition. SUMMARY
[0005] The present application provides a personnel mobile phone playing behavior detection method and device in a bright kitchen and bright stove scene, which can guarantee both the accuracy and timeliness of mobile phone playing behavior recognition.
[0006] The present application provides a personnel mobile phone playing behavior detection method in a bright kitchen and bright stove scene, comprising:
[0007] acquiring a single-frame image, detecting human key points and target objects in the single-frame image;
[0008] in a case where human key points and target objects are detected in the single-frame image, based on human skeleton joint point information and target object information extracted from the single-frame image, determining a distance between a face and the target object, a similarity between a human posture and a target posture, and a deviation degree between a face orientation and the target object;
[0009] based on the distance, the similarity and the deviation degree, determining whether the personnel behavior belongs to a target behavior;
[0010] The target object is a mobile phone, the target posture is a posture corresponding to a target behavior, and the target behavior is a mobile phone playing behavior of a person in a bright kitchen and bright stove scene.
[0011] The method for detecting the mobile phone playing behavior of the person in the bright kitchen and bright stove scene comprises the following steps of:
[0012] inputting the single frame image into a human body detection model to detect a human body key point, obtaining a first detection result output by the human body detection model, and determining whether a human body skeleton key point exists in the single frame image based on the first detection result;
[0013] inputting the single frame image into an object detection model to detect a target object, obtaining a second detection result output by the object detection model, and determining whether the target object exists in the single frame image based on the second detection result;
[0014] The human body detection model is trained based on a bottom-up algorithm and a graph theory algorithm, and the object detection model is trained based on a small target detection algorithm.
[0015] The human body detection model and the object detection model are respectively deployed on multiple processors and run in parallel.
[0016] The small target detection algorithm is a single-stage target detection algorithm.
[0017] The method for detecting the mobile phone playing behavior of the person in the bright kitchen and bright stove scene comprises the following steps of:
[0018] inputting the human body skeleton joint point information and the target object information into a trained human body posture estimation model, obtaining a distance between a face and the target object, a similarity between a human body posture and a target posture, and a deviation degree between a face orientation and the target object output by the human body posture estimation model;
[0019] The human body posture estimation model is trained based on a deep learning algorithm.
[0020] The personnel mobile phone playing behavior detection method in the kitchen and stove bright scene provided by the application, the distance, the similarity and the deviation degree are used to determine whether the personnel behavior is the target behavior, including:
[0021] The distance, the similarity and the deviation degree are used to determine the probability that the human body posture is the target posture.
[0022] When the probability is greater than a preset threshold, it is determined that the personnel behavior is the target behavior.
[0023] The personnel mobile phone playing behavior detection method in the kitchen and stove bright scene provided by the application, the distance, the similarity and the deviation degree are used to determine whether the personnel behavior is the target behavior, including:
[0024] The video data is decoded and frame extracted to obtain the single frame image.
[0025] The application further provides a personnel mobile phone playing behavior detection device in a kitchen and stove bright scene, including:
[0026] The first detection module is used to acquire a single frame image and detect human body key points and target objects in the single frame image.
[0027] The second detection module is used to determine the distance between the face and the target object, the similarity between the human body posture and the target posture, and the deviation degree between the face orientation and the target object based on the human body skeleton joint point information and the target object information extracted from the single frame image when the human body key points and the target objects are detected in the single frame image.
[0028] The judgment module is used to determine whether the personnel behavior is the target behavior based on the distance, the similarity and the deviation degree.
[0029] The target object is a mobile phone, the target posture is the posture corresponding to the target behavior, and the target behavior is the personnel mobile phone playing behavior in the kitchen and stove bright scene.
[0030] The application further provides a server including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to realize the personnel mobile phone playing behavior detection method in the kitchen and stove bright scene according to any one of the above.
[0031] The application further provides an electronic device including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to realize the personnel mobile phone playing behavior detection method in the kitchen and stove bright scene according to any one of the above.
[0032] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the personnel mobile phone playing behavior detection method in the bright kitchen and bright stove scene.
[0033] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the personnel mobile phone playing behavior detection method in the bright kitchen and bright stove scene.
[0034] The personnel mobile phone playing behavior detection method and device in the bright kitchen and bright stove scene provided by the application detect human key points and target objects in a single frame of image, determine the distance between a face and a target object, the similarity of a human posture and a target posture, and the deviation degree of a face orientation from a target object based on human skeleton joint point information and target object information, and determine whether the personnel behavior is a mobile phone playing behavior. Compared with the image classification algorithm used in the prior art to process images, the application detects human key points and target objects in a single frame of image, improves the system running efficiency, determines whether the personnel behavior is a mobile phone playing behavior in combination with the distance, the similarity, and the deviation degree, and solves the problems of poor generalization performance and poor robustness of the traditional artificial feature design method. The detection of human key points (such as face key points, hand key points, and torso key points) improves the recognition accuracy in the case of a downward shot and a side face. Therefore, the application can not only ensure the accuracy of mobile phone playing behavior recognition, but also ensure the timeliness of mobile phone playing behavior recognition. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the application or the prior art, the following will briefly introduce the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0036] Figure 1 is one of the flowcharts of the personnel mobile phone playing behavior detection method in the bright kitchen and bright stove scene provided by the application;
[0037] Figure 2 is the second flowchart of the personnel mobile phone playing behavior detection method in the bright kitchen and bright stove scene provided by the application;
[0038] Figure 3 is the flowchart of personnel key point detection provided by the application;
[0039] Figure 4 is the functional module schematic diagram of the personnel mobile phone playing behavior detection method in the bright kitchen and bright stove scene provided by the application;
[0040] Figure 5 This is a schematic diagram of the device for detecting people using mobile phones in a transparent kitchen scenario provided by the present invention.
[0041] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0043] The following is combined Figures 1-6 This invention describes a method and apparatus for detecting people using mobile phones in a transparent kitchen setting.
[0044] like Figure 1 As shown, the method for detecting people using mobile phones in a transparent kitchen scenario provided by the present invention includes:
[0045] Step 110: Acquire a single-frame image and detect key human body points and target objects in the single-frame image.
[0046] It is understandable that a single frame image can be acquired by capturing a picture of a target area using a camera device. The target area could be a scene of an open kitchen, and the target object could be a mobile phone device.
[0047] By detecting key human body points and target objects in a single frame of image, rather than detecting multiple consecutive frames, excessive computing resources can be avoided, latency can be reduced, and the timeliness of detection can be guaranteed.
[0048] Step 120: If key human body points and target objects are detected in the single frame image, based on the human skeletal joint point information and target object information extracted from the single frame image, determine the distance between the face and the target object, the similarity between the human posture and the target posture, and the degree of deviation between the face orientation and the target object.
[0049] Understandably, the human behavior detection process ends when no key human body points or target objects are detected in a single frame image.
[0050] The distance between the face and the target object, the similarity between the human body posture and the target posture, and the deviation degree of the face orientation from the target object can be determined based on the human body skeleton joint position information and the target object information extracted from the single frame image in a deep learning and big data calculation manner.
[0051] The human body key points include a plurality of preset human body skeleton key points of a torso, a plurality of preset human body skeleton key points of a face, and a plurality of preset human body skeleton key points of a human hand.
[0052] In step 130, whether the personnel behavior belongs to the target behavior is determined based on the distance, the similarity, and the deviation degree.
[0053] The target object is a mobile phone, the target posture is a posture corresponding to the target behavior, and the target behavior is a personnel mobile phone playing behavior in a bright kitchen and bright kitchen scenario.
[0054] It can be understood that when determining whether the personnel behavior belongs to the target behavior based on the distance, the similarity, and the deviation degree, the corresponding image is intercepted and retained, and an alarm is issued.
[0055] In some embodiments, the detection of the human body key points and the target object in the single frame image includes:
[0056] The single frame image is input into a human body detection model for human body key point detection to obtain a first detection result output by the human body detection model, so as to determine whether there is a human body skeleton key point in the single frame image based on the first detection result.
[0057] The single frame image is input into an object detection model for target object detection to obtain a second detection result output by the object detection model, so as to determine whether there is a target object in the single frame image based on the second detection result.
[0058] The human body detection model is trained based on a bottom-up algorithm and a graph theory algorithm, and the object detection model is trained based on a small target detection algorithm.
[0059] It can be understood that the human body detection model is based on a bottom-up method to identify human skeleton key points. Unlike the traditional top-down method (first identify the human target, and then find the key point position). The human body detection model uses the part association field technology to divide the neural network framework into two paths. One path uses a convolutional neural network to predict key points based on a confidence map. The other path also uses a convolutional neural network to obtain the part association field of each key point (which can be regarded as a two-dimensional vector recording the position and direction of the limb). The two paths are jointly learned and predicted. Finally, the key points are fused into the human body trunk, that is, these key points are connected to each other in the best way according to the graph theory method. The object detection model can be optimized for small target feature extraction. In some embodiments, the human body detection model and the object detection model are respectively deployed on multiple processors and run in parallel.
[0060] It can be understood that the human body detection model and the object detection model are respectively deployed on multiple processors and run in parallel, which can improve the detection efficiency of the human body detection model and the object detection model, and can simultaneously detect and process multiple single-frame images.
[0061] In some embodiments, the small target detection algorithm is a single-stage target detection algorithm.
[0062] It can be understood that the single-stage target detection algorithm includes but is not limited to the YOLO (You Only Look Once) series algorithm and the SSD (Single Shot MultiBox Detector) series algorithm.
[0063] The YOLO series algorithm is a target detection model algorithm. The YOLO series algorithm does not need to find the region where the target may exist in advance. The YOLO series algorithm uniformly samples different positions on the picture, uses different scales and aspect ratio clipping boxes during sampling, and then directly classifies and regresses using a convolutional neural network to extract features. The whole process only needs one step, and the advantage is fast speed.
[0064] The SSD series algorithm uses multi-scale feature maps to directly regress the target class and position.
[0065] In some embodiments, based on the human skeleton joint position information and target object information extracted from the single-frame image, the distance between the face and the target object, the similarity of the human body posture and the target posture, and the deviation degree between the face orientation and the target object are determined, including:
[0066] input the human body joint position information and the target object information into the human body posture estimation model trained to obtain a distance between the human face and the target object, a similarity between the human body posture and the target posture, and a deviation degree between the human face orientation and the target object output by the human body posture estimation model;
[0067] The human body posture estimation model is trained based on a deep learning algorithm.
[0068] It can be understood that the prediction of human body behavior by the human body posture estimation model, the distance between the human face and the target object, the similarity between the human body posture and the target posture, and the deviation degree between the human face orientation and the target object output by the human body posture estimation model are based on a deep learning algorithm network model trained based on a large amount of data, rather than based on artificially set feature values.
[0069] The deep learning algorithm can be a convolution product network, and the human body posture estimation model trained based on the deep learning algorithm can accurately and efficiently obtain the distance between the human face and the target object, the similarity between the human body posture and the target posture, and the deviation degree between the human face orientation and the target object.
[0070] The human body posture estimation model is based on a deep learning algorithm for human body posture calculation, and the features are adaptively learned through the network without artificial setting.
[0071] In some embodiments, the determining whether the personnel behavior belongs to the target behavior based on the distance, the similarity, and the deviation degree includes:
[0072] determining a probability that the human body posture is the target posture based on the distance, the similarity, and the deviation degree.
[0073] In a case where the probability is greater than a preset threshold, it is determined that the personnel behavior belongs to the target behavior.
[0074] It can be understood that the threshold value can be set according to the target behavior to be recognized. The target behavior can be a mobile phone playing behavior, and the target posture is a posture corresponding to the mobile phone playing behavior. When the probability that the human body posture is the target posture is greater than a preset threshold, it is determined that the corresponding personnel is playing a mobile phone.
[0075] In some embodiments, the obtaining a single frame image includes:
[0076] obtaining video data, decoding and frame extracting the video data to obtain the single frame image.
[0077] It can be understood that the data obtained by the camera device in the target region space is video data, and the video data is continuous multiple frame images. Therefore, the video data needs to be decoded and frame extracted at a preset frequency to obtain a single frame image.
[0078] In some embodiments, the personnel playing mobile phone behavior detection method in the bright kitchen and bright stove scene is implemented by the function modules as shown in Figure 2 The personnel playing mobile phone behavior detection method in the bright kitchen and bright stove scene is implemented by the function modules as shown in
[0079] In some embodiments, the personnel playing mobile phone behavior detection method in the bright kitchen and bright stove scene is implemented by the function modules as shown in Figure 3 The personnel playing mobile phone behavior detection method in the bright kitchen and bright stove scene is implemented by the function modules as shown in
[0080] In some embodiments, the personnel playing mobile phone behavior detection method in the bright kitchen and bright stove scene is implemented by the function modules as shown in Figure 4 The personnel playing mobile phone behavior detection method in the bright kitchen and bright stove scene is implemented by the function modules as shown in
[0081] In summary, the personnel mobile phone playing behavior detection method in the bright kitchen and bright stove scene provided by the application comprises: acquiring a single frame image, detecting human key points and target objects in the single frame image; in the case that human key points and target objects are detected in the single frame image, based on human skeleton joint point information and target object information extracted from the single frame image, determining the distance between a face and the target object, the similarity of human posture and target posture, and the deviation degree between the face orientation and the target object; based on the distance, the similarity and the deviation degree, determining whether the personnel behavior is a target behavior; wherein the target object is a mobile phone, the target posture is a posture corresponding to the target behavior, and the target behavior is personnel mobile phone playing behavior in the bright kitchen and bright stove scene.
[0082] In the personnel mobile phone playing behavior detection method in the bright kitchen and bright stove scene provided by the application, human key points and target objects in a single frame image are detected, the distance between a face and a target object, the similarity of human posture and target posture, and the deviation degree between the face orientation and the target object are determined based on human skeleton joint point information and target object information, and it is determined whether the personnel behavior is mobile phone playing behavior. Compared with the image classification algorithm used in the prior art to process images, the application detects human key points and target objects in a single frame image, improves the system running efficiency, and determines whether the personnel behavior is mobile phone playing behavior in combination with the distance, the similarity and the deviation degree, thereby solving the problems of poor generalization performance and poor robustness of the traditional artificial feature design method. The human key points (such as face key points, hand key points and torso key points) are detected, and the recognition accuracy of situations such as a downward shot and a side face is improved. Therefore, the application can not only ensure the accuracy of mobile phone playing behavior recognition, but also ensure the timeliness of mobile phone playing behavior recognition.
[0083] Further, the application uses human posture estimation and small target detection methods, which can run in parallel to infer, and through the fusion of detection information of the two technical architectures, it is comprehensively judged whether mobile phone playing behavior occurs, and the obtained result is more reliable and has more advantages in accuracy and running efficiency.
[0084] The application can optimize the model for the kitchen and refrigeration room in the bright kitchen and bright stove, has high recognition accuracy in this scene, and can also be trained for other scenes to meet different needs.
[0085] The personnel mobile phone playing behavior detection device in the bright kitchen and bright stove scene provided by the application is described below, and the personnel mobile phone playing behavior detection device described below can be correspondingly referred to the personnel mobile phone playing behavior detection method described above.
[0086] AsFigure 5 As shown, the present invention also provides a device 500 for detecting people using mobile phones in a transparent kitchen setting, comprising:
[0087] The first detection module 510 is used to acquire a single frame image and detect key human body points and target objects in the single frame image.
[0088] The second detection module 520 is used to determine the distance between the face and the target object, the similarity between the human posture and the target posture, and the degree of deviation between the face orientation and the target object based on the human skeletal joint point information and target object information extracted from the single frame image when the presence of human key points and target objects is detected in the single frame image.
[0089] The judgment module 530 is used to determine whether a person's behavior belongs to the target behavior based on the distance, the similarity, and the degree of deviation.
[0090] The target object is a mobile phone, the target posture is the posture corresponding to the target behavior, and the target behavior is the behavior of people playing with mobile phones in a transparent kitchen scenario.
[0091] The present invention also provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method for detecting people using mobile phones in a transparent kitchen scenario as described above.
[0092] The electronic device, computer program product, and storage medium provided by the present invention are described below. The electronic device, computer program product, and storage medium described below can be referred to in correspondence with the method for detecting people playing mobile phones in the open kitchen scenario described above.
[0093] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a method for detecting people using mobile phones in a transparent kitchen scenario. This method includes:
[0094] Acquire a single-frame image and detect key human body points and target objects in the single-frame image;
[0095] In a case where it is detected that the single-frame image contains human key points and a target object, based on human skeleton joint point information and target object information extracted from the single-frame image, a distance between a face and the target object, a similarity between a human posture and a target posture, and a deviation degree between a face orientation and the target object are determined.
[0096] Based on the distance, the similarity and the deviation degree, it is determined whether the personnel behavior belongs to a target behavior.
[0097] The target object is a mobile phone, the target posture is a posture corresponding to the target behavior, and the target behavior is a personnel mobile phone playing behavior in a bright kitchen and bright kitchen scenario.
[0098] In addition, the logic instructions in the memory 630 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0099] On the other hand, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and when the computer program is executed by a processor, the computer can execute the personnel mobile phone playing behavior detection method in a bright kitchen and bright kitchen scenario provided by the above-mentioned methods. The method comprises:
[0100] Obtaining a single-frame image, detecting human key points and a target object in the single-frame image;
[0101] In a case where it is detected that the single-frame image contains human key points and a target object, based on human skeleton joint point information and target object information extracted from the single-frame image, a distance between a face and the target object, a similarity between a human posture and a target posture, and a deviation degree between a face orientation and the target object are determined.
[0102] Based on the distance, the similarity and the deviation degree, it is determined whether the personnel behavior belongs to a target behavior.
[0103] The target object is a mobile phone, the target posture is a posture corresponding to the target behavior, and the target behavior is a mobile phone playing behavior of a person in a bright kitchen and bright stove scene.
[0104] In another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method for detecting the mobile phone playing behavior of the person in the bright kitchen and bright stove scene provided by the above method, the method comprising:
[0105] obtaining a single frame image, and detecting a human key point and a target object in the single frame image;
[0106] In a case where the human key point and the target object are detected in the single frame image, determining a distance between a face and the target object, a similarity between a human posture and a target posture, and a deviation degree between a face orientation and the target object based on human skeleton joint point information and target object information extracted from the single frame image;
[0107] Determining whether the person behavior belongs to a target behavior based on the distance, the similarity, and the deviation degree;
[0108] The target object is a mobile phone, the target posture is a posture corresponding to the target behavior, and the target behavior is a mobile phone playing behavior of a person in a bright kitchen and bright stove scene.
[0109] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on a plurality of network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.
[0110] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0111] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting mobile phone use by people in a transparent kitchen setting, characterized in that, include: Acquire a single-frame image and detect key human body points and target objects in the single-frame image; If key human body points and target objects are detected in the single frame image, the distance between the face and the target object, the similarity between the human body posture and the target posture, and the degree of deviation between the face orientation and the target object are determined based on the human body skeletal joint point information and target object information extracted from the single frame image. Based on the distance, the similarity, and the degree of deviation, it is determined whether the person's behavior belongs to the target behavior; Wherein, the target object is a mobile phone, the target posture is the posture corresponding to the target behavior, and the target behavior is the behavior of people playing with mobile phones in a transparent kitchen scenario; The detection of key human body points and target objects in the single frame image includes: The single-frame image is input into a human detection model to detect key points of the human body, and a first detection result is obtained from the human detection model. Based on the first detection result, it is determined whether there are human skeletal key points in the single-frame image. The single-frame image is input into an object detection model to detect target objects, and a second detection result is obtained from the object detection model. Based on the second detection result, it is determined whether there are target objects in the single-frame image. The human detection model is trained based on a bottom-up algorithm and a graph theory algorithm, and the object detection model is trained based on a small object detection algorithm. Both the human detection model and the object detection model are deployed on multiple processors and run in parallel.
2. The method for detecting people using mobile phones in a transparent kitchen scenario according to claim 1, characterized in that, The small target detection algorithm is a single-stage target detection algorithm.
3. The method for detecting people using mobile phones in a transparent kitchen scenario according to claim 1, characterized in that, The step of determining the distance between the face and the target object, the similarity between the human pose and the target pose, and the degree of deviation between the face orientation and the target object based on the human skeletal joint point information and target object information extracted from the single frame image includes: The human skeletal joint point information and the target object information are input into the trained human pose estimation model to obtain the distance between the face and the target object, the similarity between the human pose and the target pose, and the degree of deviation between the face orientation and the target object output by the human pose estimation model. The human pose estimation model is trained based on a deep learning algorithm.
4. The method for detecting people using mobile phones in a transparent kitchen scenario according to claim 1, characterized in that, The process of determining whether a person's behavior belongs to the target behavior based on the distance, the similarity, and the degree of deviation includes: Based on the distance, the similarity, and the degree of deviation, the probability that the human posture is the target posture is determined; If the probability is greater than a preset threshold, the person's behavior is determined to be the target behavior.
5. The method for detecting people using mobile phones in a transparent kitchen scenario according to any one of claims 1-4, characterized in that, The acquisition of a single frame image includes: Acquire video data, decode and extract frames from the video data to obtain the single-frame image.
6. A device for detecting people using mobile phones in a transparent kitchen setting, characterized in that, include: The first detection module is used to acquire a single-frame image and detect key human body points and target objects in the single-frame image; The second detection module is used to determine the distance between the face and the target object, the similarity between the human posture and the target posture, and the degree of deviation between the face orientation and the target object based on the human skeletal joint point information and target object information extracted from the single frame image when the presence of human key points and target objects is detected in the single frame image. The judgment module is used to determine whether a person's behavior belongs to the target behavior based on the distance, the similarity, and the degree of deviation. Wherein, the target object is a mobile phone, the target posture is the posture corresponding to the target behavior, and the target behavior is the behavior of people playing with mobile phones in a transparent kitchen scenario; The detection of key human body points and target objects in the single frame image includes: The single-frame image is input into a human detection model to detect key points of the human body, and a first detection result is obtained from the human detection model. Based on the first detection result, it is determined whether there are human skeletal key points in the single-frame image. The single-frame image is input into an object detection model to detect target objects, and a second detection result is obtained from the object detection model. Based on the second detection result, it is determined whether there are target objects in the single-frame image. The human detection model is trained based on a bottom-up algorithm and a graph theory algorithm, and the object detection model is trained based on a small object detection algorithm. Both the human detection model and the object detection model are deployed on multiple processors and run in parallel.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for detecting people using mobile phones in a transparent kitchen scenario as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for detecting people using mobile phones in a transparent kitchen scenario as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Behavior recognition method and device and computer equipment
CN113065474A