Behavior recognition method and device, electronic device, computer readable storage medium

By analyzing real-time video data, the system detects and tracks target objects, obtains the location of key skeletal points, and solves the problem of identifying abnormal behaviors such as smoking or using mobile phones, thus improving the accuracy and timeliness of behavior detection.

CN115798047BActive Publication Date: 2026-04-10CETC BIGDATA RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CETC BIGDATA RES INST CO LTD
Filing Date
2022-12-06
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Smoking or using mobile phones during work or operations can lead to a shift in attention, potentially resulting in an inability to handle emergencies in a timely manner and impacting production and daily life.

Method used

By collecting video data in real time, the system detects and tracks target objects, obtains the location of key skeletal points of individuals, determines behavior based on this location information, and issues alerts for abnormal behavior.

Benefits of technology

It improves the accuracy of detecting personnel's actions on target objects, ensuring timely identification and early warning of abnormal behavior during working hours, and preventing disruption to work performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798047B_ABST
    Figure CN115798047B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a behavior recognition method, and the specific implementation scheme is as follows: obtaining an image frame to be recognized based on real-time collected or shot video data to be recognized; detecting whether there is a target object in the image frame to be recognized in real time; in response to detecting that the image frame to be recognized has the target object, tracking the target object to obtain position information of the target object; obtaining a target skeleton key point position of a person in the image frame to be recognized; and determining a behavior of the person on the target object based on the target skeleton key point position and the position information of the target object. Through the embodiment, the accuracy of behavior detection when the person operates the target object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of computer, and in particular, to a behavior recognition method and device. BACKGROUND

[0002] With the development of mobile communication technology, mobile phones have become an important tool for modern people. In different work or work fields (for example, business hall, factory), smoking and playing mobile phone behavior is one of the behaviors that often occur in the process of personnel work or work. Once this behavior occurs, personnel can hardly ensure that they are in a normal working state, and due to the diversion of attention, unexpected events that cannot be handled in time are likely to occur, affecting normal production and life. SUMMARY

[0003] Embodiments described herein provide a behavior recognition method and device, an electronic device, and a computer readable storage medium having a computer program stored therein.

[0004] According to a first aspect of the present disclosure, a behavior recognition method is provided. In the method, based on real-time collected or shot video data to be recognized, an image frame to be recognized is obtained; whether a target object is present in the image frame to be recognized is detected in real time; in response to detecting that the target object is present in the image frame to be recognized, the target object is tracked to obtain position information of the target object; a target skeletal key point position of a person in the image frame to be recognized is acquired; based on the target skeletal key point position and the position information of the target object, a behavior of the person to the target object is determined.

[0005] In some embodiments of the present disclosure, the above-mentioned acquiring the target skeletal key point position of the person in the image frame to be recognized comprises: extracting a human skeletal key point position in the image frame to be recognized; and identifying a target skeletal key point position in the human skeletal key point position.

[0006] In some embodiments of the present disclosure, the target skeletal key point position comprises a wrist key point position and a nose key point position; the position information of the target object comprises a position of the target object and a target object identifier; and based on the target skeletal key point position and the position information of the target object, determining the behavior of the person to the target object comprises: for each target object under a target object identifier, detecting whether a distance from the wrist key point position to the position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detecting whether a distance from the position of the target object to the nose key point position is less than a second distance threshold, the first distance threshold being less than the second distance threshold; and in response to the distance from the position of the target object to the nose key point position being less than the second distance threshold, determining an abnormal behavior of the person to the target object.

[0007] In some embodiments of the present disclosure, the target skeletal key point position includes a wrist key point position and a neck key point position, the position information of the target object includes a position of the target object and an object identifier, and the behavior of the person to the target object is determined based on the target skeletal key point position and the position information of the target object, including: for each target object under an object identifier, detecting whether a distance from the wrist key point position to the position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detecting whether a distance from the position of the target object to the neck key point position is less than a third distance threshold, the first distance threshold being less than the third distance threshold; and in response to the distance from the position of the target object to the neck key point position being less than the third distance threshold, determining that the person has an abnormal behavior to the target object.

[0008] In some embodiments of the present disclosure, the method further includes labeling the person with an abnormal behavior and issuing warning information that the person has an abnormal behavior.

[0009] In some embodiments of the present disclosure, the real-time detection of whether the target object is present in the to-be-recognized image frame includes: real-time detection of the target object in the to-be-recognized image frame by using a target detection model of the target object; and the target detection model is a model obtained by adjusting a basic model.

[0010] In some embodiments of the present disclosure, the target detection model is obtained by the following steps: collecting sample video data in which different persons operate the target object; obtaining sample image data by extracting the sample video data at a preset frame rate interval; obtaining image samples by sample labeling of the target object in the sample image data by using an image labeling tool; and obtaining the target detection model by training the basic model based on the image samples.

[0011] According to a second aspect of the present disclosure, a behavior recognition device is provided. The device includes: an obtaining unit configured to obtain to-be-recognized image frames based on to-be-recognized video data collected or captured in real time; a detection unit configured to real-time detect whether a target object is present in the to-be-recognized image frames; a tracking unit configured to track the target object in response to detection of the presence of the target object in the to-be-recognized image frames, and obtain position information of the target object; an acquisition unit configured to acquire target skeletal key point positions of a person in the to-be-recognized image frames; and a determination unit configured to determine a behavior of the person to the target object based on the target skeletal key point positions and the position information of the target object.

[0012] In some embodiments of the present disclosure, the acquisition unit is further configured to: extract human skeletal key point positions in the to-be-recognized image frames; and identify target skeletal key point positions in the human skeletal key point positions.

[0013] In some embodiments of the present disclosure, the target skeleton key point position includes a wrist key point position and a nose key point position, the position information of the target object includes a position of the target object and a target object identifier, and the determination unit is further configured to: for each target object under a target object identifier, detect whether a distance from the wrist key point position to the position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detect whether a distance from the position of the target object to the nose key point position is less than a second distance threshold, the first distance threshold being less than the second distance threshold; and in response to the distance from the position of the target object to the nose key point position being less than the second distance threshold, determine that the person has an abnormal behavior on the target object.

[0014] In some embodiments of the present disclosure, the target skeleton key point position includes a wrist key point position and a neck key point position, the position information of the target object includes a position of the target object and a target object identifier, and the determination unit is further configured to: for each target object under a target object identifier, detect whether a distance from the wrist key point position to the position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detect whether a distance from the position of the target object to the neck key point position is less than a third distance threshold, the first distance threshold being less than the third distance threshold; and in response to the distance from the position of the target object to the neck key point position being less than the third distance threshold, determine that the person has an abnormal behavior on the target object.

[0015] In some embodiments of the present disclosure, the device further includes an alarm unit configured to label the person as having an abnormal behavior and issue warning information that the person has an abnormal behavior.

[0016] In some embodiments of the present disclosure, the detection unit is further configured to: detect the target object in the to-be-identified image frame in real time by using a target detection model of the target object; and the target detection model is a model obtained by adjusting a basic model.

[0017] In some embodiments of the present disclosure, the target detection model is obtained by training in the following steps: collecting sample video data in which different persons operate target objects; obtaining sample image data by extracting the sample video data at a preset frame rate interval; obtaining image samples by sample labeling the target objects in the sample image data by using an image labeling tool; and obtaining the target detection model by training a basic model based on the image samples.

[0018] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing a computer program; wherein the computer program, when executed by the at least one processor, causes the apparatus to perform the steps of the method according to the first aspect of the present disclosure.

[0019] According to a fourth aspect of the present disclosure, a computer-readable storage medium storing a computer program is provided, wherein the computer program, when executed by a processor, implements the steps of the method according to the first aspect of the present disclosure.

[0020] The behavior recognition method and device provided by the present disclosure first obtains a to-be-recognized image frame based on real-time collected or photographed to-be-recognized video data; secondly, whether the to-be-recognized image frame has a target object is detected in real time; thirdly, in response to detecting that the to-be-recognized image frame has the target object, the target object is tracked to obtain position information of the target object; fourthly, a target skeleton key point position of a person in the to-be-recognized image frame is acquired; and finally, based on the target skeleton key point position and the position information of the target object, a behavior of the person to the target object is determined. In this way, after the target object is detected, the target object is tracked and the target skeleton key point position of the person is recognized, and based on the positional relationship between the target skeleton key point and the target object, the behavior of the person to the target object is determined, thereby improving the accuracy of behavior detection when the person operates the target object. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be known that the drawings described below only relate to some embodiments of the present disclosure, rather than limiting the present disclosure, in which:

[0022] Figure 1 is a flowchart of one embodiment of the behavior recognition method according to the present disclosure;

[0023] Figure 2 is a structural schematic diagram of a target skeleton key point in an embodiment of the present disclosure;

[0024] Figure 3 is a flowchart of another embodiment of the behavior recognition method according to the present disclosure;

[0025] Figure 4 is a structural schematic diagram of one embodiment of the behavior recognition device according to the present disclosure; and

[0026] Figure 5 is a block diagram of an electronic device for implementing the behavior recognition method of the embodiments of the present disclosure. 5DETAILED DESCRIPTION

[0027] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the described...

[0028] All other embodiments of the present disclosure described herein, which can be obtained by those skilled in the art without creative effort, are also within the scope of protection of this disclosure.

[0029] See Figure 1 The diagram illustrates a flow 100 of an embodiment of a behavior recognition method according to the present disclosure, the behavior recognition method comprising the following steps:

[0030] Step 101: Obtain the image frame to be identified based on the real-time acquired or captured video data to be identified.

[0031] In this embodiment, the video data to be identified can be video surveillance data from different scenarios, such as elevators, squares, residential buildings, nursing homes, hospitals, office areas, and other public places.

[0032] The scenarios can also include private settings such as prisons and detention centers. It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of video surveillance data in different scenarios involved in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.

[0033] In this embodiment, the video data to be identified can be read sequentially according to the frame rate interval of the video, and the video data to be identified can be processed by OpenCV (a cross-platform computer vision and machine learning software library) to generate image frames to be identified. Optionally, other video software can also be used to convert the video data to be identified into image frames to be identified.

[0034] Step 102: Real-time detection of whether there is a target object in the image frame to be identified.

[0035] 5. In this embodiment, the target object is an object monitored by the execution entity on which the behavior recognition method runs, and it is also the object that the person in the image frame to be recognized is operating in the current scene or the current time period.

[0036] For example, in an office setting, the target is cigarettes. Smoking is not allowed in an office environment. If a staff member is detected smoking in the office, it indicates abnormal behavior related to cigarettes. Similarly, in the field of prosecution, the target is a mobile device, such as a mobile phone. Using a mobile phone is an operation that prison officers must prohibit during law enforcement. If an officer is detected using a mobile phone, it indicates abnormal behavior related to the phone.

[0037] In this embodiment, based on the unique shape and color features of the target object, it can be determined whether the target object is present in the image frame to be recognized through image recognition technology. For example, the mobile terminal is a cuboid-shaped object with different colors, and the cigarette is a cylindrical object with a white color. In step 103, in response to detecting the target object in the image frame to be recognized, the target object is tracked to obtain the position information of the target object.

[0038] In this embodiment, the target detection and tracking model can be used to track the target object, and the target object in the image frame to be recognized can be uniquely identified and marked. Since the target object in the image frame to be recognized can be multiple types of target objects, for example, there are both cigarettes and mobile terminals in the image frame to be recognized, the target tracking model can obtain the position information of each type of target object, wherein each position information includes the position of the target object and the target object identifier. The target object identifier is the number of different types of mobile phones, for example, the mobile phone marker numbers in different image frames to be recognized are C1, C2, C3, … Cn (n is a natural number greater than 3), and the target tracking model records the unique identifier of the target object and the position of the target object (x c , y c ) of each frame of the image frame to be recognized.

[0039] In this embodiment, the target detection and tracking model includes a target detection module and a target tracking model, wherein the target detection model can use the R-CNN (Region-CNN) model, and the target tracking model can use the deep_sort (deep_sort) model.

[0040] In step 104, the target skeleton key point position of the person in the image frame to be recognized is obtained.

[0041] In this embodiment, the person is another object monitored by the execution subject on which the behavior recognition method is run, in addition to the target object. When there is a target object in the image frame to be recognized, the target object may not be related to the person and may only be located in the image, or the target object may be operated by the person. Specifically, by obtaining the target skeleton key point position of the person, the relationship between the target object and the person can be determined.

[0042] Optionally, in order to more clearly monitor the situation of the person, before obtaining the target skeleton key point position of the person in the image frame to be recognized, the person in the image frame to be recognized can also be identified through recognition technology to confirm that the person is the person monitored by the execution subject on which the behavior recognition method is run. For example, when the person is a duty police officer, the person wearing a uniform in the image frame to be recognized is confirmed through image recognition technology.

[0043] In this embodiment, the image frame to be recognized can have at least one person, and the human skeleton key point position of each person is obtained by performing skeleton key point detection on each person. Specifically, a mature skeleton key point detection algorithm (such as the OpenPose algorithm) can be used, and optionally, the image frame to be recognized can be input into an SPPE (Single-Person Pose Estimator) network to obtain the human skeleton key point position.

[0044] In this embodiment, the target skeleton key point position is the position information corresponding to the target skeleton key point, and the target skeleton key point is the skeleton key point most related to the target object. For example, when the target object is a handheld mobile terminal, the target skeleton key point can be the left eye or right eye key point. When the target object is a cigarette, the target skeleton key point can be the left wrist or right wrist key point.

[0045] In step 105, the behavior of the person on the target object is determined based on the target skeleton key point position and the position information of the target object.

[0046] In this embodiment, after obtaining the target skeleton key point position and the position information of the target object, the abnormal behavior and normal behavior of the person operating the target object can be determined based on the actual positional relationship between the target skeleton key point and the target object.

[0047] Specifically, the above determination of the behavior of the person on the target object based on the target skeleton key point position and the position information of the target object includes: for each target object under each target object identifier, connecting the left eye key point position or the right eye key point position to the position of the target object, and determining the normal behavior of the person operating the target object in response to the length of the connecting line being greater than a preset visual distance (which can be set based on the person, for example, the preset visual distance is 300-500m).

[0048] The behavior recognition method provided by the present disclosure first obtains the image frame to be recognized based on the real-time collected or photographed video data to be recognized. Second, it detects whether there is a target object in the image frame to be recognized. Third, in response to detecting that the image frame to be recognized has a target object, the target object is tracked to obtain the position information of the target object. Fourth, the target skeleton key point position of the person in the image frame to be recognized is obtained. Finally, the behavior of the person on the target object is determined based on the target skeleton key point position and the position information of the target object. Thus, after detecting the target object, the target object is tracked and the target skeleton key point position of the person is recognized, and the behavior of the person on the target object is determined based on the positional relationship between the target skeleton key point and the target object, thereby improving the accuracy of behavior detection when the person operates the target object.

[0049] Optionally, in another embodiment of the present disclosure, the behavior recognition method comprises: obtaining a to-be-recognized image frame based on real-time collected or photographed to-be-recognized video data; determining a time point at which the to-be-recognized image frame is located; detecting whether the time point is a personnel working time; in response to the time point at which the to-be-recognized image frame is located being the personnel working time, real-time detecting whether there is a target object in the to-be-recognized image frame; in response to detecting that there is a target object in the to-be-recognized image frame, tracking the target object to obtain position information of the target object; obtaining a target skeletal key point position of a personnel in the to-be-recognized image frame; determining a behavior of the personnel on the target object based on the target skeletal key point position and the position information of the target object; and in response to the behavior of the personnel on the target object being an abnormal behavior, issuing early warning information that the personnel has the abnormal behavior.

[0050] In the embodiment, the time point can be obtained by collecting the current time, or can be obtained by collecting the time of a device that photographs the to-be-recognized video data.

[0051] In order to effectively exclude temporary viewing operations of the personnel on the target object, optionally, after determining the abnormal behavior of the personnel on the target object, the behavior recognition method can further comprise: starting timing of a time threshold, real-time collecting first video data during the timing, and determining an actual behavior of the personnel on the target object based on the target skeletal key point position and the position information of the target object in the first video data. The actual behavior of the personnel on the target object can be that the personnel operates the target object for a long time, or that the personnel operates the target object for a short time, and the personnel operating the target object for a long time is regarded as a behavior abnormal personnel.

[0052] In some optional implementations of the present disclosure, the obtaining of the target skeletal key point position of the personnel in the to-be-recognized image frame comprises: extracting a human skeletal key point position in the to-be-recognized image frame; and identifying a target skeletal key point position in the human skeletal key point position.

[0053] In the optional implementation, the human skeletal key point in the to-be-recognized image frame is all the skeletal key points presented by the personnel in the to-be-recognized image frame, and in order to effectively recognize the behavior of the personnel operating the target object, the target skeletal key point in all the skeletal key points is selected.

[0054] In the optional implementation, different numbers of human skeletal key point positions are obtained by using different skeletal key point detection algorithms, for example, 18 human skeletal key points can be obtained by using some skeletal key point detection algorithms (such as OpenPose), and 17 human skeletal key points can be obtained by using other skeletal key point detection algorithms, such as Figure 2The image shows the locations of 17 skeletal keypoints of the human body obtained through the PoseNet model. PoseNet is a deep learning model that estimates human posture by detecting body parts such as elbows, hips, wrists, knees, and ankles, and forms the skeletal structure of the posture by connecting these keypoints. Trained on the ImageNet dataset, PoseNet is primarily used for image classification and object estimation. It is a lightweight model that uses depthwise separable convolutions to deepen the network, reduce parameters, computational cost, and improve accuracy. PoseNet provides a total of 17 usable keypoints, from the eyes to the ears, knees, and ankles.

[0055] In this optional implementation, the target skeleton key points are the key points most relevant to the target object. When personnel operate the target object, the target skeleton key points play a major supporting role.

[0056] This optional implementation provides a reliable basis for selecting target skeletal key point locations by first extracting all human skeletal key point locations and then selecting target skeletal key point locations. This improves the accuracy of the selected target skeletal key point locations.

[0057] In some optional implementations of this disclosure, the aforementioned target skeletal keypoint locations include: wrist keypoint locations and nose keypoint locations; the target object's location information includes: the target object's location and a target object identifier; based on the target skeletal keypoint locations and the target object's location information, determining the person's behavior towards the target object includes: for each target object identifier, detecting whether the distance from the wrist keypoint location to the target object's location is less than a first distance threshold; in response to the distance from the wrist keypoint location to the target object's location being less than the first distance threshold, detecting whether the distance from the target object's location to the nose keypoint location is less than a second distance threshold, wherein the first distance threshold is less than the second distance threshold; in response to the distance from the target object's location to the nose keypoint location being less than the second distance threshold, determining the person's abnormal behavior towards the target object.

[0058] In this optional implementation, the wrist key point location can include: the left wrist key point location or the right wrist key point location, such as... Figure 2 The left wrist keypoint L, right wrist keypoint R, and nose keypoint N were obtained using the PoseNet model.

[0059] Based on the target object's coordinates (x) c y c ) and the location of the left wrist key point L(x) extracted from the key nodes of the human skeleton. l y l Key points on the right wrist (x)r , y r ) and the nose key point position (x n , y n ), three values are calculated using the Euclidean distance calculation method, which are the distance d1 from the target object coordinate to the left wrist key point position, the distance d2 from the target object coordinate to the right wrist key point position, and the distance d3 from the target object coordinate to the nose key point position. The behavior rules of the person operating the target object are customized by threshold setting, and it is judged whether there is an abnormal behavior.

[0060] In the optional implementation, the first distance threshold includes: a first sub-distance threshold s1 of the target object coordinate to the left wrist key point position, and a second sub-distance threshold s2 of the target object coordinate to the right wrist key point position in the state of the person operating the target object. A second distance threshold s3 of the target object coordinate to the nose key point position. When the distance d1 from the target object coordinate to the left wrist key point position in the to-be-identified video data is less than the first sub-distance threshold s1, and the distance d3 from the target object coordinate to the nose key point position is less than the threshold s3;

[0061] or the distance d2 from the target object coordinate to the left wrist key point position is less than the first sub-distance threshold s2, and the distance d3 from the target object coordinate to the nose key point position is less than the second distance threshold s3, it is determined that the on-duty person has an abnormal behavior of the person operating the target object, which indicates that the person is playing a mobile phone or smoking.

[0062] In the optional implementation, the abnormal behavior of the person operating the target object refers to an operation state of the person operating the target object that is not allowed in the current time period or the current scene. The normal behavior of the person operating the target object refers to an operation state of the person operating the target object that is allowed in the current time period or the current scene.

[0063] The method for determining the behavior of the person operating the target object provided in the optional implementation determines the abnormal behavior of the person operating the target object through the distance information between the wrist key point position, the nose key point position and the position of the target object, and provides an optional way for determining the abnormal behavior.

[0064] In some optional implementations of the present disclosure, the target skeleton key point position includes: a wrist key point position and a neck key point position; the position information of the target object includes: a position of the target object and an identification of the target object; and the behavior of the person operating the target object is determined based on the target skeleton key point position and the position information of the target object, including:

[0065] The distance from the wrist key point position to the position of the target object is detected, and whether the distance is less than a first distance threshold is determined; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, the distance from the position of the target object to the neck key point position is detected, and whether the distance is less than a third distance threshold is determined, the first distance threshold being less than the third distance threshold; in response to the distance from the position of the target object to the neck key point position being less than the third distance threshold, the abnormal behavior of the person to the target object is determined.

[0066] In this optional implementation, a mature skeleton key point detection algorithm can be used to detect the neck key point position, and further, a Euclidean distance formula can be used to calculate the distance from the position of the target object to the neck key point position.

[0067] The method for determining the behavior of a person to a target object provided by this optional implementation determines the abnormal behavior of the person to the target object through the distance information between the wrist key point position, the neck key point position and the position of the target object, and provides another optional way for determining abnormal behavior.

[0068] Referring back to Figure 3 , a flow 300 of another embodiment of the behavior recognition method according to the present disclosure is shown, which includes the following steps:

[0069] In step 301, based on real-time collected or shot video data to be recognized, an image frame to be recognized is obtained.

[0070] In step 302, whether there is a target object in the image frame to be recognized is detected in real time; if there is a target object, step 303 is performed.

[0071] In step 303, the target object is tracked to obtain the position of the target object and a target object identifier.

[0072] In step 304, the wrist key point position and the nose key point position of a person in the image frame to be recognized are obtained.

[0073] In step 305, for each target object under a target object identifier, whether the distance from the wrist key point position to the position of the target object under the target object identifier is less than a first distance threshold is detected; if it is less than the first distance threshold, step 306 is performed; otherwise, step 309 is performed.

[0074] In step 306, whether the distance from the position of the target object under the target object identifier to the nose key point position is less than a second distance threshold is detected; if it is less than the second distance threshold, step 307 is performed; otherwise, step 309 is performed.

[0075] In step 307, the abnormal behavior of the person to the target object is determined.

[0076] It should be understood that the operations and features in steps 301-307 above correspond to the operations and features in steps 101-105, respectively, and the operations and features described in the optional implementation above, and thus the descriptions of the operations and features in steps 101-105 and the optional implementation above also apply to steps 301-307, which will not be repeated here.

[0077] In step 308, the abnormal behavior of the personnel is labeled, and warning information that the personnel has abnormal behavior is sent.

[0078] In this embodiment, labeling the abnormal behavior of the personnel means recording the current image frame to be identified as an image frame of personnel with abnormal operation, and the abnormal operation personnel is the personnel who is performing abnormal behavior operation, which can be determined by the identification of the personnel, for example, when it is determined that the personnel has the behavior of playing a mobile phone, the specific human body playing the mobile phone is identified.

[0079] In this embodiment, the warning information can be related to the personnel, for example, the warning information is “xx personnel is performing abnormal behavior operation” or “the personnel at xx position is performing abnormal behavior operation”.

[0080] Optionally, the behavior recognition method provided in this embodiment can further include: intercepting and saving a segment of the to-be-identified video data that has the abnormal behavior of the personnel on the target object, to facilitate subsequent manual review and supervision of the abnormal behavior.

[0081] In step 309, the normal behavior of the personnel on the target object is determined.

[0082] The behavior recognition method provided in this embodiment detects whether the target object under different identifications is in the hand of the personnel based on the distance relationship between the wrist key point position of the personnel, the nose key point position, and the position of the target object, determines the abnormal behavior of the personnel on the target object, and sends the warning information that the personnel has abnormal behavior, thereby providing a reliable monitoring basis for monitoring the abnormal behavior of the personnel playing a mobile terminal or smoking during working hours.

[0083] In some optional implementations of the present disclosure, the real-time detection of whether the to-be-identified image frame has the target object includes: real-time detection of the target object in the to-be-identified image frame by using a target detection model of the target object; and the target detection model is a model obtained after the basic model is tuned.

[0084] In this embodiment, the base model can be a model based on YOLO (You Only Look Once, which means that only one look can identify the category and position of the object in the picture). The YOLO model can find certain specific objects in a picture, not only identify the category of the object, but also mark the position of the object. Specifically, the YOLO model can adopt a YOLOv5 model.

[0085] The method for detecting a target object provided by the optional implementation improves the reliability of target object determination by using a target detection model to perform real-time detection on the to-be-identified image. Furthermore, the target detection model is obtained by adjusting the parameters of the base model, which simplifies the step of obtaining the target detection model and saves the data calculation amount.

[0086] In some optional implementations of the present disclosure, the target detection model is obtained by training in the following steps: collecting sample video data in which different persons operate the target object; obtaining sample image data by intercepting the sample video data at a preset frame rate interval; obtaining image samples by sample labeling the target object in the sample image data using an image labeling tool; and training the base model based on the image samples to obtain the target detection model.

[0087] In the optional implementation, due to the limitation of the working area and other conditions, it can be impossible to obtain the sample data from the real environment. Therefore, simulated video data can be used as sample video data. Specifically, because it is impossible to obtain video data of the real environment for operating the target object for technical research, based on field research, real-time and historical video data of the area corresponding to the to-be-identified video data, and based on sufficient analysis of the behavior of the person operating the target object, the environment, posture, and other conditions of the person operating the target object are summarized. Therefore, the simulated video data is video data that sufficiently simulates the real person operating the target object.

[0088] Optionally, sample video data with the target object can also be obtained from an open model sample library. Based on different requirements for detecting the category of the target object, the content of the sample video data is different. For example, when the target object is a mobile terminal, the content of the sample video data is a video of a worker holding a mobile terminal. When the target object is a cigarette, the content of the sample video data is a video of a worker holding a cigarette.

[0089] When the target object is a mobile phone and the on-duty room personnel are monitored, the sample video data obtaining process is as follows: sample video data similar to the scene of using a mobile phone is collected. The sample video data may be, for example, 19 pieces of sample video data of playing a mobile phone, in which the on-duty personnel playing a mobile phone covers sitting, standing, squatting, crawling and other types, and also simulates the situation of multiple people playing a mobile phone or some people playing a mobile phone and others not playing a mobile phone, and in order to verify the effect of the model, video data of making a phone call which is difficult to distinguish from playing a mobile phone is also collected. When shooting, different monitoring angles and different time periods of on-duty situations are considered, and in this case, the shape of the mobile phone also changes to different degrees.

[0090] In the process of identifying the on-duty personnel playing a mobile phone, the mobile phone entity should be identified first. Since the mobile phone entity is a relatively small object in the video picture, and is affected by different angles, the shape of the mobile phone also changes constantly, so the mobile phone identification model needs to be labeled and the identification model needs to be constructed, trained and tested. Specifically, a target detection model based on yolov5 can be used to implement the construction of the target detection model.

[0091] The training process of the target detection model is as follows: 1) the collected sample video data is intercepted at different image data intervals at a frame rate of 25fps. 2) the image labeling tool (such as LabelImg tool) is used to complete the YOLO format labeling of the mobile phone target of the on-duty personnel, such as labeling the label as “cellphone”. The data set is divided into a training set and a verification set in a ratio of 8:2. 3) based on the pre-trained model weight, the mobile phone target detection re-identification training based on the data set is performed, and based on the above labeled data, the model is fine-tuned based on the YOLOv5 target detection pre-trained model to obtain the on-duty mobile phone target detection model.

[0092] The target detection model provided by the optional implementation manner can effectively supervise the on-duty police playing a mobile phone, accurately identify and warn the on-duty police playing a mobile phone by using deep learning video behavior detection, human key point detection and rule discrimination, and automatically save and record the video clips of playing a mobile phone, which is helpful for the procuratorial supervision personnel to further examine and judge, so as to reduce the violation of the on-duty personnel and improve the efficiency of prison law enforcement and safety management.

[0093] The target detection model training method provided by the optional implementation manner selects image samples and then trains the basic model, simplifies the target detection model obtaining process, and improves the reliability of the target detection model.

[0094] Further reference Figure 4As an implementation of the method shown in the above figures, the disclosure provides an embodiment of a behavior recognition device, which corresponds to the method embodiment shown in Figure 1 The device can be applied in various electronic devices.

[0095] As shown in Figure 4 The behavior recognition device 400 of the embodiment can include a obtaining unit 401, a detecting unit 402, a tracking unit 403, an acquiring unit 404, and a determining unit 405. The obtaining unit 401 can be configured to obtain an image frame to be recognized based on video data to be recognized collected or captured in real time. The detecting unit 402 can be configured to detect whether there is a target object in the image frame to be recognized in real time. The tracking unit 403 can be configured to track the target object to obtain position information of the target object in response to detecting that there is a target object in the image frame to be recognized. The acquiring unit 404 can be configured to acquire a target skeletal key point position of a person in the image frame to be recognized. The determining unit 405 can be configured to determine a behavior of the person to the target object based on the target skeletal key point position and the position information of the target object.

[0096] In some embodiments of the disclosure, the acquiring unit 403 is further configured to: extract a human skeletal key point position in the image frame to be recognized; and identify a target skeletal key point position in the human skeletal key point position.

[0097] In some embodiments of the disclosure, the target skeletal key point position includes a wrist key point position and a nose key point position, and the position information of the target object includes a position of the target object and an identification of the target object. The determining unit 405 is further configured to: for each target object under the identification of the target object, detect whether a distance from the wrist key point position to the position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detect whether a distance from the position of the target object to the nose key point position is less than a second distance threshold, the first distance threshold being less than the second distance threshold; and in response to the distance from the position of the target object to the nose key point position being less than the second distance threshold, determine an abnormal behavior of the person to the target object.

[0098] In some embodiments of the present disclosure, the target skeletal key point positions include a wrist key point position and a neck key point position, and the position information of the target object includes a position of the target object and a target object identifier. The determination unit 405 is further configured to: for each target object under a target object identifier, detect whether a distance from the wrist key point position to the position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detect whether a distance from the position of the target object to the neck key point position is less than a third distance threshold, the first distance threshold being less than the third distance threshold; and in response to the distance from the position of the target object to the neck key point position being less than the third distance threshold, determine that the person has an abnormal behavior on the target object.

[0099] In some embodiments of the present disclosure, the device 400 further includes an alarm unit configured to label the person as having an abnormal behavior and issue early warning information that the person has an abnormal behavior.

[0100] In some embodiments of the present disclosure, the detection unit 402 is further configured to: perform real-time detection on the target object in the to-be-identified image frame by using a target detection model of the target object; and the target detection model is a model obtained by adjusting a basic model.

[0101] In some embodiments of the present disclosure, the target detection model is obtained by training in the following steps: collecting sample video data in which different persons operate on a target object; obtaining sample image data by intercepting the sample video data at a preset frame rate interval; obtaining image samples by performing sample labeling on the target object in the sample image data by using an image labeling tool; and training a basic model based on the image samples to obtain the target detection model.

[0102] The behavior recognition device provided by the present disclosure first obtains a to-be-identified image frame based on to-be-identified video data collected or captured in real time by the obtaining unit 401. Then, the detection unit 402 detects whether there is a target object in the to-be-identified image frame in real time. In response to detecting that there is a target object in the to-be-identified image frame, the tracking unit 403 tracks the target object to obtain position information of the target object. The obtaining unit 404 obtains target skeletal key point positions of a person in the to-be-identified image frame. Finally, the determination unit 405 determines a behavior of the person on the target object based on the target skeletal key point positions and the position information of the target object. Thus, after detecting the target object, the target object is tracked and the target skeletal key point positions of the person are recognized. Based on the positional relationship between the target skeletal key points and the target object, the behavior of the person on the target object is determined, thereby improving the accuracy of behavior detection when the person operates on the target object.

[0103] Figure 5A schematic block diagram of an electronic device 500 for implementing the behavior recognition method according to the embodiments of the present disclosure is shown. As shown in Figure 5 The device 500 can include a processor 501 and a memory 502 storing a computer program. When the computer program is executed by the processor 501, the device 500 can perform the steps of the method 100 as shown in Figure 1 In one example, the device 500 can be a computer device or a cloud computing node.

[0104] In the embodiments of the present disclosure, the processor 501 can be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. The memory 502 can be any type of memory implemented using data storage technologies, including but not limited to random access memory, read only memory, semiconductor-based memory, flash memory, disk storage, etc.

[0105] In addition, in the embodiments of the present disclosure, the device 500 can also include an input device 503, such as a microphone, a keyboard, a mouse, etc., for inputting a plurality of multimedia files to be mixed. In addition, the device 500 can also include an output device 504, such as a loudspeaker, a display, etc., for outputting the mixed multimedia file.

[0106] The behavior recognition device provided by the embodiments of the present disclosure can be applied to any product with display function, such as electronic paper, mobile phone, tablet computer, television, notebook computer, digital photo frame, wearable device, navigation device, etc.

[0107] In other embodiments of the present disclosure, a computer readable storage medium storing a computer program is also provided, wherein the computer program, when executed by a processor, can implement the steps of the method as shown in Figures 1 to 3

[0108] The behavior recognition method provided by the present disclosure first obtains an image frame to be recognized based on real-time collected or photographed video data to be recognized; secondly, detects whether there is a target object in the image frame to be recognized in real time; thirdly, in response to detecting that there is a target object in the image frame to be recognized, tracks the target object to obtain position information of the target object; fourthly, obtains a target skeleton key point position of a person in the image frame to be recognized; and finally, determines a behavior of the person to the target object based on the target skeleton key point position and the position information of the target object. Thus, after detecting the target object, the target object is tracked and the target skeleton key point position of the person is recognized, and the behavior of the person to the target object is determined based on the positional relationship between the target skeleton key point and the target object, thereby improving the accuracy of behavior detection when the person operates the target object.

[0109] ​The diagrams of the flowcharts and block diagrams in the drawings show the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code which comprises one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession can in fact be executed substantially concurrently or in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowcharts, and combinations thereof, can be implemented by dedicated hardware-based systems which perform the specified functions or acts, or can be implemented by a combination of dedicated hardware and computer instructions.

[0110] As used herein and in the appended claims, the singular form of a word includes the plural, and vice versa, unless the context clearly dictates otherwise. Thus, to illustrate, when referring to a singular, the plural is generally included. Similarly, the words "comprise", "comprises" and "comprising" are to be interpreted inclusively rather than exclusively. Likewise, the terms "include", "including" and "or" should be construed as inclusive rather than exclusive, unless otherwise expressly indicated herein. Where the term "example" is used in the present document, particularly after the term "such as", "example" is merely an illustration and is not to be construed as exclusive or exhaustive, unless otherwise expressly indicated herein.

[0111] Further aspects and scope of adaptations become apparent from the description provided herein. It should be appreciated that individual aspects of the present application can be implemented alone or in combination with one or more other aspects. It should also be appreciated that the description and specific examples herein are intended to be illustrative only and are not intended to limit the scope of the present application.

[0112] The above detailed description of several embodiments of the present disclosure has been presented for the purposes of illustration and description. It is apparent to those skilled in the art that various modifications and variations can be made to the embodiments of the present disclosure without deviating from the spirit and scope of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A behavior recognition method, the method comprising: obtaining a to-be-recognized image frame based on real-time collected or shot to-be-recognized video data; detecting whether a target object is present in the to-be-recognized image frame in real time; in response to detecting that the target object is present in the to-be-recognized image frame, tracking the target object to obtain position information of the target object; obtaining a target skeletal key point position of a person in the to-be-recognized image frame; based on the target skeletal key point position and the position information of the target object, determining a behavior of the person towards the target object; the target skeletal key point position comprises a wrist key point position and a nose key point position; the position information of the target object comprises a position of the target object and a target object identifier; and the determining of the behavior of the person towards the target object based on the target skeletal key point position and the position information of the target object comprises: for each target object under a target object identifier, detecting whether a distance from the wrist key point position to a position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detecting whether a distance from the position of the target object to the nose key point position is less than a second distance threshold, the first distance threshold being less than the second distance threshold; in response to the distance from the position of the target object to the nose key point position being less than the second distance threshold, determining an abnormal behavior of the person towards the target object. the target skeletal key point position comprises a wrist key point position and a neck key point position; the position information of the target object comprises a position of the target object and a target object identifier; and the determining of the behavior of the person towards the target object based on the target skeletal key point position and the position information of the target object comprises: for each target object under a target object identifier, detecting whether a distance from the wrist key point position to a position of the target object is less than a first distance threshold; in response to the distance from the wrist key point position to the position of the target object being less than the first distance threshold, detecting whether a distance from the position of the target object to the neck key point position is less than a third distance threshold, the first distance threshold being less than the third distance threshold; in response to the distance from the position of the target object to the neck key point position being less than the third distance threshold, determining an abnormal behavior of the person towards the target object.

2. The method of claim 1, wherein, the obtaining of the target skeletal key point position of the person in the to-be-recognized image frame comprises: extracting a human skeletal key point position in the to-be-recognized image frame; identifying a target skeletal key point position in the human skeletal key point position.

3. The method of claim 2, further comprising: annotating an abnormal behavior of the person and issuing early warning information that the person has an abnormal behavior.

4. The method of claim 1, wherein, the real-time detection of whether a target object is present in the to-be-recognized image frame comprises: using a target object detection model to perform real-time detection of a target object in the to-be-recognized image frame; the target object detection model is a model obtained after parameter tuning of a basic model.

5. The method of claim 4, wherein, The target detection model is trained by the following steps: Collect sample video data of different persons operating the target object; Intercept the sample video data according to a preset frame rate interval to obtain sample image data; Sample mark the target object in the sample image data by using an image marking tool to obtain image samples; Train a basic model based on the image samples to obtain the target detection model.

6. A behavior recognition apparatus for performing a behavior recognition method according to any one of claims 1-5, the apparatus comprising: An obtaining unit configured to obtain an image frame to be recognized based on video data to be recognized collected or photographed in real time; A detection unit configured to detect whether the image frame to be recognized has a target object in real time; A tracking unit configured to track the target object to obtain position information of the target object in response to detecting that the image frame to be recognized has the target object; An acquisition unit configured to acquire a target skeleton key point position of a person in the image frame to be recognized; A determination unit configured to determine a behavior of the person to the target object based on the target skeleton key point position and the position information of the target object.

7. An electronic device comprising: at least one processor; and at least one memory having computer program stored therein; wherein the computer program, when executed by the at least one processor, causes the apparatus to perform the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, wherein, The computer program, when executed by the processor, implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Target behavior detection method and device, computer equipment, and storage medium

    CN113792595A