Method and device for detecting falling behavior, electronic device, computer storage medium

By identifying key points of the human skeleton and fall element values, the problem of high resource consumption in fall behavior detection in existing technologies has been solved, achieving efficient and accurate fall behavior detection.

CN115797973BActive Publication Date: 2025-11-25CETC BIGDATA RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211546276.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-05
Publication Date
2025-11-25
Estimated Expiration
2042-12-05

AI Technical Summary

Technical Problem

In existing technologies, fall detection methods based on models or classifiers tend to consume time and resources in complex scenarios, making it difficult to efficiently detect human fall behavior.

Method used

By identifying key points of the human skeleton in an image frame, the element values ​​of the falling element are determined, and whether the falling condition is met within a time threshold is judged, thus realizing the detection of falling behavior.

Benefits of technology

It simplifies the process of detecting falling behavior, reduces resource consumption, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797973B_ABST
    Figure CN115797973B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a falling behavior detection method, and the specific implementation scheme is as follows: obtaining an image frame based on video data collected or shot in real time; identifying human skeleton key points of at least one person in the image frame; determining an element value of a human falling element based on the human skeleton key points; in response to the element value of the image frame meeting a falling condition, starting timing of a time threshold; and based on the element value of the human falling element of the total image frame, determining a falling behavior detection result within the time threshold. Through this embodiment, the accuracy of human falling detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of computer, and in particular, to a falling behavior detection method and device, electronic equipment and computer readable storage medium. BACKGROUND

[0002] With the rapid development of today's society, there are more and more emergencies. Human abnormal behavior detection based on video monitoring data has become an important research direction of computer vision. Falling behavior is one of the human abnormal behaviors that often occur. For video monitoring data in key places such as prisons, office areas, hospitals, squares, nursing homes, building communities, elevators, and the like, related technical means of human behavior detection and recognition are used to realize the detection of human falling behavior, and then to assist relevant video supervisors to timely discover, handle and solve abnormal situations, and to protect the life and property safety of regional personnel.

[0003] In the prior art, a model or classifier based method is generally used to detect whether falling occurs. However, in actual engineering practice, complex model design is easy to cause consumption of time and resources. SUMMARY

[0004] Embodiments described herein provide a falling behavior detection method, device, electronic equipment and computer readable storage medium having a computer program stored therein.

[0005] According to a first aspect of the present disclosure, a falling behavior detection method is provided. In the method, based on real-time collected or shot video data, an image frame is obtained; human skeleton key points of at least one person in the image frame are identified; based on the human skeleton key points, an element value of a human falling element is determined; in response to the element value of the image frame satisfying a falling condition, timing of a time threshold is started; and within the time threshold, based on the element value of the human falling element of the total image frame, a falling behavior detection result is determined.

[0006] In some embodiments of the present disclosure, the human falling element includes a falling rectangular frame, and the element value includes a length value and a width value of the falling rectangular frame. The determination of the element value of the human falling element based on the human skeleton key points includes: determining minimum position points and maximum position points of the human skeleton key points of each person in the image frame in a first coordinate axis projection, the first coordinate axis being parallel to the pixel row direction of the image frame; connecting the minimum position points and the maximum position points in the first coordinate axis projection to obtain the length of the long side of the falling rectangular frame of each person and the length value of the long side; determining minimum position points and maximum position points of the human skeleton key points of each person in the first coordinate axis vertical direction projection; and connecting the minimum position points and the maximum position points in the first coordinate axis vertical direction projection to obtain the width of the wide side of the falling rectangular frame of each person and the width value of the wide side.

[0007] In some embodiments of the present disclosure, the falling condition includes that a ratio of the length value and the width value of the falling rectangle is greater than 1.

[0008] In some embodiments of the present disclosure, the human body falling element includes a first line segment and a second line segment, the element value includes a first length value of the first line segment and a second length value of the second line segment, and determining the element value of the human body falling element based on the human body skeleton key points includes: determining a first minimum distance and a first maximum distance of the human body skeleton key points of each person in the image frame in the projection of a first coordinate axis, the first coordinate axis being parallel to the pixel row direction of the image frame; determining a second minimum distance and a second maximum distance of the human body skeleton key points of each person in the image frame in the projection of a direction perpendicular to the first coordinate axis; forming a first coordinate point from the first minimum distance and the second minimum distance, forming a second coordinate point from the first maximum distance and the second maximum distance, forming a third coordinate point from the first maximum distance and the second minimum distance, and forming a fourth coordinate point from the first minimum distance and the second maximum distance; connecting the first coordinate point and the second coordinate point to obtain the first line segment and the first length value of the first line segment; and connecting the third coordinate point and the fourth coordinate point to obtain the second line segment and the second length value of the second line segment.

[0009] In some embodiments of the present disclosure, the determining the falling behavior detection result based on the element value of the human body falling element in the total image frames within the time threshold includes: detecting whether the element value of the human body falling element in each image frame in the total image frames within the time threshold satisfies the falling condition; in response to detecting that the element value of the human body falling element in each image frame satisfies the falling condition, detecting whether the number of consecutive image frames in the total image frames is greater than a set number; and in response to the number of consecutive image frames being greater than the set number, determining that the person has a falling behavior.

[0010] In some embodiments of the present disclosure, the method further includes: intercepting and storing a video segment with a person falling behavior in a video corresponding to the total image frames.

[0011] In some embodiments of the present disclosure, identifying the human body skeleton key points of at least one person in the image frame includes: determining human body skeleton node coordinates of each person in the image frame; in response to the human body skeleton node coordinates including wrist coordinates and elbow coordinates, removing the wrist coordinates and the elbow coordinates to obtain the human body skeleton key points.

[0012] According to a second aspect of the present disclosure, a falling behavior detection apparatus is provided. The apparatus comprises: a obtaining unit configured to obtain an image frame based on real-time captured or shot video data; an identifying unit configured to identify human body skeleton key points of at least one person in the image frame; an element determining unit configured to determine an element value of a human body falling element based on the human body skeleton key points; a timing unit configured to start timing a time threshold in response to the element value of the image frame satisfying a falling condition; and a result determining unit configured to determine a falling behavior detection result based on the element value of the human body falling element of the total image frames within the time threshold.

[0013] In some embodiments of the present disclosure, the human body falling element comprises a falling rectangular frame, and the element value comprises a length value and a width value of the falling rectangular frame. The element determining unit is further configured to: determine minimum position points and maximum position points of the human body skeleton key points of each person in the image frame in a first coordinate axis projection, the first coordinate axis being parallel to a pixel row direction of the image frame; connect the minimum position points and the maximum position points in the first coordinate axis projection to obtain a long side of the falling rectangular frame of each person and a length value of the long side; determine minimum position points and maximum position points of the human body skeleton key points of each person in a vertical direction of the first coordinate axis projection; and connect the minimum position points and the maximum position points in the vertical direction of the first coordinate axis projection to obtain a wide side of the falling rectangular frame of each person and a width value of the wide side.

[0014] In some embodiments of the present disclosure, the falling condition comprises that a ratio of the length value and the width value of the falling rectangular frame is greater than 1.

[0015] In some embodiments of the present disclosure, the human body falling element comprises a first line segment and a second line segment, and the element value comprises a first length value of the first line segment and a second length value of the second line segment. The element determining unit is further configured to: determine a first minimum distance and a first maximum distance of the human body skeleton key points of each person in the image frame in a first coordinate axis projection, the first coordinate axis being parallel to a pixel row direction of the image frame; determine a second minimum distance and a second maximum distance of the human body skeleton key points of each person in a vertical direction of the first coordinate axis projection; form a first coordinate point by the first minimum distance and the second minimum distance, form a second coordinate point by the first maximum distance and the second maximum distance, form a third coordinate point by the first maximum distance and the second minimum distance, and form a fourth coordinate point by the first minimum distance and the second maximum distance; connect the first coordinate point and the second coordinate point to obtain the first line segment and the first length value of the first line segment; and connect the third coordinate point and the fourth coordinate point to obtain the second line segment and the second length value of the second line segment.

[0016] In some embodiments of the present disclosure, the result determination unit is further configured to: detect whether the element values of the human falling elements in each of the total image frames satisfy the falling condition within a time threshold; in response to detecting that the element values of the human falling elements in each of the image frames satisfy the falling condition, detect whether the number of consecutive image frames in the total image frames is greater than a set number; and in response to the number of consecutive image frames being greater than the set number, determine that the person has a falling behavior.

[0017] In some embodiments of the present disclosure, the device further comprises: intercepting and storing a video segment with the falling behavior of the person in the video corresponding to the total image frames.

[0018] In some embodiments of the present disclosure, the identification unit is further configured to: determine the human body skeleton node coordinates of each person in the image frame; in response to the human body skeleton node coordinates including two wrist coordinates and two elbow coordinates, and all the wrist coordinates and all the elbow coordinates being on a straight line, remove all the wrist coordinates and all the elbow coordinates to obtain the human body skeleton key points.

[0019] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing a computer program; wherein when the computer program is executed by the at least one processor, the device performs the steps of the method according to the first aspect of the present disclosure.

[0020] According to a fourth aspect of the present disclosure, a computer-readable storage medium storing a computer program is provided, wherein the computer program, when executed by a processor, implements the steps of the method according to the first aspect of the present disclosure.

[0021] The falling behavior detection method and device provided by the present disclosure first obtains an image frame based on real-time collected or photographed video data; secondly, identifies human body skeleton key points of at least one person in the image frame; thirdly, determines element values of human falling elements based on the human body skeleton key points; and finally, in response to the element values of the image frame satisfying a falling condition, starts timing a time threshold; and within the time threshold, determines a falling behavior detection result based on the element values of the human falling elements in the total image frames. Thus, by judging whether the element values of the falling elements satisfy the falling condition, the detection of the falling behavior is simply and quickly realized, which effectively reduces resource consumption and improves the accuracy of human falling detection, compared with the falling behavior detection using a model. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below, and it should be known that the drawings described below only relate to some embodiments of the present disclosure, but not limit the present disclosure, wherein:

[0023] Figure 1 is a flowchart of an embodiment of a falling behavior detection method according to the present disclosure;

[0024] Figure 2 is a flowchart of another embodiment of a falling behavior detection method according to the present disclosure;

[0025] Figure 3 is a diagram of a correspondence between human body skeleton key points and a falling rectangle frame in an embodiment of the present disclosure;

[0026] Figure 4 is a structural schematic diagram of an embodiment of a falling behavior detection device according to the present disclosure; and

[0027] Figure 5 is a block diagram of an electronic device for implementing a falling behavior detection method in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are part of, rather than all of, the embodiments of the present disclosure. Based on the described embodiments of the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without any inventive effort also fall within the scope of protection of the present disclosure.

[0029] In order to solve the problem of large amount of calculation in detecting human falling behavior by using a model or a classifier in the prior art, the present disclosure provides a simple and efficient method for detecting human falling behavior, as shown in Figure 1 which shows a flowchart 100 of an embodiment of a falling behavior detection method according to the present disclosure, the falling behavior detection method comprising the following steps:

[0030] Step 101: obtaining an image frame based on real-time collected or shot video data.

[0031] In the present embodiment, the video data can be video monitoring data in different scenes, and the different scenes can be elevator, square, community building, nursing home, hospital, office area, etc.

[0032] In the present embodiment, the video data can be read in sequence according to the frame rate interval of the video, and the video data can be processed by OpenCV (cross-platform computer vision and machine learning software library) to generate picture frames. Alternatively, other video software can also be used to convert the video data into image frames.

[0033] Step 102: identifying human body skeleton key points of at least one person in the image frame.

[0034] In this embodiment, there can be at least one person in the image frame, and the position information of the skeletal key points of each person can be obtained by performing skeletal key point detection on the image of each person. Specifically, a mature skeletal key point detection algorithm (such as OpenPose) can be used, and optionally, the image frame can be input into an SPPE (Single-Person Pose Estimator) network to obtain the human skeletal key points.

[0035] In this embodiment, after obtaining all the human skeletal key points of at least one person, part of the human skeletal key points can be removed to facilitate the analysis of the falling condition.

[0036] In step 103, the element value of the human falling element is determined based on the human skeletal key points.

[0037] In this embodiment, the human falling element is the key information to distinguish whether the human falls or not, and the human falling element includes a time element and a posture element, wherein the element value of the time element is different time points, and the element value of the posture element is a predetermined position.

[0038] In step 104, the timing of the time threshold is started in response to the element value of the image frame satisfying the falling condition.

[0039] In this embodiment, the falling conditions corresponding to different human falling elements can be different. For example, when the human falling element includes a time element and a posture element, the falling condition includes that the person maintains a fixed position at a predetermined time.

[0040] In this embodiment, starting the timing of the time threshold means starting the timing from zero until the timing stops when the time reaches the time threshold. It should be noted that within the time threshold range of starting the timing, the execution of the falling behavior detection method runs on the video data read in the frame rate interval of the video, and processes the read video data to obtain the total image frames within the time threshold range. The total image frames within the time threshold range can include multiple image frames, each image frame can include at least one person, and the human skeletal key point detection can be performed on all the people in the total image frames to determine whether each image frame in the total image frames satisfies the falling condition.

[0041] In this embodiment, the time threshold can be adaptively set based on the falling conditions of different people in different scenes. For example, in a nursing home scene, the corresponding time threshold for an old person falling can be 5 minutes, and in a hospital scene, the corresponding time threshold for a young person falling can be 15 minutes.

[0042] Step 105, within the time threshold, based on the element value of the human body falling element in the total image frame, the falling behavior detection result is determined.

[0043] In this embodiment, the total image frame is all image frames obtained after processing the video data within the time range corresponding to the time threshold. When the element value of the falling element corresponding to the human body skeleton key point of at least one person in the total image frame meets the falling condition, it is determined that the person is in a falling state within a certain time, and it can be judged that the person falls.

[0044] In this embodiment, the above determination of the falling behavior detection result based on the element value of the human body falling element in the total image frame within the time threshold includes: detecting whether the element value of the human body falling element in each image frame in the total image frame meets the falling condition; if the falling condition is met, it is determined that the video has a person falling behavior.

[0045] Optionally, the above determination of the falling behavior detection result based on the element value of the human body falling element in the total image frame within the time threshold can include: within the time threshold, detecting whether the number of total image frames is greater than a set number; if the number of total image frames is greater than the set number, detecting whether the element value of the human body falling element in each image frame in the total image frame meets the falling condition; if the falling condition is met, it is determined that the video has a person falling behavior.

[0046] The falling behavior detection method provided in this embodiment first obtains an image frame based on real-time collected or shot video data; secondly, identifies the human body skeleton key point of at least one person in the image frame; thirdly, determines the element value of the human body falling element based on the human body skeleton key point; finally, in response to the element value of the image frame meeting the falling condition, the time threshold is started to be counted; and within the time threshold, the falling behavior detection result is determined based on the element value of the human body falling element in the total image frame. Thus, by judging whether the element value of the falling element meets the falling condition, the detection of the falling behavior is simply and quickly realized, and compared with the falling behavior detection by using a model, the resource consumption is effectively reduced, and the accuracy of the human body falling detection is improved.

[0047] In order to better save the human body falling behavior data, the disclosure provides another falling behavior detection method, which is shown in Figure 2 which shows the flow 200 of another embodiment of the falling behavior detection method according to the disclosure, which includes the following steps:

[0048] Step 201, based on real-time collected or shot video data, an image frame is obtained.

[0049] Step 202, identify the human body skeleton key point of at least one person in the image frame.

[0050] In step 203, the element value of the human body falling element is determined based on the human body key points.

[0051] In step 204, in response to the element value of the image frame satisfying the falling condition, the time threshold is started.

[0052] It should be understood that the operations and features in the above steps 201-204 correspond to the operations and features in steps 101-104 respectively, and thus the description of the operations and features in steps 101-104 is also applicable to steps 201-204, which will not be repeated here.

[0053] In step 205, within the time threshold, it is detected whether the element value of the human body falling element in each image frame in the total image frames satisfies the falling condition; if yes, step 206 is performed.

[0054] In the embodiment, when the element value of the human body falling element in each image frame in the total image frames satisfies the falling condition, it is determined that the human body is in a falling state, and thus it can be determined that the person falls.

[0055] In step 206, in response to detecting that the element value of the human body falling element in each image frame satisfies the falling condition, it is detected whether the number of consecutive image frames in the total image frames is greater than a set number; if yes, step 207 is performed.

[0056] In the embodiment, the set number can be determined based on the frame rate interval of the video and the time threshold, and the set number can be less than or equal to the value of the time threshold divided by the frame rate interval. When the set number is equal to the value of the time threshold divided by the frame rate interval, it is determined that the total image frames have no frame loss when the number of consecutive image frames in the total image frames is greater than the set number.

[0057] In step 207, it is determined that the person has a falling behavior.

[0058] In step 208, the video segment with the falling behavior of the person is intercepted and stored.

[0059] In the embodiment, the video obtained within the time threshold is a continuous video, and after the image frame satisfying the falling condition is determined, the image frame satisfying the falling condition can be directly generated into a video segment by using a video editing software or a video template, wherein the video editing software or the video template are mature technologies, which will not be repeated here.

[0060] The falling behavior detection method provided in the embodiment determines whether the person has a falling behavior based on the total image frames of the video data within the time threshold, and when the person has a falling behavior, the video segment with the falling behavior of the person is intercepted and stored, which can provide a reliable basis for the behavior analysis of the falling person.

[0061] In some optional implementations of this embodiment, the human body falling to the ground element includes: a falling rectangle, and the element values ​​include the length and width of the falling rectangle. The above-mentioned determination of the element values ​​of the human body falling to the ground element based on human skeletal key points includes: determining the minimum and maximum position points of the human skeletal key points of each person in the image frame projected onto the first coordinate axis, wherein the first coordinate axis is parallel to the pixel row direction of the image frame; connecting the minimum and maximum position points projected onto the first coordinate axis to obtain the long side and the length value of the long side of the falling rectangle of each person; determining the minimum and maximum position points of the human skeletal key points of each person in the image frame projected onto the first coordinate axis in the vertical direction; connecting the minimum and maximum position points projected onto the first coordinate axis in the vertical direction to obtain the wide side and the width value of the wide side of the falling rectangle of each person.

[0062] like Figure 3 As shown, there are 18 key points on the human skeleton: nose 0, neck 1, right shoulder 2, right elbow 3, right wrist 4, left shoulder 5, left elbow 6, left wrist 7, right hip 8, right knee 9, right ankle 10, left hip 11, left knee 12, left ankle 13, right eye 14, left eye 15, right ear 16, and left ear 17. The inverted rectangle k is determined based on the minimum and maximum positions of the key points on the first coordinate axis x, and the minimum and maximum positions of the key points on the human skeleton y in the direction perpendicular to the first coordinate axis.

[0063] The method for determining the element values ​​of a human body falling to the ground provided in this embodiment uses the length and width of the falling rectangle as the falling element and the length and width values ​​of the falling rectangle as the element values. This provides an optional implementation for determining the element values ​​of a human body falling to the ground and improves the reliability of human body falling detection.

[0064] Regarding the method for determining the element values ​​of a human body falling to the ground provided in the above embodiments, in one embodiment of this disclosure, the falling condition includes: the ratio of the length value to the width value of the falling rectangle is greater than 1.

[0065] The falling condition provided in this embodiment uses the ratio of the length to the width of the falling rectangle as a means to determine whether the falling condition is met, providing an optional method for implementing the falling condition of a human body.

[0066] Optionally, weight values ​​are assigned to the length and width of the falling rectangle under different scenario requirements. For example, in a sports scenario (for athletes on a sports field), the weight value of the length is less than the weight value of the width. The falling condition may include: the ratio of the length value multiplied by the length weight value and the width value multiplied by the width weight value of the falling rectangle is greater than 1.

[0067] Optionally, the falling condition includes that the length value of the falling rectangle frame is greater than the width value of the falling rectangle frame.

[0068] In some optional implementations of the embodiment, the human body falling element includes a first line segment and a second line segment, and the element value includes a first length value of the first line segment and a second length value of the second line segment. The method for determining the element value of the human body falling element based on the human body skeleton key points includes:

[0069] The first minimum distance and the first maximum distance of the human body skeleton key points of each person in the image frame in the projection of the first coordinate axis are determined, and the first coordinate axis is parallel to the pixel row direction of the image frame. The second minimum distance and the second maximum distance of the human body skeleton key points of each person in the image frame in the projection of the vertical direction of the first coordinate axis are determined. The first coordinate point is composed of the first minimum distance and the second minimum distance, the second coordinate point is composed of the first maximum distance and the second maximum distance, the third coordinate point is composed of the first maximum distance and the second minimum distance, and the fourth coordinate point is composed of the first minimum distance and the second maximum distance. The first line segment and the first length value of the first line segment are obtained by connecting the first coordinate point and the second coordinate point. The second line segment and the second length value of the second line segment are obtained by connecting the third coordinate point and the fourth coordinate point.

[0070] In the optional implementation, the first line segment is obtained based on the first minimum distance of the projection of the human body skeleton key points on the first coordinate axis, the second minimum distance of the projection of the vertical direction of the first coordinate axis, the first coordinate point, and the first maximum distance of the projection of the human body skeleton key points on the first coordinate axis, the second maximum distance of the projection of the vertical direction of the first coordinate axis, and the second coordinate point. Specifically, the first line segment is obtained by connecting the first coordinate point and the second coordinate point.

[0071] In the optional implementation, the second line segment is obtained based on the first maximum distance of the projection of the human body skeleton key points on the first coordinate axis, the second minimum distance of the projection of the vertical direction of the first coordinate axis, the third coordinate point, and the first minimum distance of the projection of the human body skeleton key points on the first coordinate axis, the second maximum distance of the projection of the vertical direction of the first coordinate axis, and the fourth coordinate point. Specifically, the second line segment is obtained by connecting the third coordinate point and the fourth coordinate point.

[0072] The method for determining the element value of the human body falling element provided in the embodiment provides another optional implementation for determining the element value of the human body falling element by taking the first line segment and the second line segment as the falling element and taking the first length value of the first line segment and the second length value of the second line segment as the element value, thereby improving the reliability of human body falling detection.

[0073] For the method for determining the element value of the human body falling element provided in the above embodiment, the falling condition can include that the ratio of the first length value and the second length value is greater than 1.

[0074] The ratio of the first length value and the second length value is taken as a means for judging whether the falling condition is met, and another optional way is provided for implementing the falling condition of the human body.

[0075] Optionally, different scene requirements are provided for the first line segment and the second line segment, respectively, such as in a sports scene (for personnel, it is a player on the sports field), the weight value of the first line segment is less than the weight value of the second line segment, and the falling condition can include: the ratio of the first length value multiplied by the weight value of the first line segment and the second length value multiplied by the weight value of the second line segment is greater than 1.

[0076] Optionally, the falling condition can also include: the first length value is greater than the second length value.

[0077] In this embodiment, as an example, the human body skeleton key point can be all 18 key points as shown in Figure 3 .

[0078] In order to exclude the interference of the human body after stretching the arms to the full length (the length of the arms of some people is higher than the height of the human body) on the falling condition, in some optional implementations of this embodiment, the identifying the human body skeleton key point of at least one person in the image frame includes: determining the human body skeleton node coordinates of each person in the image frame; in response to the human body skeleton node coordinates including two wrist coordinates and two elbow coordinates, and all wrist coordinates and all elbow coordinates being on a straight line, removing the wrist coordinates and the elbow coordinates to obtain the human body skeleton key point. As shown in Figure 3 , the two wrist coordinates are the right wrist 4 coordinate and the left wrist 7 coordinate, and the two elbow coordinates are the right elbow 3 coordinate and the left elbow 6 coordinate.

[0079] The method for identifying the human body skeleton key point provided by this optional implementation removes the wrist coordinates and the elbow coordinates when the human body skeleton node coordinates in the image frame have two wrist coordinates and two elbow coordinates, and all wrist coordinates and all elbow coordinates are on a straight line, thereby excluding the interference of the length of the arms being greater than the height of the human body when standing on the falling condition.

[0080] Optionally, the identifying the human body skeleton key point of at least one person in the image frame includes: determining the human body skeleton node coordinates of each person in the image frame; in response to the human body skeleton node coordinates including two wrist coordinates and two elbow coordinates, and all wrist coordinates and all elbow coordinates being on a straight line, removing any one of the two wrist coordinates and any one of the two elbow coordinates to obtain the human body skeleton key point.

[0081] Optionally, the identifying the human skeleton key points of the at least one person in the image frame comprises: detecting a scene where the person in the image frame is located through environment information of the person in the image frame; in response to the scene where the person is located not belonging to a motion scene, determining human skeleton node coordinates of each person in the image frame, and taking the human skeleton node coordinates as the human skeleton key points.

[0082] Optionally, the identifying the human skeleton key points of the at least one person in the image frame further comprises: in response to the scene where the person is located belonging to a motion scene, when the human skeleton node coordinates include two wrist coordinates and two elbow coordinates, removing the wrist coordinates and the elbow coordinates to obtain the human skeleton key points.

[0083] With reference to Figure 4 , as an implementation of the method shown in Figure 1 , the present application provides a falling behavior detection device, which corresponds to the method embodiment shown in Figure 1 , and can be applied to various electronic devices.

[0084] As shown in Figure 4 , the falling behavior detection device 400 of the embodiment can include a obtaining unit 401, an identifying unit 402, an element determining unit 403, a timing unit 404, and a result determining unit 405. The obtaining unit 401 can be configured to obtain an image frame based on video data collected or captured in real time. The identifying unit 402 can be configured to identify human skeleton key points of at least one person in the image frame. The element determining unit 403 can be configured to determine an element value of a human falling element based on the human skeleton key points. The timing unit 404 can be configured to start timing a time threshold in response to the element value of the image frame meeting a falling condition. The result determining unit 405 can be configured to determine a falling behavior detection result based on the element value of the human falling element of the total image frame within the time threshold.

[0085] In some optional implementations of the present disclosure, the human falling element includes a falling rectangular frame, and the element value includes a length value and a width value of the falling rectangular frame. The element determining unit 403 is further configured to: determine minimum position points and maximum position points of the human skeleton key points of each person in the image frame in a first coordinate axis projection, the first coordinate axis being parallel to a pixel row direction of the image frame; connect the minimum position points and the maximum position points in the first coordinate axis projection to obtain a long side of the falling rectangular frame of each person and a length value of the long side; determine minimum position points and maximum position points of the human skeleton key points of each person in a vertical direction of the first coordinate axis projection; and connect the minimum position points and the maximum position points in the vertical direction of the first coordinate axis projection to obtain a wide side of the falling rectangular frame of each person and a width value of the wide side.

[0086] In some embodiments of the present disclosure, the falling condition includes that a ratio of the length value and the width value of the falling rectangle is greater than 1.

[0087] In some embodiments of the present disclosure, the human body falling element includes a first line segment and a second line segment, the element value includes a first length value of the first line segment and a second length value of the second line segment, and the element determination unit 403 is further configured to: determine a first minimum distance and a first maximum distance of the human body skeleton key points of each person in the image frame in the first coordinate axis projection, the first coordinate axis being parallel to the pixel row direction of the image frame; determine a second minimum distance and a second maximum distance of the human body skeleton key points of each person in the image frame in the vertical direction of the first coordinate axis; form a first coordinate point by the first minimum distance and the second minimum distance, form a second coordinate point by the first maximum distance and the second maximum distance, form a third coordinate point by the first maximum distance and the second minimum distance, and form a fourth coordinate point by the first minimum distance and the second maximum distance; connect the first coordinate point and the second coordinate point to obtain the first line segment and the first length value of the first line segment; and connect the third coordinate point and the fourth coordinate point to obtain the second line segment and the second length value of the second line segment.

[0088] In some embodiments of the present disclosure, the result determination unit 405 is further configured to: within a time threshold, detect whether the element values of the human body falling elements of each image frame in the total image frames all satisfy the falling condition; in response to detecting that the element values of the human body falling elements of each image frame all satisfy the falling condition, detect whether the number of consecutive image frames in the total image frames is greater than a set number; and in response to the number of consecutive image frames being greater than the set number, determine that the personnel falling behavior exists.

[0089] In some embodiments of the present disclosure, the device 400 further includes: intercepting and storing a video segment with the personnel falling behavior in a video corresponding to the total image frames.

[0090] In some embodiments of the present disclosure, the recognition unit 402 is further configured to: determine the human body skeleton node coordinates of each person in the image frame; in response to the human body skeleton node coordinates including two wrist coordinates and two elbow coordinates, and all the wrist coordinates and all the elbow coordinates being on a straight line, remove all the wrist coordinates and all the elbow coordinates to obtain the human body skeleton key points.

[0091] The falling behavior detection device provided by the embodiment first obtains an image frame based on video data collected or photographed in real time by the obtaining unit 401; secondly, the recognizing unit 402 recognizes at least one human body skeleton key point in the image frame; thirdly, the element determining unit 403 determines an element value of a human body falling element based on the human body skeleton key point; fourthly, the timing unit 404 starts timing a time threshold in response to the element value of the image frame meeting a falling condition; and finally, the result determining unit 405 determines a falling behavior detection result based on the element value of the human body falling element of the total image frame within the time threshold. Thus, by judging whether the element value of the falling element meets the falling condition, the falling behavior is simply and quickly detected, resource consumption is effectively reduced compared with the falling behavior detection by using a model, and the accuracy of human body falling detection is improved.

[0092] Figure 5 A schematic block diagram of an apparatus 500 for mixing multimedia files according to an embodiment of the present disclosure is shown. As shown, the apparatus 500 can include a processor 501 and a memory 502 storing a computer program. When the computer program is executed by the processor 501, the apparatus 500 can perform the steps of the method shown as Figure 5 or Figure 1 or Figure 2 In one example, the apparatus 500 can be a computer device or a cloud computing node.

[0093] In embodiments of the present disclosure, the processor 501 can be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. The memory 502 can be any type of memory implemented using data storage technologies, including but not limited to random access memory, read only memory, semiconductor-based memory, flash memory, disk storage, etc.

[0094] In addition, in embodiments of the present disclosure, the apparatus 500 can also include an input device 503, such as a microphone, a keyboard, a mouse, etc., for inputting a plurality of multimedia files to be mixed. In addition, the apparatus 500 can also include an output device 504, such as a loudspeaker, a display, etc., for outputting the mixed multimedia file.

[0095] The display device provided by the embodiments of the present disclosure can be applied to any product with display function, for example, electronic paper, mobile phone, tablet computer, television, notebook computer, digital photo frame, wearable device or navigator, etc.

[0096] In other embodiments of the present disclosure, a computer readable storage medium storing a computer program is also provided, wherein the computer program, when executed by a processor, can implement the steps of the method shown as Figures 1 to 2 .

[0097] The method for detecting falling behavior provided by the present disclosure first obtains an image frame based on real-time collected or shot video data; second, identifies human body skeleton key points in the image frame; third, determines a falling element value of a human body falling element based on the human body skeleton key points; and finally, in response to the falling element value of the image frame satisfying a falling condition, starts timing of a time threshold; and within the time threshold, determines a falling behavior detection result based on the falling element value of the human body falling element of the total image frame. Thus, by judging whether the falling element value satisfies the falling condition, the detection of the falling behavior is simply and quickly realized, and compared with the falling behavior detection using a model, the resource consumption is effectively reduced, and the accuracy of the human body falling detection is improved.

[0098] The flowcharts and block diagrams in the drawings show the architectural, functional and operational aspects of possible implementations of apparatuses and methods in accordance with various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment or a portion of code which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figures. For example, two blocks shown in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It is also noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0099] Unless the context clearly indicates otherwise, as used herein and in the appended claims, the singular form of a word includes the plural and vice versa. Thus, the use of the singular will include its interpretation as the plural and vice versa. Similarly, the words "comprise", "comprises" and "comprising" are to be interpreted inclusively rather than exclusively. Likewise, the terms "include", "including" and "or" should be construed as inclusive rather than exclusive. Where the term "example" is used in the following description, particularly in the context of a series of terms, "example" is merely an example of one or more of the terms and should not be construed as a complete listing of all of the terms, nor should it be construed as excluding additional terms not specifically listed.

[0100] Further aspects and scope of adaptation become apparent from the description provided herein. It should be appreciated that individual aspects of the present disclosure can be implemented alone or in combination with one or more other aspects. It should also be appreciated that the description and specific examples herein are intended to be for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0101] The above detailed description of several embodiments of the disclosure sets forth both the permitted and specific modifications and variations. It is to be understood that those skilled in the art can make various modifications and variations without departing from the spirit and scope of the disclosure. The scope of protection of the disclosure is defined by the appended claims.

Claims

1. A method for detecting falling behavior, the method comprising: Image frames are obtained based on real-time acquired or captured video data; Identify key points of the human skeleton of at least one person in the image frame; Based on the aforementioned key points of the human skeleton, determine the element values ​​of the human falling element; In response to the element values ​​of the image frame satisfying the inversion condition, timing of the time threshold begins; Within the time threshold, the detection result of the falling behavior is determined based on the element values ​​of the human body falling elements in the total image frames; The human body falling element includes: a first line segment and a second line segment; the element value includes a first length value of the first line segment and a second length value of the second line segment; the element value for determining the human body falling element based on the key points of the human skeleton includes: Determine the first minimum distance and the first maximum distance of the human skeleton key points of each person in the image frame projected onto the first coordinate axis, wherein the first coordinate axis is parallel to the pixel row direction of the image frame; Determine the second minimum distance and the second maximum distance of the projection of the human skeleton key points of each person in the image frame onto the vertical direction of the first coordinate axis; A first coordinate point is formed by the first minimum distance and the second minimum distance; a second coordinate point is formed by the first maximum distance and the second maximum distance; a third coordinate point is formed by the first maximum distance and the second minimum distance; and a fourth coordinate point is formed by the first minimum distance and the second maximum distance. Connect the first coordinate point and the second coordinate point to obtain the first line segment and the first length value of the first line segment; Connect the third coordinate point and the fourth coordinate point to obtain the second line segment and the second length value of the second line segment.

2. The method according to claim 1, wherein, The human body falling element includes: a falling rectangle, and the element values ​​include the length and width of the falling rectangle. The element values ​​for determining the human body falling element based on the human skeletal key points include: Determine the minimum and maximum position points of the human skeleton key points of each person in the image frame projected onto the first coordinate axis, where the first coordinate axis is parallel to the pixel row direction of the image frame; Connect the minimum and maximum position points projected on the first coordinate axis to obtain the long side of the rectangle of each person's fall and the length value of the long side; Determine the minimum and maximum position points of the human skeleton key points of each person in the image frame projected in the direction perpendicular to the first coordinate axis; By connecting the minimum and maximum position points projected vertically along the first coordinate axis, the width of the rectangle of each person's fall and the width value of the width are obtained.

3. The method according to claim 2, wherein, The conditions for falling to the ground include: the ratio of the length value to the width value of the falling rectangle is greater than 1.

4. The method according to claim 1, wherein, Within the time threshold, determining the fall detection result based on the element values ​​of the human body falling elements in the total image frames includes: Within the time threshold, it is detected whether the element values ​​of the human body falling to the ground in each image frame of the total image frames all meet the falling condition. In response to the detection that the element values ​​of the human body falling to the ground in each image frame all meet the falling condition, it is detected whether the number of consecutive image frames of that type in the total number of image frames is greater than a set number. If the number of consecutive image frames exceeds a set number, it is determined that there is a person falling to the ground.

5. The method according to claim 4, further comprising: Extract and store video segments from the corresponding total image frames that show people falling to the ground.

6. The method according to any one of claims 1-5, wherein, The identification of key points of the human skeleton of at least one person in the image frame includes: Determine the coordinates of the human skeleton nodes of each person in the image frame; In response to the human skeleton node coordinates including two wrist coordinates and two elbow coordinates, and when all wrist coordinates and all elbow coordinates are in a straight line, all wrist coordinates and all elbow coordinates are removed to obtain the human skeleton key points.

7. A fall detection device, the device comprising: The unit is configured to obtain image frames based on real-time acquired or captured video data. The recognition unit is configured to recognize key points of the human skeleton of at least one person in the image frame; The element determination unit is configured to determine the element value of the human body falling to the ground element based on the key points of the human skeleton. The timing unit is configured to start timing a time threshold in response to the element value of the image frame satisfying the fall condition; The result determination unit is configured to determine the fall behavior detection result based on the element values ​​of the human fall element in the total image frames within the time threshold. The human body falling element includes: a first line segment and a second line segment; the element value includes a first length value of the first line segment and a second length value of the second line segment; the element value for determining the human body falling element based on the key points of the human skeleton includes: Determine the first minimum distance and the first maximum distance of the human skeleton key points of each person in the image frame projected onto the first coordinate axis, wherein the first coordinate axis is parallel to the pixel row direction of the image frame; Determine the second minimum distance and the second maximum distance of the projection of the human skeleton key points of each person in the image frame onto the vertical direction of the first coordinate axis; A first coordinate point is formed by the first minimum distance and the second minimum distance; a second coordinate point is formed by the first maximum distance and the second maximum distance; a third coordinate point is formed by the first maximum distance and the second minimum distance; and a fourth coordinate point is formed by the first minimum distance and the second maximum distance. Connect the first coordinate point and the second coordinate point to obtain the first line segment and the first length value of the first line segment; Connect the third coordinate point and the fourth coordinate point to obtain the second line segment and the second length value of the second line segment.

8. An electronic device, comprising: At least one processor; as well as At least one memory storing a computer program; When the computer program is executed by the at least one processor, the processor performs the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, wherein, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Human body tumble detection method and device and terminal equipment

    CN113392681A

  • Construction site safety behavior monitoring method and device based on deep learning, and electronic equipment

    CN114973335A