Fall detection method, apparatus, device, and medium

By fusing the positional information of multiple frames of skeletal keypoint data into fall detection and updating it with distance information, the impact of camera distance on detection accuracy is resolved, thereby improving the robustness and accuracy of fall detection.

CN118279812BActive Publication Date: 2026-08-25SHENZHEN MICROBT ELECTRONICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211733608.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-08-25
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

Existing fall detection methods have inconsistent accuracy depending on the distance between pedestrians and the camera equipment, especially at long distances, which affects the robustness and accuracy of the detection.

Method used

By acquiring skeletal keypoint data from multiple video frames, fusing positional information to determine the range of motion, and updating the confidence level with distance information, fall detection is performed using the distance between skeletal keypoints within the range of motion and the starting point, thus reducing the impact of distance on detection.

Benefits of technology

It improves the robustness and accuracy of fall detection, especially when pedestrians are far from the camera, avoiding a decrease in detection accuracy and enhancing the recognition of differences in different movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118279812B_ABST
    Figure CN118279812B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a fall detection method, device, equipment and medium, wherein the method specifically comprises: acquiring N video frames to be detected; determining key point data of a skeleton key point in the N video frames respectively; fusing first position information of the skeleton key points in the N video frames to obtain action range information of a preset target in the N video frames; updating the first position information of the skeleton key points in the N video frames according to the action range information to obtain second position information after updating; updating confidence information to first distance information of the skeleton key points according to the second position information of the skeleton key points in the N video frames; and detecting whether the preset target has a falling action according to the second position information and the first distance information of the skeleton key points in the N video frames. The embodiment of the application can improve the robustness of fall detection and improve the accuracy of fall detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a fall detection method, apparatus, device, and medium. Background Technology

[0002] In video monitoring scenarios, timely detection of pedestrians and other pre-set targets falling and alerting relevant personnel can effectively mitigate the consequences of accidental falls, improve service quality in relevant locations (such as shopping malls and subways), and better protect pedestrian safety.

[0003] Current fall detection methods typically use camera equipment to capture target videos containing pedestrians, identify skeletal key points in the video frames, and detect whether the pedestrian has fallen based on the skeletal key points.

[0004] In practical applications, the distance between pedestrians and the camera equipment can easily affect the accuracy of fall detection. For example, when pedestrians are far from the camera equipment, the accuracy of fall detection is lower; while when pedestrians are close to the camera equipment, the accuracy of fall detection is higher. Summary of the Invention

[0005] This application provides a fall detection method that can improve the robustness and accuracy of fall detection.

[0006] Accordingly, embodiments of this application also provide a fall detection device, an electronic device, and a machine-readable medium to ensure the implementation and application of the above methods.

[0007] To address the aforementioned problems, this application discloses a fall detection method, the method comprising:

[0008] Obtain N video frames to be detected; N is a positive integer;

[0009] Key point data of skeletal key points in N video frames are determined respectively; the key point data includes: first location information and confidence information;

[0010] The first position information of the skeletal key points in N video frames is fused to obtain the motion range information of the preset target in N video frames;

[0011] Based on the motion range information, the first position information of the skeletal key points in the N video frames is updated to obtain the updated second position information;

[0012] Based on the second position information of the skeletal keypoints in N video frames, the confidence information is updated to the first distance information of the skeletal keypoints; the first distance information represents the distance of the skeletal keypoint relative to the starting point of the action range, obtained based on the second position information.

[0013] Based on the second position information and first distance information of the skeletal key points in N video frames, detect whether the preset target has a falling action.

[0014] To address the aforementioned problems, this application discloses a fall detection device, the device comprising:

[0015] The video frame acquisition module is used to acquire N video frames to be detected; N is a positive integer.

[0016] A key point determination module is used to determine key point data of skeletal key points in N video frames respectively; the key point data includes: first location information and confidence information;

[0017] The fusion module is used to fuse the first position information of the skeletal key points in N video frames to obtain the motion range information of the preset target in N video frames;

[0018] The first update module is used to update the first position information of the skeletal key points in the N video frames according to the motion range information, so as to obtain the updated second position information.

[0019] The second update module is used to update the confidence information to the first distance information of the skeletal keypoints based on the second position information of the skeletal keypoints in N video frames; the first distance information represents the distance of the skeletal keypoint relative to the starting point of the action range, obtained based on the second position information.

[0020] The first detection module is used to detect whether a preset target has a falling action based on the second position information and the first distance information of the skeletal key points in N video frames.

[0021] Optionally, the fusion module includes:

[0022] The upper and lower limit determination module is used to determine the lower and upper limits of the horizontal coordinate, as well as the lower and upper limits of the vertical coordinate, based on the first position information of the skeletal key points in N video frames.

[0023] The coordinate range determination module is used to determine the range information of the horizontal coordinate based on the lower limit and the upper limit of the horizontal coordinate, and to determine the range information of the vertical coordinate based on the lower limit and the upper limit of the vertical coordinate.

[0024] Optionally, the first update module includes:

[0025] The horizontal coordinate update module is used to determine the horizontal coordinate of the second position information based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information.

[0026] The horizontal coordinate update module is used to determine the vertical coordinate of the second position information based on the second difference between the vertical coordinate of the first position information and the lower limit of the vertical coordinate, as well as the vertical coordinate range information.

[0027] Optionally, the device further includes:

[0028] The distance determination module is used to determine the first distance information of the skeletal key points based on the squares of the horizontal and vertical coordinates of the second position information and adjustment parameters.

[0029] Optionally, the device further includes:

[0030] The center of gravity determination module is used to determine the center of gravity position of the pedestrian in the i-th video frame based on the second position information of the preset skeletal key points in the i-th video frame; i is a positive integer;

[0031] The third update module is used to update the second position information of the skeletal key points in the i-th video frame according to the center position of the preset target in the i-th video frame, so as to obtain the third position information of the skeletal key points in the i-th video frame.

[0032] The fourth update module is used to update the confidence information to the second distance information of the skeletal keypoints based on the third position information of the skeletal keypoints in N video frames; the second distance information represents the distance of the skeletal keypoint relative to the starting point of the action range, obtained based on the third position information.

[0033] The second detection module is used to detect whether a preset target has a falling action based on the third position information and the second distance information of the skeletal key points in N video frames.

[0034] Optionally, the third update module is specifically used to determine the third position information of the skeletal key points in the i-th video frame based on the third difference between the second position information of the skeletal key points in the i-th video frame and the center of gravity position of the preset target.

[0035] This application also discloses an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform the method described in this application.

[0036] This application also discloses a machine-readable medium storing executable code thereon, which, when executed, causes a processor to perform the method described in this application.

[0037] The embodiments of this application have the following advantages:

[0038] In the technical solution of this application embodiment, the motion range information of the preset target in N video frames can characterize the motion range of the preset target in N video frames. Based on this motion range information, this application embodiment updates the first position information of the skeletal key points in the N video frames. The updating principle of the skeletal key points in this application embodiment is: updating the motion range of the skeletal key points relative to the preset target in N video frames. Since this updating principle is not limited by the distance between the preset target and the camera device, this application embodiment can reduce the impact of the distance between the preset target and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the preset target is far from the camera device, the updating of the motion range of the skeletal key points relative to the preset target in N video frames in this application embodiment can avoid the problem of the updated value of the skeletal key points being a small value, thus reducing the impact of the distance between the preset target and the camera device on the accuracy of fall detection.

[0039] This embodiment updates the confidence score to the first distance information of the skeletal keypoints based on the second position information of the skeletal keypoints in N video frames, and uses the first distance information of the skeletal keypoints relative to the starting point of the movement range in the fall detection process. Since the first distance information can be used to distinguish the motion path lengths of different skeletal keypoints, this embodiment can improve the difference between different movements based on the distinguishable motion path length, thereby improving the accuracy of fall detection. For example, for a fall, the motion path length of the upper limb skeletal keypoints is usually greater than that of the lower limb skeletal keypoints; while for a squatting movement, the motion path lengths of the upper limb skeletal keypoints are roughly equivalent to those of the lower limb skeletal keypoints, etc. Therefore, this embodiment can improve the difference between different movements based on the distinguishable motion path length. Attached Figure Description

[0040] Figure 1 This is a schematic flowchart of the fall detection method according to an embodiment of this application;

[0041] Figure 2 This is a schematic diagram of the skeletal key points of N video frames according to an embodiment of this application;

[0042] Figure 3 This is a schematic flowchart of the fall detection method according to an embodiment of this application;

[0043] Figure 4 This is a schematic diagram of the structure of a fall detection device according to an embodiment of this application;

[0044] Figure 5 This is a schematic diagram of the structure of an apparatus provided in one embodiment of this application. Detailed Implementation

[0045] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] The embodiments of this application can be applied to fall detection scenarios. In a fall detection scenario, a camera device can be used to capture target video containing a preset target, and the presence of a fall action can be detected in the video.

[0047] Fall detection scenarios can include home settings or public places. In home settings, accidental falls pose a significant threat to the physical and psychological well-being of the elderly, and timely detection and assistance for fallen elderly individuals can greatly reduce disability and mortality rates. In public places, promptly detecting falls of pre-defined individuals and alerting relevant personnel can improve service quality in public places (such as shopping malls and subways) and better protect the safety of pre-defined individuals. It is understood that the embodiments of this application do not limit the specific fall detection scenarios.

[0048] Current fall detection methods identify skeletal keypoints in video frames and use these keypoints to detect whether a target is undergoing a fall. In this process, the skeletal keypoints are typically normalized first, specifically by updating their dimensions relative to the width and height of the video frame. A technical problem with this normalization is that when the target is far from the camera, the normalized values ​​for the skeletal keypoints tend to be smaller, which can negatively impact the accuracy of fall detection.

[0049] To address the technical problem that the distance between a preset target and a camera device can easily affect the accuracy of fall detection, this application provides a fall detection method. This method specifically includes: acquiring N video frames to be detected; N being a positive integer; determining key point data of skeletal keypoints in each of the N video frames; the key point data includes: first position information and confidence information; fusing the first position information of the skeletal keypoints in the N video frames to obtain the movement range information of the preset target in the N video frames; updating the first position information of the skeletal keypoints in the N video frames according to the movement range information to obtain updated second position information; updating the confidence information to first distance information of the skeletal keypoints according to the second position information of the skeletal keypoints in the N video frames; the first distance information represents the distance of the skeletal keypoint relative to the starting point of the movement range, obtained based on the second position information; and detecting whether the preset target has a fall action based on the second position information and the first distance information of the skeletal keypoints in the N video frames.

[0050] In this embodiment, the motion range information of the preset target in N video frames can characterize the motion range of the preset target in N video frames. Based on this motion range information, this embodiment updates the first position information of the skeletal key points in the N video frames. The updating principle for the skeletal key points in this embodiment is: updating the motion range of the skeletal key points relative to the preset target in N video frames. Since this updating principle is not limited by the distance between the preset target and the camera device, this embodiment can reduce the impact of the distance between the preset target and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the preset target is far from the camera device, the updating of the motion range of the skeletal key points relative to the preset target in N video frames in this embodiment can avoid the problem of the updated value of the skeletal key points being a small value, thus reducing the impact of the distance between the preset target and the camera device on the accuracy of fall detection.

[0051] In this technical field, a key point detection model is typically used first to determine the key point data of the skeleton key points in N video frames; then, the key point data is input into a fall detection model, which detects whether a preset target has a falling action.

[0052] In practical applications, keypoint detection models in traditional technologies typically output keypoint data that includes three dimensions: X, Y, and score. X represents the horizontal coordinate of the first position, Y represents the vertical coordinate of the first position, and score represents the confidence score of the skeletal keypoint. The confidence score represents the probability that the pixel corresponding to the skeletal keypoint is indeed a skeletal keypoint. The score ∈ (0,1). These three dimensions (X, Y, and score) can then be used in calculations within fall detection models.

[0053] Since the confidence information represented by the score in traditional technologies is not related to human actions, the confidence information usually only serves as a placeholder in subsequent fall detection models. Experiments show that even if the scores of all skeletal keypoints in N video frames are set to values ​​between 0.8 and 1.0, it has virtually no impact on the detection results of the fall detection model.

[0054] In this embodiment of the application, the confidence information score is updated to the first distance information of the skeletal keypoints based on the second position information of the skeletal keypoints in N video frames; the first distance information can characterize the distance of the skeletal keypoints relative to the starting point of the action range, obtained from the second position information.

[0055] When a person performs an action, the movement paths of the skeletal key points in different parts of the body are usually different. For example, in the process of a person falling forward, the movement path length of the skeletal key points of the upper limbs is usually longer than that of the skeletal key points of the lower limbs. This application embodiment uses the first distance information of the skeletal key points relative to the starting point of the movement range for the fall detection process. Because it can distinguish the movement path lengths of different skeletal key points, it can improve the difference between different actions based on the distinguishable movement path length, thus improving the accuracy of fall detection. For example, for a falling action, the movement path length of the skeletal key points of the upper limbs is usually longer than that of the skeletal key points of the lower limbs; while for a squatting action, the movement path lengths of the skeletal key points of the upper limbs and lower limbs are basically equal, etc. Therefore, this application embodiment can improve the difference between different actions based on the distinguishable movement path length.

[0056] Method Example 1

[0057] refer to Figure 1 The diagram illustrates a step-by-step flowchart of a fall detection method according to an embodiment of this application. The method may specifically include the following steps:

[0058] Step 101: Obtain N video frames to be detected; N is a positive integer;

[0059] Step 102: Determine the key point data of the skeleton key points in N video frames respectively; the key point data may include: first position information and confidence information;

[0060] Step 103: Fuse the first position information of the skeletal key points in N video frames to obtain the motion range information of the preset target in N video frames;

[0061] Step 104: Based on the motion range information, update the first position information of the skeletal key points in the N video frames to obtain the updated second position information.

[0062] Step 105: Based on the second position information of the skeletal keypoints in N video frames, update the confidence information to the first distance information of the skeletal keypoints; the first distance information represents the distance of the skeletal keypoints relative to the starting point of the action range, obtained based on the second position information.

[0063] Step 106: Based on the second position information and first distance information of the skeletal key points in N video frames, detect whether the preset target has a falling action.

[0064] In step 101, the N video frames can originate from target videos containing preset targets captured by a camera device. The preset targets can be objects with the ability to perform actions, such as pedestrians, animals, and robots.

[0065] The length of the target video can be determined by those skilled in the art based on actual application requirements. Assuming the target video is 2 seconds long and contains 25 video frames per second, the value of N can be 50. Of course, this application does not limit the specific value of N.

[0066] In step 102, skeletal key points, also known as human body key points, correspond to the joints of the human body, such as the joints of the neck, shoulder, elbow, wrist, waist, knee, and ankle. After locating and recognizing the skeletal key points, the positional information of the skeletal key points in three-dimensional space can be confirmed, and then a certain algorithm can be used to determine what kind of action the human body of the identified preset target is involved in.

[0067] In practical applications, keypoint detection models can be used to determine keypoint data of skeletal keypoints in N video frames. In one example, the keypoint detection model can first perform preset target detection on the target video to obtain the detection box of the preset target in the video frame; then, it can perform skeletal keypoint detection on the preset target in the detection box to obtain the keypoint data of the skeletal keypoints.

[0068] This application does not limit the number M of different types of skeletal key points. For example, there may be 14 or 17 types of skeletal key points.

[0069] Keypoint detection models typically output keypoint data comprising three dimensions (X, Y, and score). X represents the x-coordinate of the first location information, Y represents the y-coordinate of the first location information, and the score represents the confidence score of the skeletal keypoint. The confidence score characterizes the probability that the pixel corresponding to the skeletal keypoint is indeed a skeletal keypoint.

[0070] The first position information of a skeletal keypoint can refer to the initial position information of the skeletal keypoint in each video frame. Assuming a planar coordinate system is used, with the width of the video frame corresponding to the x-coordinate (X) and the height corresponding to the y-coordinate (Y), the coordinates corresponding to the first position information can be represented as (X, Y). The origin of the coordinate system can be determined by those skilled in the art based on actual application requirements. For example, the origin could be the lower left vertex of the rectangle corresponding to each video frame; of course, the origin could also be any other arbitrary position, such as the upper left vertex of the rectangle corresponding to each video frame, or a position outside the video frame, etc.

[0071] Reference Figure 2 The diagram illustrates the skeletal key points of N video frames according to an embodiment of this application. The correspondence between the skeletal key points and their numerical numbers is as follows: nose-0, right eye-1, left eye-2, right ear-3, left ear-4, right shoulder-5, left shoulder-6, right elbow-7, left elbow-8, right wrist-9, left wrist-10, right waist-11, left waist-12, right knee-13, left knee-14, right ankle-15, left ankle-16.

[0072] In practical applications, the first position information of the skeletal keypoints in each video frame can be within the entire video frame's size range [W, H]. For example, if W = 1920 and H = 1080, then X ∈ (0, 1920) and Y ∈ (0, 1080). During a user's action, this first position information can change over time. Taking a falling action as an example, the first position information of at least some skeletal keypoints can change in the direction of the fall as the action occurs.

[0073] Figure 2 This illustrates the first position information of the first video frame and the Nth video frame out of N video frames, assuming a falling action. For ease of explanation, Figure 2 It is believed that the skeletal keypoints with the same numerical number in the first video frame and the Nth video frame correspond to different coordinate points.

[0074] It is understandable that, in the case of a fall or other actions (such as squatting), the same numerical code can correspond to the same or different first position information in different video frames. For example, during a certain period of the fall, the position of the ankle may not change, so the numerical code 15 can correspond to the same first position information in some of the N video frames. Similarly, during a squat, the position of the ankle may remain unchanged, so the numerical code 15 can correspond to the same first position information in all of the N video frames.

[0075] In step 103, each video frame contains M skeletal keypoints. This embodiment of the application can fuse the M skeletal keypoints contained in N video frames. The above fusion can summarize the first position information of the M skeletal keypoints in the N video frames. Assuming the summarized first position information is called the summary set, the summary set can include M*N pieces of first position information. Since the same numerical code can correspond to the same first position information in different video frames, the M*N pieces of first position information can contain duplicate coordinate points. This embodiment of the application does not limit the number of times duplicate coordinate points are recorded in the summary set; they can be recorded once or multiple times.

[0076] In one implementation, the motion range information may include: horizontal coordinate range information and vertical coordinate range information. Accordingly, the process of fusing the first position information of skeletal keypoints in N video frames may specifically include: determining the lower and upper limits of the horizontal coordinate, and the lower and upper limits of the vertical coordinate, based on the first position information of the skeletal keypoints in the N video frames; determining the horizontal coordinate range information based on the lower and upper limits of the horizontal coordinate; and determining the vertical coordinate range information based on the lower and upper limits of the vertical coordinate.

[0077] like Figure 2 As shown, this embodiment of the application can sort the x-coordinates of the coordinate points in the summary set in ascending order to obtain the lower limit Xmin and the upper limit Xmax of the x-coordinates. This embodiment of the application can also sort the y-coordinates of the coordinate points in the summary set in ascending order to obtain the lower limit Ymin and the upper limit Ymax of the y-coordinates. Thus, the x-coordinate range information can be [Xmin, Xmax], and the y-coordinate range information can be [Ymin, Ymax]. The x-coordinate range information can further include: width range information W' obtained based on the difference between the lower limit Xmin and the upper limit Xmax of the x-coordinates. The y-coordinate range information can further include: height range information H' obtained based on the difference between the lower limit Ymin and the upper limit Ymax of the y-coordinates.

[0078] In another implementation, the motion range information may include the standard deviation and mean of the first position information of the skeletal keypoints in N video frames. Accordingly, the standard deviation and mean of the x-coordinates of the coordinate points in the summary set can be calculated, and the standard deviation and mean of the y-coordinates of the coordinate points in the summary set can also be calculated.

[0079] In step 104, the motion range information of the preset target in N video frames can characterize the motion range of the preset target in N video frames. Based on this motion range information, this embodiment updates the first position information of the skeletal keypoints in the N video frames. The principle for updating the skeletal keypoints is: updating the motion range of the skeletal keypoints relative to the preset target in the N video frames.

[0080] This application embodiment can utilize a normalization method to update the first position information of skeletal keypoints in the N video frames based on the motion range information. The normalization method can map both the horizontal and vertical coordinates of the first position information to the range of [0,1] or [-1,1] based on the motion range information.

[0081] In one implementation, updating the first position information of the skeletal keypoints in the N video frames may specifically include:

[0082] The horizontal coordinate of the second position information is determined based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information. For example, the horizontal coordinate of the second position information can be obtained based on the ratio of the first difference to the width range.

[0083] The ordinate of the second position information is determined based on the second difference between the ordinate of the first position information and the lower limit of the ordinate, as well as the ordinate range information. For example, the abscissa of the second position information can be obtained based on the ratio of the second difference to the height range.

[0084] In another implementation, updating the first position information of the skeletal key points in the N video frames may specifically include: determining the abscissa of the second position information based on the fourth difference between the abscissa of the first position information and the mean of the abscissa of the first position information, and the standard deviation of the abscissa of the first position information; and determining the ordinate of the second position information based on the fifth difference between the ordinate of the first position information and the mean of the ordinate of the first position information, and the standard deviation of the ordinate of the first position information.

[0085] For example, the formula for calculating the x-coordinate of the second location information is as follows: Where x represents the x-coordinate of the second position information, X represents the x-coordinate of the first position information, μ represents the mean of the x-coordinate of the first position information, and σ represents the standard deviation of the x-coordinate of the first position information.

[0086] In step 105, the second position information is the position information updated based on the position range information. The position range information may correspond to a new coordinate origin, which can be called the starting point of the range of motion. The new coordinate system may be related to the range of motion of the human body; the new coordinate origin can be the starting point of the range of motion or the point where the X and Y axes are smallest within the range of motion. Figure 2 As shown, the new coordinate origin (0,0) corresponds to Xmin and Ymin.

[0087] In this embodiment, the confidence information can be updated to the first distance information of the skeletal keypoints based on the second position information of the skeletal keypoints in N video frames and the starting point of the action range.

[0088] Assuming the coordinates corresponding to the second location information within the location range information can be represented as (x, y), then the first distance information D1 can be the product of the distance between the coordinate point corresponding to the second location information and the starting point and the adjustment parameter. The adjustment parameter can be used to adjust the range of the first distance information.

[0089] In this embodiment, the first distance information of the skeletal keypoint can be determined based on the squares of the abscissa and ordinate of the second position information, as well as adjustment parameters. For example, when x∈[0,1] and y∈[0,1], the distance between the coordinate point corresponding to the second position information and the starting point can be... The parameters can then be adjusted as follows: It is understood that the embodiments of this application do not limit the specific adjustment parameters.

[0090] In step 106, the second position information (x, y) and the first distance information D1 can be used as the updated keypoint data. In other words, the keypoint data can be updated from (X, Y, score) to (x, y, D1).

[0091] In this embodiment of the application, (x, y, D1) can be input into the fall detection model, and the fall detection model can output the corresponding detection results.

[0092] The fall detection module can employ a graph convolutional neural network. A graph convolutional neural network defines each skeletal keypoint as a graph node, the physical connection between two skeletal keypoints as an edge in the graph, and adds temporal edges between the same nodes in adjacent video frames. This allows the behavior to be determined by... Figure 2The spatiotemporal relationships in graphs are used to represent this. The input to a graph convolutional neural network is the coordinate vector corresponding to the second position information of a graph node. Graph convolutional neural networks can extract high-level features based on a series of spatiotemporal graph convolution operations and use a classifier to obtain the corresponding behavior classification. For example, the classifier corresponds to two categories: falling actions and non-falling actions.

[0093] The fall detection module can employ a gated recurrent neural network. The gated recurrent neural network extracts temporal features from the second position information of skeletal keypoints in N video frames, and inputs the output vector of the hidden layer into a fully connected layer for processing to obtain the detection result.

[0094] If the detection results indicate that the preset target has fallen, a prompt message can be sent to the relevant device. For example, in a home setting, the relevant device could be a family member's device. Or, in a public place setting, the relevant device could be the manager's device. Of course, corresponding voice prompts can also be played to attract the attention of those nearby.

[0095] In summary, the fall detection method of this application updates the first position information of skeletal key points in N video frames based on the motion range information. Since the update principle for skeletal key points in this application is based on updating the motion range of the skeletal key points relative to a preset target in N video frames, and this update principle is not limited by the distance between the preset target and the camera device, this application can reduce the impact of the distance between the preset target and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the preset target is far from the camera device, the update of the motion range of the skeletal key points relative to the preset target in N video frames in this application can avoid the problem of the updated value of the skeletal key points being too small, thus reducing the impact of the distance between the preset target and the camera device on the accuracy of fall detection.

[0096] Since this update principle is not limited by the distance between the preset target and the camera device, the embodiments of this application can reduce the requirements for training data and detection data. For example, related technologies usually require that the training data and detection data correspond to the same application scenario, while the embodiments of this application do not require that the training data and detection data correspond to the same application scenario; thus, the embodiments of this application can reduce the difficulty of obtaining training data, thereby increasing the richness of the training data.

[0097] Furthermore, in this embodiment, the confidence score is updated to the first distance information of the skeletal keypoints based on the second position information of the skeletal keypoints in N video frames, and the first distance information of the skeletal keypoints relative to the starting point of the action range is used in the fall detection process. Since the first distance information can be used to distinguish the motion path lengths of different skeletal keypoints, this embodiment can improve the difference between different actions based on the distinguishable motion path length, thereby improving the accuracy of fall detection. For example, for a fall, the motion path length of the upper limb skeletal keypoints is usually greater than that of the lower limb skeletal keypoints; while for a squatting action, the motion path lengths of the upper limb skeletal keypoints are roughly equivalent to those of the lower limb skeletal keypoints, etc. Therefore, this embodiment can improve the difference between different actions based on the distinguishable motion path length.

[0098] Method Example 2

[0099] refer to Figure 3 The diagram illustrates a step-by-step flowchart of a fall detection method according to an embodiment of this application. The method may specifically include the following steps:

[0100] Step 301: Obtain N video frames to be detected; N is a positive integer;

[0101] Step 302: Determine the key point data of the skeleton key points in N video frames respectively; the key point data may include: first position information and confidence information;

[0102] Step 303: Fuse the first position information of the skeletal key points in N video frames to obtain the motion range information of the preset target in N video frames;

[0103] Step 304: Based on the motion range information, update the first position information of the skeletal key points in the N video frames to obtain the updated second position information;

[0104] Step 305: Determine the centroid position of the preset target in the i-th video frame based on the second position information of the preset skeletal key points in the i-th video frame; i can be a positive integer.

[0105] Step 306: Based on the center of gravity position of the preset target in the i-th video frame, update the second position information of the skeletal key points in the i-th video frame to obtain the third position information of the skeletal key points in the i-th video frame.

[0106] Step 307: Based on the third position information of the skeletal keypoints in N video frames, update the confidence information to the second distance information of the skeletal keypoints; the second distance information specifically represents the distance of the skeletal keypoints relative to the starting point of the action range, obtained based on the third position information.

[0107] Step 308: Based on the third position information and second distance information of the skeletal key points in N video frames, detect whether the preset target has a falling action.

[0108] In this embodiment of the application, when updating the first position information of the skeletal keypoints in the N video frames to the second position information, the second position information of the skeletal keypoints in the i-th video frame is also updated according to the centroid position of the preset target in the i-th video frame. Wherein, 1≤i≤N.

[0109] Human skeletal key points typically represent the positional information of a single joint. However, multiple skeletal key points in a person should be closely related. Related technologies primarily utilize the positional information of skeletal key points without considering or using the relationships between them.

[0110] In this embodiment, updating the skeletal keypoints of a video frame relative to the center of gravity results in more closely related skeletal keypoints. Thus, during fall detection, the coordinate changes of multiple skeletal keypoints are treated as a whole, rather than as scattered independent points. This allows the fall detection model to capture the intrinsic connection information of the skeletal keypoints, thereby improving the accuracy of fall detection.

[0111] This application embodiment can determine the center of gravity position of a preset target in the i-th video frame based on the second position information (x, y) of preset skeletal keypoints in the i-th video frame. The preset skeletal keypoints can be determined by those skilled in the art according to actual application requirements. For example, the preset skeletal keypoints may include four skeletal keypoints: right shoulder-5, left shoulder-6, right waist-11, and left waist-12. This application embodiment can average the second position information of the four skeletal keypoints to obtain the center of gravity position. Specifically, the x-coordinate of the center of gravity position can be the average of the x-coordinates of the second position information of the four skeletal keypoints, and the y-coordinate of the center of gravity position can be the average of the y-coordinates of the second position information of the four skeletal keypoints.

[0112] The process of updating the second position information of the skeletal keypoints in the i-th video frame can specifically include: determining the third position information (x', y') of the skeletal keypoints in the i-th video frame based on the third difference between the second position information of the skeletal keypoints in the i-th video frame and the centroid position of the preset target. The horizontal coordinate of the third position information can be the difference between the horizontal coordinate of the second position information and the horizontal coordinate of the centroid position, and the vertical coordinate of the third position information can be the difference between the vertical coordinate of the second position information and the vertical coordinate of the centroid position.

[0113] When updating the second position information of the skeletal keypoints in the N video frames to the third position information, the embodiments of this application can update the confidence information to the second distance information of the skeletal keypoints based on the third position information (x', y') of the skeletal keypoints in the N video frames; the second distance information specifically represents the distance of the skeletal keypoints relative to the starting point of the action range, obtained based on the third position information.

[0114] In this embodiment, the confidence information can be updated to the second distance information of the skeletal keypoints based on the third position information of the skeletal keypoints in N video frames and the starting point of the action range.

[0115] Assuming the coordinates of the third location information corresponding to the location range information can be represented as (x', y'), then the second distance information D2 can be the product of the distance between the coordinate point corresponding to the third location information and the starting point and the adjustment parameter. The adjustment parameter can be used to adjust the range of the first distance information.

[0116] In this embodiment, the first distance information of the skeletal keypoint can be determined based on the squares of the abscissa and ordinate of the second position information, as well as adjustment parameters. For example, when x'∈[0,1] and y'∈[0,1], the distance between the coordinate point corresponding to the second position information and the starting point can be... The parameters can then be adjusted as follows: It is understood that the embodiments of this application do not limit the specific adjustment parameters.

[0117] In this embodiment, the third location information (x', y') and the second distance information D2 can be used as the updated keypoint data. In other words, the keypoint data can be updated from (X, Y, score) to (x', y', D2) first.

[0118] In this embodiment of the application, (x', y', D2) can be input into the fall detection model, and the fall detection model can output the corresponding detection results.

[0119] In summary, the fall detection method of this application updates the first position information of skeletal key points in N video frames based on the motion range information. Since the update principle for skeletal key points in this application is based on updating the motion range of the skeletal key points relative to a preset target in N video frames, and this update principle is not limited by the distance between the preset target and the camera device, this application can reduce the impact of the distance between the preset target and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the preset target is far from the camera device, the update of the motion range of the skeletal key points relative to the preset target in N video frames in this application can avoid the problem of the updated value of the skeletal key points being too small, thus reducing the impact of the distance between the preset target and the camera device on the accuracy of fall detection.

[0120] Furthermore, by updating the skeletal keypoints of a video frame relative to the center of gravity in this embodiment, more closely related skeletal keypoints can be obtained. In this way, during the fall detection process, the coordinate changes of multiple skeletal keypoints will be regarded as a whole, rather than scattered independent points. This allows the fall detection model to capture the intrinsic connection information of the skeletal keypoints, thereby improving the accuracy of fall detection.

[0121] Furthermore, in this embodiment, the confidence score is updated to the second distance information of the skeletal keypoints based on the third position information of the skeletal keypoints in N video frames, and the second distance information of the skeletal keypoints relative to the starting point of the movement range is used in the fall detection process. Since the second distance information can be used to distinguish the motion path lengths of different skeletal keypoints, this embodiment can improve the difference between different movements based on the distinguishable motion path length, thereby improving the accuracy of fall detection. For example, for a fall, the motion path length of the upper limb skeletal keypoints is usually greater than that of the lower limb skeletal keypoints; while for a squatting movement, the motion path lengths of the upper limb skeletal keypoints are roughly equivalent to those of the lower limb skeletal keypoints, etc. Therefore, this embodiment can improve the difference between different movements based on the distinguishable motion path length.

[0122] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.

[0123] Based on the above embodiments, this embodiment also provides a fall detection device, referring to... Figure 4 The device may specifically include: a video frame acquisition module 401, a key point determination module 402, a fusion module 403, a first update module 404, a second update module 405, and a first detection module 406.

[0124] Among them, the video frame acquisition module 401 is used to acquire N video frames to be detected; N is a positive integer;

[0125] The key point determination module 402 is used to determine key point data of skeletal key points in N video frames respectively; the key point data includes: first location information and confidence information;

[0126] The fusion module 403 is used to fuse the first position information of the skeletal key points in N video frames to obtain the motion range information of the preset target in the N video frames.

[0127] The first update module 404 is used to update the first position information of the skeletal key points in the N video frames according to the motion range information, so as to obtain the updated second position information.

[0128] The second update module 405 is used to update the confidence information to the first distance information of the skeletal key points based on the second position information of the skeletal key points in N video frames; the first distance information represents the distance of the skeletal key point relative to the starting point of the action range, obtained based on the second position information.

[0129] The first detection module 406 is used to detect whether a preset target has a falling action based on the second position information and the first distance information of the skeletal key points in N video frames.

[0130] Optionally, the fusion module 403 may specifically include:

[0131] The upper and lower limit determination module is used to determine the lower and upper limits of the horizontal coordinate, as well as the lower and upper limits of the vertical coordinate, based on the first position information of the skeletal key points in N video frames.

[0132] The coordinate range determination module is used to determine the range information of the horizontal coordinate based on the lower limit and the upper limit of the horizontal coordinate, and to determine the range information of the vertical coordinate based on the lower limit and the upper limit of the vertical coordinate.

[0133] Optionally, the first update module 404 may specifically include:

[0134] The horizontal coordinate update module is used to determine the horizontal coordinate of the second position information based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information.

[0135] The horizontal coordinate update module is used to determine the vertical coordinate of the second position information based on the second difference between the vertical coordinate of the first position information and the lower limit of the vertical coordinate, as well as the vertical coordinate range information.

[0136] Optionally, the device may further include:

[0137] The distance determination module is used to determine the first distance information of the skeletal key points based on the squares of the horizontal and vertical coordinates of the second position information and adjustment parameters.

[0138] Optionally, the device may further include:

[0139] The center of gravity determination module is used to determine the center of gravity position of the pedestrian in the i-th video frame based on the second position information of the preset skeletal key points in the i-th video frame; i is a positive integer;

[0140] The third update module is used to update the second position information of the skeletal key points in the i-th video frame according to the center position of the preset target in the i-th video frame, so as to obtain the third position information of the skeletal key points in the i-th video frame.

[0141] The fourth update module is used to update the confidence information to the second distance information of the skeletal keypoints based on the third position information of the skeletal keypoints in N video frames; the second distance information represents the distance of the skeletal keypoint relative to the starting point of the action range, obtained based on the third position information.

[0142] The second detection module is used to detect whether a preset target has a falling action based on the third position information and the second distance information of the skeletal key points in N video frames.

[0143] Optionally, the third update module is specifically used to determine the third position information of the skeletal key points in the i-th video frame based on the third difference between the second position information of the skeletal key points in the i-th video frame and the center of gravity position of the preset target.

[0144] In summary, the fall detection device of this application updates the first position information of skeletal key points in N video frames based on the motion range information. Since the updating principle of the skeletal key points in this application is based on updating the motion range of the skeletal key points relative to a preset target in N video frames, and since this updating principle is not limited by the distance between the preset target and the camera device, this application can reduce the impact of the distance between the preset target and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the preset target is far from the camera device, the updating of the motion range of the skeletal key points relative to the preset target in N video frames in this application can avoid the problem of the updated value of the skeletal key points being a small value, thus reducing the impact of the distance between the preset target and the camera device on the accuracy of fall detection.

[0145] Since this update principle is not limited by the distance between the preset target and the camera device, the embodiments of this application can reduce the requirements for training data and detection data. For example, related technologies usually require that the training data and detection data correspond to the same application scenario, while the embodiments of this application do not require that the training data and detection data correspond to the same application scenario; thus, the embodiments of this application can reduce the difficulty of obtaining training data, thereby increasing the richness of the training data.

[0146] Furthermore, in this embodiment, the confidence score is updated to the first distance information of the skeletal keypoints based on the second position information of the skeletal keypoints in N video frames, and the first distance information of the skeletal keypoints relative to the starting point of the action range is used in the fall detection process. Since the first distance information can be used to distinguish the motion path lengths of different skeletal keypoints, this embodiment can improve the difference between different actions based on the distinguishable motion path length, thereby improving the accuracy of fall detection. For example, for a fall, the motion path length of the upper limb skeletal keypoints is usually greater than that of the lower limb skeletal keypoints; while for a squatting action, the motion path lengths of the upper limb skeletal keypoints are roughly equivalent to those of the lower limb skeletal keypoints, etc. Therefore, this embodiment can improve the difference between different actions based on the distinguishable motion path length.

[0147] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.

[0148] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In this application, the electronic device includes various types of devices such as terminal devices and servers (clusters).

[0149] The embodiments of this disclosure can be implemented as an apparatus configured as desired using any suitable hardware, firmware, software, or any combination thereof, including electronic devices such as terminal devices and servers (clusters). Figure 5 An exemplary apparatus 1100 is schematically shown that can be used to implement the various embodiments described in this application.

[0150] In one embodiment, Figure 5 An exemplary device 1100 is shown, which includes one or more processors 1102, a control module (chipset) 1104 coupled to at least one of the processors 1102, a memory 1106 coupled to the control module 1104, a non-volatile memory (NVM) / storage device 1108 coupled to the control module 1104, one or more input / output devices 1110 coupled to the control module 1104, and a network interface 1112 coupled to the control module 1104.

[0151] Processor 1102 may include one or more single-core or multi-core processors, and processor 1102 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1100 can serve as a terminal device, server (cluster), or other device as described in the embodiments of this application.

[0152] In some embodiments, apparatus 1100 may include one or more computer-readable media (e.g., memory 1106 or NVM / storage device 1108) having instructions 1114 and one or more processors 1102 that are combined with the one or more computer-readable media and configured to execute instructions 1114 to implement a module thereby performing the actions described in this disclosure.

[0153] In one embodiment, the control module 1104 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1102 and / or any suitable device or component communicating with the control module 1104.

[0154] The control module 1104 may include a memory controller module to provide an interface to the memory 1106. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0155] Memory 1106 may be used, for example, to load and store data and / or instructions 1114 for device 1100. In one embodiment, memory 1106 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1106 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0156] In one embodiment, the control module 1104 may include one or more input / output controllers to provide interfaces to the NVM / storage device 1108 and (one or more) input / output devices 1110.

[0157] For example, NVM / storage device 1108 may be used to store data and / or instructions 1114. NVM / storage device 1108 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).

[0158] NVM / storage device 1108 may include storage resources that are physically part of a device on which device 1100 is mounted, or that can be accessed by the device without needing to be part of the device. For example, NVM / storage device 1108 may be accessed via a network via one or more input / output devices 1110.

[0159] One or more input / output devices 1110 may provide an interface for device 1100 to communicate with any other suitable device. Input / output devices 1110 may include communication components, audio components, sensor components, etc. Network interface 1112 may provide an interface for device 1100 to communicate via one or more networks. Device 1100 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.

[0160] In one embodiment, at least one of the processors 1102 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1104. In one embodiment, at least one of the processors 1102 may be logically packaged with one or more controllers of the control module 1104 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1102 may be integrated with the logic of one or more controllers of the control module 1104 on the same die. In one embodiment, at least one of the processors 1102 may be integrated with the logic of one or more controllers of the control module 1104 on the same die to form a system-on-a-chip (SoC).

[0161] In various embodiments, device 1100 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, device 1100 may have more or fewer components and / or different architectures. For example, in some embodiments, device 1100 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0162] The detection device can use a main control chip as a processor or control module, and sensor data, position information, etc. can be stored in a memory or NVM / storage device. The sensor group can be used as an input / output device, and the communication interface can include a network interface.

[0163] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0164] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0165] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0166] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.

[0167] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.

[0168] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0169] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0170] The foregoing has provided a detailed description of a fall detection method and apparatus, an electronic device, and a machine-readable medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A fall detection method, characterized in that, The method includes: Obtain N video frames to be detected; N is a positive integer; Key point data of skeletal key points in N video frames are determined respectively; the key point data includes: first location information and confidence information; The first position information of the skeletal key points in N video frames is fused to obtain the motion range information of the preset target in N video frames; Based on the motion range information, the first position information of the skeletal key points in the N video frames is updated to obtain the updated second position information; Based on the second position information of skeletal keypoints in N video frames, the confidence information is updated to the first distance information of the skeletal keypoints. The first distance information represents the distance of the skeletal keypoint relative to the starting point of the action range, obtained from the second position information. The first distance information is the product of the distance between the coordinate point corresponding to the second position information and the starting point and an adjustment parameter, wherein the adjustment parameter is used to adjust the range of the first distance information. The update refers to updating the keypoint data from (X, Y, score) to (x, y, D1). Wherein, X represents the abscissa of the first position information, Y represents the ordinate of the first position information, score represents the confidence information of the skeletal keypoint, x represents the abscissa of the second position information, y represents the ordinate of the second position information, and D1 represents the first distance information. Based on the second position information and first distance information of the skeletal key points in N video frames, detect whether the preset target has a falling action; The fusion of the first position information of skeletal key points in N video frames includes: Based on the first position information of the skeletal key points in N video frames, determine the lower limit and upper limit of the horizontal coordinate, as well as the lower limit and upper limit of the vertical coordinate. The range of the horizontal axis is determined based on its lower and upper limits, and the range of the vertical axis is determined based on its lower and upper limits.

2. The method according to claim 1, characterized in that, The step of updating the first position information of the skeletal key points in the N video frames includes: The horizontal coordinate of the second position information is determined based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information. The ordinate of the second position information is determined based on the second difference between the ordinate of the first position information and the lower limit of the ordinate, as well as the ordinate range information.

3. The method according to claim 1, characterized in that, The method further includes: Based on the squares of the horizontal and vertical coordinates of the second location information, and the adjustment parameters, the first distance information of the skeletal key points is determined.

4. The method according to claim 1, characterized in that, The method further includes: Based on the second position information of the preset skeletal key points in the i-th video frame, determine the centroid position of the preset target in the i-th video frame; i is a positive integer; Based on the centroid position of the preset target in the i-th video frame, the second position information of the skeletal key points in the i-th video frame is updated to obtain the third position information of the skeletal key points in the i-th video frame. Based on the third position information of the skeletal keypoints in N video frames, the confidence information is updated to the second distance information of the skeletal keypoints; the second distance information represents the distance of the skeletal keypoint relative to the starting point of the action range, obtained based on the third position information. Based on the third position information and second distance information of the skeletal key points in N video frames, detect whether the preset target has a falling action.

5. The method according to claim 4, characterized in that, The update of the second position information of the skeletal keypoints in the i-th video frame includes: The third position information of the skeletal keypoints in the i-th video frame is determined based on the third difference between the second position information of the skeletal keypoints in the i-th video frame and the centroid position of the preset target.

6. A fall detection device, characterized in that, The device includes: The video frame acquisition module is used to acquire N video frames to be detected; N is a positive integer. A key point determination module is used to determine key point data of skeletal key points in N video frames respectively; the key point data includes: first location information and confidence information; The fusion module is used to fuse the first position information of the skeletal key points in N video frames to obtain the motion range information of the preset target in N video frames; The first update module is used to update the first position information of the skeletal key points in the N video frames according to the motion range information, so as to obtain the updated second position information. The second update module is used to update the confidence information to the first distance information of the skeletal keypoints based on the second position information of the skeletal keypoints in N video frames. The first distance information represents the distance of the skeletal keypoint relative to the starting point of the action range, obtained based on the second position information. The first distance information is the product of the distance between the coordinate point corresponding to the second position information and the starting point and an adjustment parameter, wherein the adjustment parameter is used to adjust the range of the first distance information. The update refers to updating the keypoint data from (X, Y, score) to (x, y, D1). Wherein, X represents the abscissa of the first position information, Y represents the ordinate of the first position information, score represents the confidence information of the skeletal keypoint, x represents the abscissa of the second position information, y represents the ordinate of the second position information, and D1 represents the first distance information. The first detection module is used to detect whether a preset target has a falling action based on the second position information and the first distance information of the skeletal key points in N video frames. The fusion module includes: The upper and lower limit determination module is used to determine the lower and upper limits of the horizontal coordinate, as well as the lower and upper limits of the vertical coordinate, based on the first position information of the skeletal key points in N video frames. The coordinate range determination module is used to determine the range information of the horizontal coordinate based on the lower limit and the upper limit of the horizontal coordinate, and to determine the range information of the vertical coordinate based on the lower limit and the upper limit of the vertical coordinate.

7. An electronic device, characterized in that, include: processor; and A memory having executable code stored thereon, which, when executed, causes the processor to perform the method as described in any one of claims 1-5.

8. A machine-readable medium having executable code stored thereon, which, when executed, causes a processor to perform the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Behavior detection method, device and system

    CN112686075A

  • Person falling detection method and device and electronic equipment

    CN112766168A