Fall detection method, apparatus, device, and medium
Patent Information
- Application Number
- CN202211738983.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-12-30
AI Technical Summary
[0004]在实际应用中,行人距离摄像设备的远近,容易影响摔倒检测的准确度
[0033]在本申请实施例的技术方案中,行人在N个视频帧中的位置范围信息可以表征行人在N个视频帧中的活动范围。本申请实施例根据该位置范围信息,对该N个视频帧中骨骼关键点的第一位置信息进行更新。由于本申请实施例针对骨骼关键点的更新原理是:骨骼关键点相对于行人在N个视频帧中的活动范围的更新,由于该更新原理可以不受行人距离摄像设备的远近的限制,故本申请实施例能够降低行人距离摄像设备的远近对于摔倒检测的准确度的影响,进而能够提高摔倒检测的鲁棒性。例如,在行人距离摄像设备较远的情况下,本申请实施例中骨骼关键点相对于行人在N个视频帧中的活动范围的更新,可以避免出现骨骼关键点对应的更新后数值为较小数值的问题,因此能够降低行人距离摄像设备远对于摔倒检测的准确度的影响。
Smart Images

Figure CN118279814B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a fall detection method, apparatus, device, and medium. Background Technology
[0002] In video surveillance scenarios, timely detection of pedestrian falls and alerting relevant personnel can effectively mitigate the consequences of accidental falls, improve service quality in relevant locations (such as shopping malls and subways), and better protect pedestrian safety.
[0003] Current fall detection methods typically use camera equipment to capture target videos containing pedestrians, identify skeletal key points in the video frames, and detect whether the pedestrian has fallen based on the skeletal key points.
[0004] In practical applications, the distance between pedestrians and the camera equipment can easily affect the accuracy of fall detection. For example, when pedestrians are far from the camera equipment, the accuracy of fall detection is lower; while when pedestrians are close to the camera equipment, the accuracy of fall detection is higher. Summary of the Invention
[0005] This application provides a fall detection method that can improve the robustness of fall detection.
[0006] Accordingly, embodiments of this application also provide a fall detection device, an electronic device, and a machine-readable medium to ensure the implementation and application of the above methods.
[0007] To address the aforementioned problems, this application discloses a fall detection method, the method comprising:
[0008] Obtain N video frames to be detected; N is a positive integer;
[0009] Determine the first position information of the skeletal keypoints in N video frames respectively;
[0010] The first position information of the skeletal key points in N video frames is fused to obtain the position range information of the pedestrian in N video frames;
[0011] Based on the location range information, the first location information of the skeletal key points in the N video frames is updated to obtain the updated second location information;
[0012] Based on the second position information of skeletal key points in N video frames, detect whether a pedestrian has fallen.
[0013] To address the aforementioned problems, this application discloses a fall detection device, the device comprising:
[0014] The video frame acquisition module is used to acquire N video frames to be detected; N is a positive integer.
[0015] The key point determination module is used to determine the first position information of the skeletal key points in N video frames respectively;
[0016] The fusion module is used to fuse the first position information of the skeletal key points in N video frames to obtain the position range information of the pedestrian in N video frames.
[0017] The first update module is used to update the first position information of the skeletal key points in the N video frames according to the position range information, so as to obtain the updated second position information.
[0018] The first detection module is used to detect whether a pedestrian has fallen based on the second position information of the skeletal key points in N video frames.
[0019] Optionally, the fusion module includes:
[0020] The upper and lower limit determination module is used to determine the lower and upper limits of the horizontal coordinate, as well as the lower and upper limits of the vertical coordinate, based on the first position information of the skeletal key points in N video frames.
[0021] The coordinate range determination module is used to determine the range information of the horizontal coordinate based on the lower limit and the upper limit of the horizontal coordinate, and to determine the range information of the vertical coordinate based on the lower limit and the upper limit of the vertical coordinate.
[0022] Optionally, the first update module includes:
[0023] The horizontal coordinate update module is used to determine the horizontal coordinate of the second position information based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information.
[0024] The horizontal coordinate update module is used to determine the vertical coordinate of the second position information based on the second difference between the vertical coordinate of the first position information and the lower limit of the vertical coordinate, as well as the vertical coordinate range information.
[0025] Optionally, the device further includes:
[0026] The center of gravity determination module is used to determine the center of gravity position of the pedestrian in the i-th video frame based on the second position information of the preset skeletal key points in the i-th video frame; i is a positive integer;
[0027] The second update module is used to update the second position information of the skeletal key points in the i-th video frame based on the center of gravity position of the pedestrian in the i-th video frame, so as to obtain the third position information of the skeletal key points in the i-th video frame.
[0028] The second detection module is used to detect whether a pedestrian has fallen based on the third position information of the skeletal key points in N video frames.
[0029] Optionally, the second update module is specifically used to determine the third position information of the skeletal key points in the i-th video frame based on the third difference between the second position information of the skeletal key points in the i-th video frame and the position of the pedestrian's center of gravity.
[0030] This application also discloses an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform the method described in this application.
[0031] This application also discloses a machine-readable medium storing executable code thereon, which, when executed, causes a processor to perform the method described in this application.
[0032] The embodiments of this application have the following advantages:
[0033] In the technical solution of this application embodiment, the position range information of a pedestrian in N video frames can characterize the pedestrian's activity range in N video frames. Based on this position range information, this application embodiment updates the first position information of skeletal key points in the N video frames. Since the updating principle of the skeletal key points in this application embodiment is based on the update of the skeletal key points relative to the pedestrian's activity range in N video frames, and since this updating principle is not limited by the distance between the pedestrian and the camera device, this application embodiment can reduce the impact of the distance between the pedestrian and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the pedestrian is far from the camera device, the update of the skeletal key points relative to the pedestrian's activity range in N video frames in this application embodiment can avoid the problem of the updated value of the skeletal key points being a small value, thus reducing the impact of the distance between the pedestrian and the camera device on the accuracy of fall detection. Attached Figure Description
[0034] Figure 1 This is a schematic flowchart of the fall detection method according to an embodiment of this application;
[0035] Figure 2 This is a schematic diagram of the skeletal key points of N video frames according to an embodiment of this application;
[0036] Figure 3 This is a schematic flowchart of the fall detection method according to an embodiment of this application;
[0037] Figure 4This is a schematic diagram of the structure of a fall detection device according to an embodiment of this application;
[0038] Figure 5 This is a schematic diagram of the structure of an apparatus provided in one embodiment of this application. Detailed Implementation
[0039] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0040] The embodiments of this application can be applied to fall detection scenarios. In fall detection scenarios, camera equipment can be used to capture target videos containing pedestrians and detect whether there is a fall in the video.
[0041] Fall detection scenarios can include home settings or public places. In home settings, accidental falls can cause significant physical and psychological harm to the elderly, and timely detection and assistance for fallen elderly individuals can greatly reduce disability and mortality rates. In public places, timely detection of pedestrian falls and alerting relevant personnel can improve service quality in public places (such as shopping malls and subways) and better protect pedestrian safety. It is understood that the embodiments of this application do not limit the specific fall detection scenarios.
[0042] Current fall detection methods identify skeletal keypoints in video frames and detect whether a pedestrian has fallen based on these keypoints. In this process, the skeletal keypoints are typically normalized first, specifically by updating their dimensions relative to the width and height of the video frame. A technical problem with this normalization is that when the pedestrian is far from the camera, the normalized values for the skeletal keypoints tend to be smaller, which can negatively impact the accuracy of fall detection.
[0043] To address the technical problem that the distance between pedestrians and camera equipment can affect the accuracy of fall detection, this application provides a fall detection method. Specifically, the method may include: acquiring N video frames to be detected; N can be a positive integer; determining the first position information of skeletal key points in each of the N video frames; fusing the first position information of the skeletal key points in the N video frames to obtain the position range information of the pedestrian in the N video frames; updating the first position information of the skeletal key points in the N video frames based on the position range information to obtain updated second position information; and detecting whether the pedestrian has fallen based on the second position information of the skeletal key points in the N video frames.
[0044] In this embodiment, the position range information of a pedestrian in N video frames can characterize the pedestrian's activity range in those N video frames. Based on this position range information, this embodiment updates the first position information of skeletal keypoints in the N video frames. Since the updating principle for skeletal keypoints in this embodiment is based on the update of the skeletal keypoints relative to the pedestrian's activity range in the N video frames, and this updating principle is not limited by the distance between the pedestrian and the camera, this embodiment can reduce the impact of the pedestrian's distance from the camera on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the pedestrian is far from the camera, the updating of the skeletal keypoints relative to the pedestrian's activity range in the N video frames in this embodiment can avoid the problem of the updated values of the skeletal keypoints being too small, thus reducing the impact of the distance between the pedestrian and the camera on the accuracy of fall detection.
[0045] Method Example 1
[0046] refer to Figure 1 The diagram illustrates a step-by-step flowchart of a fall detection method according to an embodiment of this application. The method may specifically include the following steps:
[0047] Step 101: Obtain N video frames to be detected; N can be a positive integer;
[0048] Step 102: Determine the first position information of the skeletal key points in N video frames respectively;
[0049] Step 103: Fuse the first position information of the skeletal key points in N video frames to obtain the position range information of the pedestrian in N video frames;
[0050] Step 104: Based on the location range information, update the first location information of the skeletal key points in the N video frames to obtain the updated second location information;
[0051] Step 105: Detect whether a pedestrian has fallen based on the second position information of the skeletal key points in N video frames.
[0052] In step 101, the N video frames can originate from target videos containing pedestrians captured by a camera device.
[0053] The length of the target video can be determined by those skilled in the art based on actual application requirements. Assuming the target video is 2 seconds long and contains 25 video frames per second, the value of N can be 50. Of course, this application does not limit the specific value of N.
[0054] In step 102, skeletal key points, also known as human body key points, correspond to the joints of the human body, such as the joints of the neck, shoulder, elbow, wrist, waist, knee, and ankle. After the skeletal key points are located and identified, certain algorithms can be used to determine what kind of movements the pedestrian is involved in.
[0055] In practical applications, pedestrian detection can be performed on the target video first to obtain the detection box of the pedestrian in the video frame; then, skeletal keypoint detection can be performed on the pedestrian in the detection box to obtain the first position information of the skeletal keypoint.
[0056] This application does not limit the number M of different types of skeletal key points. For example, there may be 14 or 17 types of skeletal key points.
[0057] The first position information of a skeletal keypoint can refer to the initial position information of the skeletal keypoint in each video frame. Assuming a planar coordinate system is used, with the width of the video frame corresponding to the x-coordinate (X) and the height corresponding to the y-coordinate (Y), the coordinates corresponding to the first position information can be represented as (X1, Y1). The origin of the coordinate system can be determined by those skilled in the art based on actual application requirements. For example, the origin could be the lower left vertex of the rectangle corresponding to each video frame; of course, the origin could also be any other arbitrary position, such as the upper left vertex of the rectangle corresponding to each video frame, or a position outside the video frame, etc.
[0058] Reference Figure 2 The diagram illustrates the skeletal key points of N video frames according to an embodiment of this application. The correspondence between the skeletal key points and their numerical numbers is as follows: nose-0, right eye-1, left eye-2, right ear-3, left ear-4, right shoulder-5, left shoulder-6, right elbow-7, left elbow-8, right wrist-9, left wrist-10, right waist-11, left waist-12, right knee-13, left knee-14, right ankle-15, left ankle-16.
[0059] In practical applications, the first position information of the skeletal keypoints in each video frame can be within the entire video frame's size range [W, H], for example, W = 1920, H = 1080, etc. During a user's action, the first position information can change over time. Taking a falling action as an example, the first position information of at least some skeletal keypoints can change in the direction of the fall as the action occurs.
[0060] Figure 2 This illustrates the first position information of the first video frame and the Nth video frame out of N video frames, assuming a falling action. For ease of explanation, Figure 2 It is believed that the skeletal keypoints with the same numerical number in the first video frame and the Nth video frame correspond to different coordinate points.
[0061] It is understandable that, in the case of a fall or other actions (such as squatting), the same numerical code can correspond to the same or different first position information in different video frames. For example, during a certain period of the fall, the position of the ankle may not change, so the numerical code 15 can correspond to the same first position information in some of the N video frames. Similarly, during a squat, the position of the ankle may remain unchanged, so the numerical code 15 can correspond to the same first position information in all of the N video frames.
[0062] In step 103, each video frame contains M skeletal keypoints. This embodiment of the application can fuse the M skeletal keypoints contained in N video frames. The above fusion can summarize the first position information of the M skeletal keypoints in the N video frames. Assuming the summarized first position information is called the summary set, the summary set can include M*N pieces of first position information. Since the same numerical code can correspond to the same first position information in different video frames, the M*N pieces of first position information can contain duplicate coordinate points. This embodiment of the application does not limit the number of times duplicate coordinate points are recorded in the summary set; they can be recorded once or multiple times.
[0063] In one implementation, the location range information may include: horizontal coordinate range information and vertical coordinate range information. Accordingly, the process of fusing the first location information of skeletal keypoints in N video frames may specifically include: determining the lower and upper limits of the horizontal coordinates, and the lower and upper limits of the vertical coordinates, based on the first location information of the skeletal keypoints in the N video frames; determining the horizontal coordinate range information based on the lower and upper limits of the horizontal coordinates; and determining the vertical coordinate range information based on the lower and upper limits of the vertical coordinates.
[0064] like Figure 2 As shown, in this embodiment, the horizontal coordinates of the coordinate points in the summary set can be sorted in ascending order to obtain the lower limit Xmin and the upper limit Xmax of the horizontal coordinates. Similarly, the vertical coordinates of the coordinate points in the summary set can be sorted in ascending order to obtain the lower limit Ymin and the upper limit Ymax of the vertical coordinates. Thus, the horizontal coordinate range information can be [Xmin, Xmax], and the vertical coordinate range information can be [Ymin, Ymax]. The horizontal coordinate range information may further include: width range information W obtained based on the difference between the lower limit Xmin and the upper limit Xmax of the horizontal coordinates. The vertical coordinate range information may further include: height range information H obtained based on the difference between the lower limit Ymin and the upper limit Ymax of the vertical coordinates.
[0065] In another implementation, the location range information may include the standard deviation and mean of the first location information of the skeletal keypoints in N video frames. Accordingly, the standard deviation and mean of the x-coordinates of the coordinate points in the summary set can be calculated, and the standard deviation and mean of the y-coordinates of the coordinate points in the summary set can also be calculated.
[0066] In step 104, the pedestrian's position range information in N video frames can characterize the pedestrian's activity range in N video frames. Based on this position range information, this embodiment updates the first position information of the skeletal key points in the N video frames. The principle behind updating the skeletal key points is: updating the skeletal key points relative to the pedestrian's activity range in the N video frames.
[0067] This application embodiment can utilize a normalization method to update the first position information of skeletal keypoints in the N video frames based on the position range information. The normalization method can map the first position information to the range [0,1] or [-1,1] based on the position range information.
[0068] In one implementation, updating the first position information of the skeletal keypoints in the N video frames may specifically include:
[0069] The horizontal coordinate of the second position information is determined based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information. For example, the horizontal coordinate of the second position information can be obtained based on the ratio of the first difference to the width range.
[0070] The ordinate of the second position information is determined based on the second difference between the ordinate of the first position information and the lower limit of the ordinate, as well as the ordinate range information. For example, the abscissa of the second position information can be obtained based on the ratio of the second difference to the height range.
[0071] In another implementation, updating the first position information of the skeletal key points in the N video frames may specifically include: determining the abscissa of the second position information based on the fourth difference between the abscissa of the first position information and the mean of the abscissa of the first position information, and the standard deviation of the abscissa of the first position information; and determining the ordinate of the second position information based on the fifth difference between the ordinate of the first position information and the mean of the ordinate of the first position information, and the standard deviation of the ordinate of the first position information.
[0072] For example, the formula for calculating the x-coordinate of the second location information is as follows: Where x represents the x-coordinate of the second position information, X represents the x-coordinate of the first position information, μ represents the mean of the x-coordinate of the first position information, and σ represents the standard deviation of the x-coordinate of the first position information.
[0073] In step 105, the second position information of the skeletal key points in N video frames can be input into the fall detection model, and the fall detection model can output the corresponding detection results.
[0074] The fall detection module can employ a graph convolutional neural network. A graph convolutional neural network defines each skeletal keypoint as a graph node, the physical connection between two skeletal keypoints as an edge in the graph, and adds temporal edges between the same nodes in adjacent video frames. This allows the behavior to be determined by... Figure 2 The spatiotemporal relationships in graphs are used to represent this. The input to a graph convolutional neural network is the coordinate vector corresponding to the second position information of a graph node. Graph convolutional neural networks can extract high-level features based on a series of spatiotemporal graph convolution operations and use a classifier to obtain the corresponding behavior classification. For example, the classifier corresponds to two categories: falling behavior and non-falling behavior.
[0075] The fall detection module can employ a gated recurrent neural network. The gated recurrent neural network extracts temporal features from the second position information of skeletal keypoints in N video frames, and inputs the output vector of the hidden layer into a fully connected layer for processing to obtain the detection result.
[0076] If the detection results indicate that a pedestrian has fallen, a notification message can be sent to the relevant device. For example, in a home setting, the device could be a family member's device. Or, in a public place, the device could be the manager's device. Alternatively, a corresponding voice notification can be played to attract the attention of those nearby.
[0077] In summary, the fall detection method of this application updates the first position information of skeletal key points in N video frames based on the position range information. Since the update principle of the skeletal key points in this application is based on the update of the skeletal key points relative to the pedestrian's activity range in N video frames, and since this update principle is not limited by the distance between the pedestrian and the camera device, this application can reduce the impact of the distance between the pedestrian and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the pedestrian is far from the camera device, the update of the skeletal key points relative to the pedestrian's activity range in N video frames in this application can avoid the problem of the updated value of the skeletal key points being too small, thus reducing the impact of the distance between the pedestrian and the camera device on the accuracy of fall detection.
[0078] Since this update principle is not limited by the distance between pedestrians and the camera device, the embodiments of this application can reduce the requirements for training data and detection data. For example, related technologies usually require that the training data and detection data correspond to the same application scenario, while the embodiments of this application do not require that the training data and detection data correspond to the same application scenario; thus, the embodiments of this application can reduce the difficulty of obtaining training data, thereby increasing the richness of the training data.
[0079] Method Example 2
[0080] refer to Figure 3 The diagram illustrates a step-by-step flowchart of a fall detection method according to an embodiment of this application. The method may specifically include the following steps:
[0081] Step 301: Obtain N video frames to be detected; N can be a positive integer;
[0082] Step 302: Determine the first position information of the skeletal key points in N video frames respectively;
[0083] Step 303: Fuse the first position information of the skeletal key points in N video frames to obtain the position range information of the pedestrian in N video frames;
[0084] Step 304: Based on the location range information, update the first location information of the skeletal key points in the N video frames to obtain the updated second location information;
[0085] Step 305: Determine the center of gravity position of the pedestrian in the i-th video frame based on the second position information of the preset skeletal key points in the i-th video frame; i is a positive integer;
[0086] Step 306: Based on the center of gravity position of the pedestrian in the i-th video frame, update the second position information of the skeletal key points in the i-th video frame to obtain the third position information of the skeletal key points in the i-th video frame.
[0087] Step 307: Detect whether a pedestrian has fallen based on the third position information of the skeletal key points in N video frames.
[0088] In this embodiment of the application, when updating the first position information of the skeletal key points in the N video frames to the second position information, the second position information of the skeletal key points in the i-th video frame is also updated according to the center of gravity position of the pedestrian in the i-th video frame. Wherein, 1≤i≤N.
[0089] While key points on the human skeleton are individual location information, multiple key points on a person's skeleton should be closely related. Related technologies primarily utilize the location information of key points on the skeleton, without considering or using the relationships between them.
[0090] In this embodiment, updating the skeletal keypoints of a video frame relative to the center of gravity results in more closely related skeletal keypoints. Thus, during fall detection, the coordinate changes of multiple skeletal keypoints are treated as a whole, rather than as scattered independent points. This allows the fall detection model to capture the intrinsic connection information of the skeletal keypoints, thereby improving the accuracy of fall detection.
[0091] This application embodiment can determine the center of gravity position of a pedestrian in the i-th video frame based on the second position information of preset skeletal key points in the i-th video frame. The preset skeletal key points can be determined by those skilled in the art according to actual application requirements. For example, the preset skeletal key points may include four skeletal key points: right shoulder-5, left shoulder-6, right waist-11, and left waist-12. This application embodiment can average the second position information of the four skeletal key points to obtain the center of gravity position. Specifically, the horizontal coordinate of the center of gravity position can be the average of the horizontal coordinates of the second position information of the four skeletal key points, and the vertical coordinate of the center of gravity position can be the average of the vertical coordinates of the second position information of the four skeletal key points.
[0092] The process of updating the second position information of the skeletal keypoints in the i-th video frame can specifically include: determining the third position information of the skeletal keypoints in the i-th video frame based on the third difference between the second position information of the skeletal keypoints and the pedestrian's center of gravity position. The horizontal coordinate of the third position information can be the difference between the horizontal coordinate of the second position information and the horizontal coordinate of the center of gravity position, and the vertical coordinate of the third position information can be the difference between the vertical coordinate of the second position information and the vertical coordinate of the center of gravity position.
[0093] In this embodiment, the third position information of skeletal key points in N video frames can be input into the fall detection model, and the fall detection model can output the corresponding detection results.
[0094] In summary, the fall detection method of this application updates the first position information of skeletal key points in N video frames based on the position range information. Since the update principle of the skeletal key points in this application is based on updating the skeletal key points relative to the pedestrian's activity range in N video frames, and this update principle is not limited by the distance between the pedestrian and the camera device, this application can reduce the impact of the pedestrian's distance from the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the pedestrian is far from the camera device, the update of the skeletal key points relative to the pedestrian's activity range in N video frames in this application can avoid the problem of the updated values of the skeletal key points being too small, thus reducing the impact of the distance between the pedestrian and the camera device on the accuracy of fall detection.
[0095] Furthermore, by updating the skeletal keypoints of a video frame relative to the center of gravity in this embodiment, more closely related skeletal keypoints can be obtained. In this way, during the fall detection process, the coordinate changes of multiple skeletal keypoints will be regarded as a whole, rather than scattered independent points. This allows the fall detection model to capture the intrinsic connection information of the skeletal keypoints, thereby improving the accuracy of fall detection.
[0096] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.
[0097] Based on the above embodiments, this embodiment also provides a fall detection device, referring to... Figure 4 The device may specifically include: a video frame acquisition module 401, a key point determination module 402, a fusion module 403, a first update module 404, and a first detection module 405.
[0098] Among them, the video frame acquisition module 401 is used to acquire N video frames to be detected; N is a positive integer;
[0099] The key point determination module 402 is used to determine the first position information of the skeletal key points in N video frames respectively;
[0100] The fusion module 403 is used to fuse the first position information of the skeletal key points in N video frames to obtain the position range information of the pedestrian in N video frames.
[0101] The first update module 404 is used to update the first position information of the skeletal key points in the N video frames according to the position range information, so as to obtain the updated second position information.
[0102] The first detection module 405 is used to detect whether a pedestrian has fallen based on the second position information of the skeletal key points in N video frames.
[0103] Optionally, the fusion module 403 may specifically include:
[0104] The upper and lower limit determination module is used to determine the lower and upper limits of the horizontal coordinate, as well as the lower and upper limits of the vertical coordinate, based on the first position information of the skeletal key points in N video frames.
[0105] The coordinate range determination module is used to determine the range information of the horizontal coordinate based on the lower limit and the upper limit of the horizontal coordinate, and to determine the range information of the vertical coordinate based on the lower limit and the upper limit of the vertical coordinate.
[0106] Optionally, the first update module 404 may specifically include:
[0107] The horizontal coordinate update module is used to determine the horizontal coordinate of the second position information based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information.
[0108] The horizontal coordinate update module is used to determine the vertical coordinate of the second position information based on the second difference between the vertical coordinate of the first position information and the lower limit of the vertical coordinate, as well as the vertical coordinate range information.
[0109] Optionally, the device may further include:
[0110] The center of gravity determination module is used to determine the center of gravity position of the pedestrian in the i-th video frame based on the second position information of the preset skeletal key points in the i-th video frame; i is a positive integer;
[0111] The second update module is used to update the second position information of the skeletal key points in the i-th video frame based on the center of gravity position of the pedestrian in the i-th video frame, so as to obtain the third position information of the skeletal key points in the i-th video frame.
[0112] The second detection module is used to detect whether a pedestrian has fallen based on the third position information of the skeletal key points in N video frames.
[0113] Optionally, the second update module is specifically used to determine the third position information of the skeletal key points in the i-th video frame based on the third difference between the second position information of the skeletal key points in the i-th video frame and the position of the pedestrian's center of gravity.
[0114] In summary, the fall detection device of this application updates the first position information of skeletal key points in N video frames based on the position range information. Since the update principle of the skeletal key points in this application is based on the update of the skeletal key points relative to the pedestrian's activity range in N video frames, and since this update principle is not limited by the distance between the pedestrian and the camera device, this application can reduce the impact of the distance between the pedestrian and the camera device on the accuracy of fall detection, thereby improving the robustness of fall detection. For example, when the pedestrian is far from the camera device, the update of the skeletal key points relative to the pedestrian's activity range in N video frames in this application can avoid the problem of the updated value of the skeletal key points being too small, thus reducing the impact of the distance between the pedestrian and the camera device on the accuracy of fall detection.
[0115] Since this update principle is not limited by the distance between pedestrians and the camera device, the embodiments of this application can reduce the requirements for training data and detection data. For example, related technologies usually require that the training data and detection data correspond to the same application scenario, while the embodiments of this application do not require that the training data and detection data correspond to the same application scenario; thus, the embodiments of this application can reduce the difficulty of obtaining training data, thereby increasing the richness of the training data.
[0116] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.
[0117] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In this application, the electronic device includes various types of devices such as terminal devices and servers (clusters).
[0118] The embodiments of this disclosure can be implemented as an apparatus configured as desired using any suitable hardware, firmware, software, or any combination thereof, including electronic devices such as terminal devices and servers (clusters). Figure 5 An exemplary apparatus 1100 is schematically shown that can be used to implement the various embodiments described in this application.
[0119] In one embodiment, Figure 5An exemplary device 1100 is shown, which includes one or more processors 1102, a control module (chipset) 1104 coupled to at least one of the processors 1102, a memory 1106 coupled to the control module 1104, a non-volatile memory (NVM) / storage device 1108 coupled to the control module 1104, one or more input / output devices 1110 coupled to the control module 1104, and a network interface 1112 coupled to the control module 1104.
[0120] Processor 1102 may include one or more single-core or multi-core processors, and processor 1102 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1100 can serve as a terminal device, server (cluster), or other device as described in the embodiments of this application.
[0121] In some embodiments, apparatus 1100 may include one or more computer-readable media (e.g., memory 1106 or NVM / storage device 1108) having instructions 1114 and one or more processors 1102 that are combined with the one or more computer-readable media and configured to execute instructions 1114 to implement a module thereby performing the actions described in this disclosure.
[0122] In one embodiment, the control module 1104 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1102 and / or any suitable device or component communicating with the control module 1104.
[0123] The control module 1104 may include a memory controller module to provide an interface to the memory 1106. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0124] Memory 1106 may be used, for example, to load and store data and / or instructions 1114 for device 1100. In one embodiment, memory 1106 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1106 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0125] In one embodiment, the control module 1104 may include one or more input / output controllers to provide interfaces to the NVM / storage device 1108 and (one or more) input / output devices 1110.
[0126] For example, NVM / storage device 1108 may be used to store data and / or instructions 1114. NVM / storage device 1108 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).
[0127] NVM / storage device 1108 may include storage resources that are physically part of a device on which device 1100 is mounted, or that can be accessed by the device without needing to be part of the device. For example, NVM / storage device 1108 may be accessed via a network via one or more input / output devices 1110.
[0128] One or more input / output devices 1110 may provide an interface for device 1100 to communicate with any other suitable device. Input / output devices 1110 may include communication components, audio components, sensor components, etc. Network interface 1112 may provide an interface for device 1100 to communicate via one or more networks. Device 1100 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.
[0129] In one embodiment, at least one of the processors 1102 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1104. In one embodiment, at least one of the processors 1102 may be logically packaged with one or more controllers of the control module 1104 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1102 may be integrated with the logic of one or more controllers of the control module 1104 on the same die. In one embodiment, at least one of the processors 1102 may be integrated with the logic of one or more controllers of the control module 1104 on the same die to form a system-on-a-chip (SoC).
[0130] In various embodiments, device 1100 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, device 1100 may have more or fewer components and / or different architectures. For example, in some embodiments, device 1100 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0131] The detection device can use a main control chip as a processor or control module, and sensor data, position information, etc. can be stored in a memory or NVM / storage device. The sensor group can be used as an input / output device, and the communication interface can include a network interface.
[0132] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0133] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0134] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0135] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.
[0136] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.
[0137] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0138] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0139] The foregoing has provided a detailed description of a fall detection method and apparatus, an electronic device, and a machine-readable medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A fall detection method, characterized in that, The method includes: Obtain N video frames to be detected; N is a positive integer; Determine the first position information of the skeletal keypoints in N video frames respectively; The first position information of the skeletal key points in N video frames is fused to obtain the position range information of the pedestrian in N video frames; Based on the location range information, the first location information of the skeletal key points in the N video frames is updated to obtain the updated second location information; Based on the second position information of skeletal key points in N video frames, detect whether a pedestrian has fallen. The process of fusing the first position information of skeletal key points in N video frames includes: Based on the first position information of the skeletal key points in N video frames, determine the lower limit and upper limit of the horizontal coordinate, as well as the lower limit and upper limit of the vertical coordinate. The range of the horizontal axis is determined based on the lower and upper limits of the horizontal axis, and the range of the vertical axis is determined based on the lower and upper limits of the vertical axis. The method further includes: Based on the second position information of the preset skeletal key points in the i-th video frame, determine the center of gravity position of the pedestrian in the i-th video frame; i is a positive integer; Based on the pedestrian's center of gravity position in the i-th video frame, the second position information of the skeletal key points in the i-th video frame is updated to obtain the third position information of the skeletal key points in the i-th video frame. Based on the third position information of skeletal key points in N video frames, detect whether a pedestrian has fallen.
2. The method according to claim 1, characterized in that, The step of updating the first position information of the skeletal key points in the N video frames includes: The horizontal coordinate of the second position information is determined based on the first difference between the horizontal coordinate of the first position information and the lower limit of the horizontal coordinate, as well as the horizontal coordinate range information. The ordinate of the second position information is determined based on the second difference between the ordinate of the first position information and the lower limit of the ordinate, as well as the ordinate range information.
3. The method according to claim 1, characterized in that, The update of the second position information of the skeletal keypoints in the i-th video frame includes: The third position information of the skeletal keypoints in the i-th video frame is determined based on the third difference between the second position information of the skeletal keypoints in the i-th video frame and the position of the pedestrian's center of gravity.
4. A fall detection device, characterized in that, The device includes: The video frame acquisition module is used to acquire N video frames to be detected; N is a positive integer. The key point determination module is used to determine the first position information of the skeletal key points in N video frames respectively; The fusion module is used to fuse the first position information of the skeletal key points in N video frames to obtain the position range information of the pedestrian in N video frames. The first update module is used to update the first position information of the skeletal key points in the N video frames according to the position range information, so as to obtain the updated second position information. The first detection module is used to detect whether a pedestrian has fallen based on the second position information of the skeletal key points in N video frames; The fusion module includes: The upper and lower limit determination module is used to determine the lower and upper limits of the horizontal coordinate, as well as the lower and upper limits of the vertical coordinate, based on the first position information of the skeletal key points in N video frames. The coordinate range determination module is used to determine the range information of the horizontal coordinate based on the lower limit and the upper limit of the horizontal coordinate, and to determine the range information of the vertical coordinate based on the lower limit and the upper limit of the vertical coordinate. The device further includes: The center of gravity determination module is used to determine the center of gravity position of the pedestrian in the i-th video frame based on the second position information of the preset skeletal key points in the i-th video frame; i is a positive integer; The second update module is used to update the second position information of the skeletal key points in the i-th video frame based on the center of gravity position of the pedestrian in the i-th video frame, so as to obtain the third position information of the skeletal key points in the i-th video frame. The second detection module is used to detect whether a pedestrian has fallen based on the third position information of the skeletal key points in N video frames.
5. An electronic device, characterized in that, include: processor; and A memory having executable code stored thereon, which, when executed, causes the processor to perform the method as described in any one of claims 1-3.
6. A machine-readable medium having executable code stored thereon, which, when executed, causes a processor to perform the method as described in any one of claims 1-3.
Citation Information
Patent Citations
Fall detection method and system
CN111461042A
Image processing method and device, electronic equipment and computer storage medium
CN112418153A
Behavior detection method, device and system
CN112686075A