Personnel falling-to-ground detection method and device, electronic equipment and storage medium

Through the combination of depth camera and human posture estimation model, the collapse state is identified and screened, which solves the problem of increased cost of sensors and low manual monitoring efficiency, and achieves efficient and accurate collapse detection.

CN120260122APending Publication Date: 2025-07-04HITACHI BUILDING TECH GUANGZHOU CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325663.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, wearing sensors to detect that people fall down increases costs and affects work efficiency. However, manual video monitoring is inefficient and prone to errors and misses, and it is impossible to detect falls or falls in the production area in a timely manner.

Method used

The depth camera is used to collect video data, identify the key points of the human body and the coordinates of the detection frame through video frame extraction, determine the three-dimensional coordinates based on the depth value, filter out the fallen state, and perform double detection through the human body posture estimation model to improve accuracy.

Benefits of technology

No staff need to wear sensors, reduce costs, improve work efficiency, reduce misjudgment, and achieve efficient and accurate ground-fall detection for a long time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260122A_ABST
    Figure CN120260122A_ABST
Patent Text Reader

Abstract

The invention discloses a personnel falling-down detection method and device, electronic equipment and a storage medium, human body key points and human body detection frame coordinates are obtained through frame extraction and recognition of video data collected by a depth camera, three-dimensional coordinates of the human body key points are obtained in combination with depth values, and whether personnel fall down on the ground or not is judged through the three-dimensional coordinates. A person does not need to wear a sensor to monitor the falling of the person, the cost is reduced, the inconvenience of limb movement caused by wearing the sensor is avoided, the working efficiency of the worker is improved, the person does not need to watch a video to monitor the falling of the person, the detection efficiency of the falling of the person is improved, the falling of the person can be continuously detected for a long time, and mistakes and omissions are not likely to happen. The human body posture estimation is performed through the human body posture estimation model, then the target human body detection frame is obtained by screening the human body detection frames with the posture types of falling to the ground, and the falling judgment is performed on the human body in the target human body detection frame, so that the accuracy of human body falling detection is improved through double falling detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision processing, and in particular, to a method, device, electronic device and storage medium for detecting a person falling to the ground. Background Art

[0002] With the continuous advancement of the industrialization process, the scale of factories and production workshops has been continuously expanding, and the number of staff has also been increasing day by day. In production and manufacturing activities, the incidents of staff falling or tripping due to their own health problems or accidents occur frequently. If not discovered and rescued in time, serious safety accidents will be caused.

[0003] In the prior art, most of the methods for monitoring whether the staff in the production area fall to the ground are through two ways. One way is to configure speed sensors, acceleration sensors, locators, etc. at the key parts of the staff, and detect whether the staff fall to the ground through the sensors. The other way is to collect videos of the production area through cameras, and the staff in the monitoring room observe the videos manually to find that the staff fall to the ground.

[0004] The above-mentioned first method of wearing sensors by staff not only increases the cost, but also causes inconvenience to the staff's limb movements due to wearing sensors, affecting work efficiency. The second method of manually watching videos depends on manual experience, has low efficiency, and is prone to risks such as omissions and errors. Summary of the Invention

[0005] The present invention provides a method, device, electronic device and storage medium for detecting a person falling to the ground, so as to solve the problems that detecting a person falling to the ground by wearing sensors increases the cost and affects work efficiency, and the low efficiency of manually watching videos is prone to omissions and errors in discovering that a person falls to the ground.

[0006] In a first aspect, the present invention provides a method for detecting a person falling to the ground, including:

[0007] Extracting frames from the video data collected by the depth camera to obtain video frames, where the video frames include depth values;

[0008] Inputting the video frames into a pre-trained human pose estimation model to obtain the human key points, human detection box coordinates and pose types of each detected human body;

[0009] Determining the three-dimensional coordinates of the human key points according to the depth values;

[0010] Based on the human detection box coordinates, screening the human detection boxes with the pose type of falling to the ground to obtain target human detection boxes;

[0011] Determining whether the human body in the target human detection box is in a state of falling to the ground according to the three-dimensional coordinates of the human key points of the human body in the target human detection box;

[0012] If so, generate an alarm message for the person falling to the ground in the target human detection box.

[0013] In a second aspect, the present invention provides a person falling detection device, including:

[0014] A video frame extraction module, configured to extract frames from the video data collected by the depth camera to obtain a video frame, where the video frame includes depth values;

[0015] A human body detection module, configured to input the video frame into a pre-trained human body pose estimation model to obtain the human body key points, human body detection box coordinates, and pose types of each detected human body;

[0016] A key point three-dimensional coordinate determination module, configured to determine the three-dimensional coordinates of the human body key points according to the depth values;

[0017] A human body screening module, configured to screen the human body detection boxes with a falling pose based on the human body detection box coordinates to obtain a target human body detection box;

[0018] A falling state determination module, configured to determine whether the human body in the target human body detection box is in a falling state according to the three-dimensional coordinates of the human body key points in the target human body detection box; if so, execute the falling alarm module;

[0019] A falling alarm module, configured to generate an alarm message for the person falling to the ground in the target human body detection box.

[0020] In a third aspect, the present invention provides an electronic device, where the electronic device includes:

[0021] At least one processor; and

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the person falling detection method according to the first aspect of the present invention.

[0024] In a fourth aspect, the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the person falling detection method according to the first aspect of the present invention when executed by a processor.

[0025] After the present invention collects video data through a depth camera and identifies a person falling to the ground from the video data, on the one hand, it is not necessary for the staff in the production area to wear sensors, which not only reduces costs but also avoids the inconvenience of limb movements caused by wearing sensors, and can improve the work efficiency of the staff. On the other hand, by extracting frames from the video data to identify the video frame to obtain the human key points, the coordinates of the human detection box, and combining the depth value to obtain the three-dimensional coordinates of the human key points, it is determined whether a person has fallen to the ground through the three-dimensional coordinates. There is no need for manual viewing of the video to monitor whether a person has fallen to the ground, which improves the detection efficiency of a person falling to the ground, and can continuously detect whether a person has fallen to the ground for a long time, and it is not easy to make mistakes or omissions. On the other hand, first, the human pose estimation model is used to perform human pose estimation on the extracted video frame, and then the human detection box with the pose type of falling to the ground is further screened to obtain the target human detection box, and then the person in the target human detection box is judged whether to fall to the ground, which can avoid the human pose estimation model misjudging a person squatting on the ground as a person falling to the ground. Through double fall detection, the accuracy of human fall detection can be improved.

[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0028] Figure 1 is a flowchart of a method for detecting a person falling to the ground provided in Embodiment 1 of the present invention;

[0029] Figure 2 is a flowchart of a method for detecting a person falling to the ground provided in Embodiment 2 of the present invention;

[0030] Figure 3 is a flowchart of an example for detecting a person falling to the ground;

[0031] Figure 4 is a schematic structural diagram of a device for detecting a person falling to the ground provided in Embodiment 3 of the present invention;

[0032] Figure 5 is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0034] Embodiment 1

[0035] Figure 1 The figure is a flowchart of a method for detecting a person falling to the ground provided in Embodiment 1 of the present invention. This embodiment is applicable to detecting whether there is a person falling to the ground in a target area. This method can be executed by a person falling to the ground detection device, and the person falling to the ground detection device can be implemented in the form of hardware and / or software and can be configured in an electronic device. As Figure 1 shown, the method for detecting a person falling to the ground includes:

[0036] S101. Extract frames from the video data collected by the depth camera to obtain a video frame, and the video frame includes depth values.

[0037] This embodiment can be applied to the scenario of monitoring whether a person in a target area falls to the ground. Among them, the target area can be the area where the business activities of a market entity are located. Exemplarily, taking a manufacturing enterprise as an example, the target area can be areas such as the production workshop area, the office area, and the elevator.

[0038] In practical applications, a depth camera can be arranged in the target area. The depth camera can be a binocular camera or a structured light camera, so that the depth camera can collect video data of the target area. For example, the resolution of the depth camera can be set to 640×480 or 1280×720, and the frame rate can be set to 30 frames per second to collect video data of the target area. For the video data collected by the depth camera, frames can be extracted to obtain a video frame. Exemplarily, frames can be extracted from the video data according to a preset period, or key frames (I frames) in the video data can be extracted to obtain a video frame. And since the data is collected by the depth camera, the depth value of each video frame can be obtained at the same time. The depth value can be the depth value associated with each pixel point in the video frame. Among them, the method for the depth camera to obtain the depth value can refer to the prior art and will not be elaborated here.

[0039] S102. Input the video frame into a pre-trained human pose estimation model to obtain the human key points, human detection box coordinates, and pose types of each detected human body.

[0040] In this embodiment, the human pose estimation model can be a neural network model that, when inputting a video frame, detects the human body in the frame to obtain human key points, the coordinates of the human detection box, and the pose type of the human body. The human pose estimation model can be trained using images of human bodies in various poses. The training method can refer to the supervised training method of the neural network model, which will not be elaborated in this embodiment.

[0041] Among them, the number of human key points can be 15, 21, 28, etc. Exemplarily, the human key points can include the center of the head, both eyes, both ears, both shoulders, both elbows, both hands, both hips, both knees, and both feet. The coordinates of the human detection box can be the coordinates of the four corner points of the rectangular human detection box. The pose type can include but is not limited to standing, walking, falling to the ground, squatting, etc.

[0042] After inputting the video frame into the human pose estimation model, the human pose estimation model can output the identity ID, pose type, human key points, and the coordinates of the human detection box corresponding to each human detection box.

[0043] S103. Determine the three-dimensional coordinates of the human key points according to the depth value.

[0044] The human key points output by the human pose estimation model can include the x and y coordinates. The depth value associated with the x and y coordinates of the human key points and the x and y coordinates of each pixel in the video frame can be found, and the x and y coordinates of the human key points are fused with the depth value of the video frame to obtain the three-dimensional coordinates (x, y, z) of each human key point.

[0045] It should be noted that since the depth camera continuously collects video data, the video data collected by the depth camera is continuously framed to obtain a video frame and input into the human pose estimation model. That is, S101 - S103 are continuously performed, and the video data is continuously framed to obtain the human key points, the coordinates of the human detection box, the pose type, and the three-dimensional coordinates of the human key points of the detected human body.

[0046] S104. Screen the human detection boxes with the pose type of falling to the ground based on the coordinates of the human detection box to obtain the target human detection box.

[0047] In one embodiment, first, human detection frames with a falling posture type are screened out from all human detection frames. Then, the height and width of the human detection frame are calculated based on the coordinates of the four corner points of the human detection frame with a falling posture type, and the ratio of the calculated height to the width is used as the aspect ratio. The smaller the aspect ratio, the closer the height of the human body in the video frame is to the width or even less than the width, which better indicates that the human body is in a falling state. Human detection frames with a falling posture type and an aspect ratio less than a preset threshold can be determined as target human detection frames, so that human detection frames that are estimated to be falling by the human pose estimation model but are actually human squatting can be filtered out.

[0048] S105. Determine whether the human body in the target human detection frame is in a falling state according to the three-dimensional coordinates of the human key points of the human body in the target human detection frame.

[0049] In one embodiment, after the human body is first detected and the detection frame of the human body is determined to be a target human detection frame, a vector can be calculated through the three-dimensional coordinates of the human key points in the target human detection frame, and the vector is added to the vector record as the first vector. After continuing to extract frames, if the detection frame of the human body is still determined to be a target human detection frame, a vector is calculated through the three-dimensional coordinates of the human key points in the target human detection frame obtained from the current frame extraction, and the vector is added to the vector record as the second vector. It is determined that the human body is detected to be in a falling state in two consecutive frame extractions. In the vector record, the system time is used as the start time and the end time. Then, continue to extract frames. If the detection frame of the human body is still determined to be a target human detection frame after continuing to extract frames, a vector is calculated through the three-dimensional coordinates of the human key points in the target human detection frame obtained from the current frame extraction, the calculated vector is used to replace the second vector in the vector record, and the system time is used to replace the end time. Calculate the similarity between the first vector and the second vector, and calculate the time difference between the end time and the start time. If the similarity is greater than the similarity threshold and the time difference is greater than the time threshold, it is determined that the human body is detected to be in a falling state and the duration is greater than the threshold, and it can be determined that the person has fallen.

[0050] If the similarity is less than the similarity threshold, further, the mean value of the coordinates z (height) of the human key points corresponding to the first vector and the second vector in the vector record can be calculated, or the mean value of z in the three-dimensional coordinates of the human key points of the human body obtained from the frame extraction between the start time and the end time can be calculated, and whether the human body is in a falling state is further determined through the mean value.

[0051] In another embodiment, after the human detection frame of the human body is determined to be the target human detection frame, if the human detection frames of a human body in N consecutive video frames after frame extraction are all determined to be the target human detection frames, and the average value of z in the three-dimensional coordinates of the human key points in the human detection frames of the N video frames is less than the threshold, it can also be determined that a human body is determined to fall to the ground N consecutive times of frame extraction, and it can be determined that the human body is in a fallen state.

[0052] S106. Generate an alarm message for a human body falling to the ground in the target human detection frame.

[0053] When it is determined that the human body in the target human detection frame is in a fallen state, an alarm message for a person falling to the ground can be generated. Exemplarily, a video segment including the human body falling to the ground can be intercepted from the video data, and an alarm message including the video segment and the address of the target area where the person falls to the ground can be pushed to the safety production system or to the terminal device of the safety monitoring personnel.

[0054] After the present invention collects video data through a depth camera and identifies a person falling to the ground through the video data. On the one hand, it is not necessary for the staff in the production area to wear sensors, which not only reduces the cost but also avoids the inconvenience of limb movements caused by wearing sensors, and can improve the work efficiency of the staff. On the other hand, the video frames are extracted from the video data to identify the human key points, the coordinates of the human detection frame in the video frame, and the three-dimensional coordinates of the human key points are obtained by combining the depth value. Whether a person falls to the ground is judged through the three-dimensional coordinates, without manual viewing of the video to monitor whether a person falls to the ground, which improves the detection efficiency of a person falling to the ground, and can continuously detect whether a person falls to the ground for a long time, and it is not easy to make mistakes or omissions. On the other hand, first, the human pose estimation model is used to perform human pose estimation on the extracted video frames, and then the human detection frames with the pose type of falling to the ground are further screened to obtain the target human detection frames, and then the human body in the target human detection frame is judged for falling to the ground, which can avoid the human pose estimation model misjudging a squatting human body as a fallen human body, and the accuracy of human body falling detection can be improved through double falling detection.

[0055] Embodiment 2

[0056] Figure 2 The flowchart of a method for detecting a person falling to the ground provided by the second embodiment of the present invention is optimized on the basis of the first embodiment above, such as Figure 2 shown, the method for detecting a person falling to the ground includes:

[0057] S201. Extract frames from the video data collected by the depth camera to obtain a video frame, and the video frame includes a depth value.

[0058] In this embodiment, a depth camera can be set in the target area where personnel fall detection is required. After collecting video data of the target area through the depth camera, frames are extracted to obtain a video frame. Each pixel point in the video frame can be associated with a depth value. Exemplarily, in addition to the channels for recording the values of the three colors R, G, and B for each pixel, there can also be a channel data for representing the depth value.

[0059] Among them, the depth camera can be a camera using infrared structured light or time-of-flight (TOF) technology. The resolution of the depth camera can be 640×480 or 1280×720. Of course, it can also be other resolutions, and the frame rate can be 30 frames per second. Frames can be extracted from the video data at a fixed frame rate or a dynamic frame rate to obtain a video frame.

[0060] S202. Input the video frame into a pre-trained human pose estimation model to obtain the human key points, human detection box coordinates, and pose types of each detected human body.

[0061] The human pose estimation model in this embodiment can be a neural network model based on OpenPose, PoseNet, etc. After being trained by a supervised training method, the human pose estimation model can identify the human key points, human detection box coordinates, and pose types of the human body in the video frame.

[0062] Among them, the human key points at least include human key points such as the center of the head, both eyes, both ears, both shoulders, both elbows, both hands, both hips, both knees, and both feet. The pose types include but are not limited to postures such as standing, walking, falling, and squatting. The human detection box coordinates can be the coordinates of the four corner points of the rectangular detection box.

[0063] S203. Determine the three-dimensional coordinates of the human key points according to the depth value.

[0064] Specifically, the depth value of the video frame can be fused with the estimated human key points, and the three-dimensional coordinates of the human key points can be obtained by using a three-dimensional reconstruction algorithm, that is, finally, the identity ID, pose type, human detection box coordinates, and three-dimensional coordinates of the human key points of each person in the video frame are obtained.

[0065] S204. Calculate the aspect ratio of the human detection box with the pose type of falling based on the human detection box coordinates.

[0066] The aspect ratio can be the ratio of the height to the width of the human detection box. Exemplarily, the vertices of the human detection box are: the upper left vertex coordinate A(x1, y1), the upper right vertex coordinate B(x2, y1), the lower right vertex coordinate C(x2, y2), and the lower left vertex coordinate D(x1, y2). The preset aspect ratio is M(0 < M ≤ 1) = 1. Calculate the height of the human detection box as h = y2 - y1, the width as w = x2 - x1, and the aspect ratio as m = h / w.

[0067] S205. Determine the human detection boxes with aspect ratios less than the preset threshold in the human detection boxes with the posture type of falling to the ground as the target human detection boxes.

[0068] For a human body, the smaller the aspect ratio of its human detection box, the more it indicates that the human body is in a falling-to-the-ground state. A threshold M (such as set to 1) can be pre-configured. For the human detection boxes with the posture type of falling to the ground estimated by the human pose estimation model, if the aspect ratio m of the human detection box ≤ M, it indicates that the human detection box is the detection box of the human body falling to the ground, and this human detection box can be determined as the target human detection box, and S206 is executed to further confirm whether it is a human body falling to the ground. If the aspect ratio m of the human detection box > M, it indicates that although the posture type estimated by the human pose estimation model is falling to the ground, it is actually a human body in a squatting posture, and this part of the human detection boxes can be filtered out, so as to realize the preliminary screening of the human detection boxes, which can not only avoid the problem of inaccurate recognition of the video screen by the human pose estimation model and improve the accuracy of personnel falling to the ground, but also reduce the number of human detection boxes that need to be further confirmed whether it is a human body falling to the ground subsequently, and improve the efficiency of human body falling-to-the-ground detection.

[0069] S206. Calculate the current vector of the human body in the target human detection box by using the three-dimensional coordinates of the human body key points in the target human detection box.

[0070] In one embodiment, the three-dimensional coordinates of the human body key points can be directly used as the current vector of a human body. Exemplarily, the human body key points are sorted in a preset order (from the head to the feet), and the three-dimensional coordinates of the human body key points are spliced according to the sorting to obtain the current vector of the human body.

[0071] In another embodiment, a key point matrix can be constructed by using the three-dimensional coordinates of the human body key points in the target human detection box, the three-dimensional coordinates in the key point matrix are normalized to obtain the normalized key point matrix, and the normalized key point matrix is flattened into a one-dimensional array to obtain the current vector of the human body in the target human detection box.

[0072] Exemplarily, for the human body key points corresponding to a human body i, the constructed key point matrix P i is as follows:

[0073]

[0074] x i1 、y i1 、z i1 represent the x, y, and z coordinates of the first key point of the i-th human body. Further, the three-dimensional coordinates of the human body key points are normalized through the following formula to obtain the normalized key point matrix P norm,i :

[0075]

[0076] where min(P i ) represents the minimum coordinate value among the human body key points of the i-th human body (which can be the minimum value among the x, y, and z coordinate values), and max(P i ) represents the maximum coordinate value among the human body key points of the i-th human body (which can be the maximum value among the x, y, and z coordinate values). Finally, the flattened key point matrix P norm,i is flattened to obtain the current vector a of the human body in each target human body detection box i =[flat(P norm,i )]. It should be noted that in one example, flattening can be the splicing of the coordinates of each normalized human body key point.

[0077] S207. Add the current vector to the vector record, and delete the vectors and times of the human bodies in the human body detection boxes outside the target human body detection box in the vector record. The vector record is used to record the vectors and times of the human bodies in the target human body detection boxes obtained after each frame extraction.

[0078] The vector record is used to record the vectors and times of the human bodies in the target human body detection boxes obtained after human body pose estimation on the video frames obtained by each frame extraction. The time includes the start time and the end time. The start time is the time when the human body detection box of the human body is determined as the target human body detection box for the second time after being determined as the target human body detection box for the first time, and the end time can be the time when it is determined as the target human body detection box again after being determined as the target human body detection box for the second time.

[0079] In this embodiment, after each frame extraction, when executing S201 - S206 to obtain the current vector of the human body in each target human body detection box, it can be determined whether there are a first vector and a second vector of the human body in the target human body detection box in the pre-configured vector record. The first vector is added to the vector record before the second vector.

[0080] If there are no first and second vectors of the person in the target human detection box, add the current vector as the first vector to the vector record. Specifically, if there are no first and second vectors of the person in the target human detection box in the vector record, it indicates that the person is first detected as being in the fallen posture type. The current vector can be added to the vector record as the first vector a.

[0081] If the first vector of the person in the target human detection box exists but the second vector does not, add the current vector as the second vector to the vector record, and add the current time as the start time and end time to the vector record. Specifically, for the person in the target human detection box, if the first vector a exists in the vector record but the second vector does not, it indicates that the person was recognized as fallen after the previous frame extraction, and the first vector a obtained after the previous frame extraction was added to the vector record. After the current frame extraction, the person is also recognized as fallen, that is, the person is recognized as fallen in two consecutive frame extractions. The current vector can be added to the vector record as the second vector b, and the current time is added to the vector record as the start time and end time.

[0082] If the first and second vectors of the person in the target human detection box exist, replace the second vector with the current vector in the vector record, and replace the end time with the current time. Specifically, for the person in the target human detection box, if the first vector a and the second vector b exist in the vector record, it indicates that the person is recognized as fallen in consecutive frame extractions and is also recognized as fallen after the current frame extraction. The second vector b in the vector record can be replaced with the current vector, and the end time can be sampled with the current time.

[0083] It should be noted that if the human detection box of the person is not determined as the target human detection box after the current frame extraction, that is, the posture type is not fallen or the posture type is fallen but is not determined as the target human detection box after calculating the aspect ratio filter, it indicates that the person is not in the fallen state. The vectors and times of the people in the human detection boxes other than the target human detection box can be deleted from the vector record, so that only the vectors and times of the initially fallen people or the people recognized as fallen in at least two consecutive frame extractions are recorded in the vector record.

[0084] The following uses an example to illustrate the process of updating the vector record:

[0085] Suppose a person A is detected in the first frame extraction, and its human detection box is confirmed as the target human detection box and vector a is obtained, generating a vector record: person A, vector a;

[0086] Assume that human body A is also detected in the second frame extraction. If its human detection box is not confirmed as the target human detection box, it indicates that after human body A falls to the ground, human body A stands up again. Delete human body A and its vector a from the vector record; if its human detection box is confirmed as the target human detection box, it indicates that human body A is also detected falling to the ground in the second frame extraction and vector b is obtained. Then the vector record is updated to: human body A, vector a, vector b, start time T1 (current time), end time T2 (current time).

[0087] Assume that human body A is also detected in the third frame extraction. If its human detection box is not confirmed as the target human detection box, it indicates that after human body A falls to the ground, human body A stands up again. Delete human body A, its vector a, vector b, start time, and end time from the vector record; if its human detection box is confirmed as the target human detection box, it indicates that human body A is also detected falling to the ground in the third frame extraction and the current vector is obtained. Then the vector record is updated to: human body A, vector a, vector b (replaced by the current vector), start time T1, end time T2 (current time).

[0088] And so on, that is, each human body in the vector record has at most vector a and vector b at most, and vector b and end time T2 are updated. When the human detection box of a human body is not determined as the target human detection box, it indicates that the human body stands up again after falling to the ground, and the vector and time of the human body are deleted from the vector record.

[0089] S208. Determine whether the human body in the target human detection box is in a fallen state based on the vector and time of the human body in the vector record.

[0090] Specifically, after each frame extraction updates the vector record, the first to-be-confirmed human body with the first vector and the second vector can be found in the vector record. For each first to-be-confirmed human body, calculate the similarity between the first vector and the second vector of the first to-be-confirmed human body, and determine the second to-be-confirmed human body and the third to-be-confirmed human body from the first to-be-confirmed human bodies. The second to-be-confirmed human body is the first to-be-confirmed human body with a similarity greater than or equal to the preset similarity threshold, and the third to-be-confirmed human body is the first to-be-confirmed human body with a similarity less than the similarity threshold. For each second to-be-confirmed human body, calculate the time difference between the end time and the start time of the vector of the second to-be-confirmed human body in the vector record. The start time is the system time when the second vector is added to the vector record, and the end time is the system time when the second vector is updated. When the time difference is greater than or equal to the first time threshold, determine that the second to-be-confirmed human body is a fallen human body without limb movement, and delete the vector and time of the fallen human body from the vector record. Among them, the time difference between the end time and the start time represents the duration of the human body falling to the ground.

[0091] Among them, the similarity between the first vector and the second vector can be calculated by the following formula:

[0092]

[0093] a 1i represents the first vector of the i-th human body, b 1i represents the second vector of the i-th human body, and n is the vector length of the first vector and the second vector.

[0094] Taking the following vector record Φ after one frame extraction as an example:

[0095] Φ = [(id1, a1, b1, T1, T2), (id2, a2, b2, T1, T2), (id3, a3), …, (idn, an, bn, T1, T2)].

[0096] First, search for the human body with the first vector a and the second vector b in the vector record Φ as the first human body to be confirmed. In the above example, id1, id2, and idn are the first human bodies to be confirmed. After calculating the similarity cosθ of the first human body to be confirmed, the human body with the similarity cosθ greater than or equal to the similarity threshold (0.95) in the first human body to be confirmed is determined as the second human body to be confirmed R1. Further calculate the time difference T = T2 - T1 between the end time T2 and the start time T1 of each second human body to be confirmed R1. If the time difference T is greater than the first time threshold (such as 10 seconds), it indicates that the second human body to be confirmed R1 has fallen to the ground and the duration is greater than the first time threshold. It can be determined that the second human body to be confirmed has fallen to the ground and has no limb movement (the similarity cosθ greater than or equal to the similarity threshold means that the posture change is small or unchanged). The vector and time of this human body can be deleted from the vector record, and S209 is executed.

[0097] For each third human body to be confirmed with the similarity cosθ less than the similarity threshold (0.95), use the three-dimensional coordinates of the human body key points corresponding to the first vector and the second vector of the third human body to be confirmed to determine whether the third human body to be confirmed has fallen to the ground.

[0098] Specifically, the average human height can be calculated using the three-dimensional coordinates of the corresponding human key points of the first vector and the second vector of the third human to be confirmed. For example, the three-dimensional coordinates of the three-dimensional key points obtained from the video frame of the first vector, and the average height Zav can be calculated using the height coordinates z in the three-dimensional coordinates of the three-dimensional key points obtained from the video frame of the second vector. Of course, the average height Zav can also be calculated using the three-dimensional coordinates of the three-dimensional key points obtained from multiple video frames between the video frame of the first vector and the video frame of the second vector. Among the third humans to be confirmed with a similarity cosθ less than the similarity threshold (0.95), the fourth human to be confirmed with an average human height Zav less than the first height threshold (such as h = 50 cm) is determined, and the fifth human to be confirmed with an average human height Zav greater than or equal to the first height threshold is determined. For each fourth human to be confirmed with an average human height Zav less than the first height threshold (such as h = 50 cm), calculate the time difference T between the end time T2 and the start time T1 of the vector of the fourth human to be confirmed in the vector record. When the time difference T is greater than or equal to the second time threshold (such as 15 seconds), it is determined that the fourth human to be confirmed is a fallen human with limb movements (a similarity cosθ less than the similarity threshold (0.95) indicates a difference in posture and the presence of limb movements). Delete the vector and time of the fallen human in the vector record and execute S209.

[0099] For each fifth human to be confirmed with an average human height Zav greater than or equal to the first height threshold, determine the three-dimensional coordinates (referring to the height z coordinate) of the two hip bone key points of the fifth human to be confirmed. When the three-dimensional coordinate of any one hip bone key point is less than the second height threshold (such as 10 cm), of course, it can also be that the three-dimensional coordinate of any one hip bone key point is within a preset range (such as 5 cm > z > 10 cm). Calculate the time difference T between the end time T2 and the start time T1 of the vector of the fifth human to be confirmed in the vector record. When the time difference T is greater than or equal to the third time threshold (such as 10 seconds), it is determined that the fifth human to be confirmed is a fallen human and is in a sitting state with limb activities after falling. Delete the vector and time of the fallen human in the vector record and execute S209.

[0100] In this embodiment, the state of the human body with the posture type of falling is recorded through the first vector, the second vector, the start time, and the end time of the human body in the vector record, and the similarity between the first vector and the second vector is calculated. For the human body with a similarity greater than the similarity threshold, the duration of falling is calculated through the start time and the end time to determine that the human body has fallen. For the human body with a similarity less than the similarity threshold, it is determined whether the human body has fallen by calculating the average height and the height of the hip bone key points in combination with the duration, which can more accurately and multi-levelly confirm the falling of the human body, improve the accuracy of human body falling detection, and can identify falling states such as whether there are limb movements.

[0101] S209. Generate an alarm message for a person falling to the ground in the target human detection box.

[0102] When it is determined that the human in the target human detection box is in a fallen state, an alarm message for a person falling to the ground can be generated. Exemplarily, a video segment including the human falling to the ground can be intercepted from the video data, and an alarm message including the video segment, the state of the person falling to the ground (whether there are pose actions), and the address of the target area can be pushed to the safety production system or to the terminal device of the safety monitoring personnel.

[0103] Figure 3 The following is a flowchart of an example of the person falling to the ground detection of the present invention. As Figure 3 shown, after the depth camera acquires video data of the target area and extracts frames to obtain a video frame, on the one hand, depth data is obtained through the video frame, and on the other hand, human key points, detection box coordinates, pose types, and the identity ID of the human are identified from the video frame through a deep neural network. After fusing the depth data, the three-dimensional coordinates of the human key points, detection box coordinates, pose types, and the identity ID of the human are obtained. Further, the aspect ratio is calculated through the detection box coordinates, and the human detection box with an aspect ratio less than the threshold is retained. A vector is calculated through the three-dimensional coordinates of the human key points in the retained human detection box, and the identity ID of the human and vector A are recorded. Frames are continuously extracted for detection and the identity ID and vector are recorded. After each frame extraction ends, the human including vector A and vector B is searched for from the records, and the cosine value COSθ of vector A and vector B with the same identity ID is calculated. If 0.95 ≤ COSθ ≤ 1.0 and the time difference between the calculated recorded vector B and vector A is greater than the threshold, it is determined that the person has fallen to the ground, and the human falling to the ground information is pushed. If COSθ < 0.95, the average value of the height coordinate z in the three-dimensional coordinates of the human key points is calculated. If the average value of the height coordinate z is less than the height threshold and the time difference between the recorded vector B and vector A is greater than the threshold, it is determined that the person has fallen to the ground, and the human falling to the ground information is pushed. If the average value of the height coordinate z is greater than or equal to the height threshold, it is determined whether the height coordinate z in the three-dimensional coordinates of any one of the hip key points in the human key points is less than the height threshold. If the average value of the height coordinate z is less than the height threshold and the time difference between the recorded vector B and vector A is greater than the threshold, it is determined that the person has fallen to the ground, and the human falling to the ground information is pushed. If not, it is determined whether the depth camera has stopped acquiring video frames. If not, video frames are continuously extracted to obtain video frames. If so, the human falling to the ground detection is ended.

[0104] In this embodiment, through the recognition of the video frames extracted from the video data collected by the depth camera, the three-dimensional coordinates of the human body key points, the coordinates of the human body detection frame, and the posture type are obtained by combining the depth data of the video frame. The target human body detection frame is obtained by filtering through the aspect ratio of the height and width of the human body detection frame coordinates. The current vector of the human body in the target human body detection frame is calculated using the three-dimensional coordinates of the human body key points in the target human body detection frame, and the current vector is added to the vector record. In addition, the vectors and time of the human bodies in the human body detection frames outside the target human body detection frame are deleted from the vector record. Whether the human body in the target human body detection frame is in a fallen state is determined based on the vectors and time of the human body in the vector record. If so, an alarm message indicating that the human body in the target human body detection frame has fallen is generated. After the human body posture estimation model predicts that the human body has fallen, it is further continuously determined whether the human body has fallen by vectorizing the key point three-dimensional coordinates. There is no need for the human body to wear a sensor to detect whether the human body has fallen, nor is it necessary for an operator to watch the video surveillance to check whether the human body has fallen. This reduces costs and improves the detection efficiency of human body falls. Moreover, it can continuously detect human body falls for a long time and is not prone to errors or omissions. In addition, first, the human body posture estimation model is used to perform human body posture estimation on the extracted video frames, and then the human body detection frames with the posture type of fallen are further screened to obtain the target human body detection frame. Then, whether the human body in the target human body detection frame has fallen is judged based on the human body vector, which can avoid the human body posture estimation model misjudging a squatting human body as a fallen human body. The accuracy of human body fall detection can be improved through double fall detection.

[0105] Embodiment III

[0106] Figure 4 FIG. is a schematic structural diagram of a personnel fall detection device provided in Embodiment III of the present invention. As Figure 4 shown, the personnel fall detection device includes:

[0107] A video frame extraction module 401, configured to extract video frames from the video data collected by the depth camera to obtain a video frame, where the video frame includes depth values;

[0108] A human body detection module 402, configured to input the video frame into a pre-trained human body posture estimation model to obtain the human body key points, the human body detection frame coordinates, and the posture type of each detected human body;

[0109] A key point three-dimensional coordinate determination module 403, configured to determine the three-dimensional coordinates of the human body key points according to the depth values;

[0110] A human body screening module 404, configured to screen the human body detection frames with the posture type of fallen based on the human body detection frame coordinates to obtain a target human body detection frame;

[0111] The falling state determination module 405 is configured to determine whether the human body in the target human body detection frame is in a falling state according to the three-dimensional coordinates of the human body key points of the human body in the target human body detection frame; if so, execute the falling alarm module 406;

[0112] The falling alarm module 406 is configured to generate an alarm message for the human body falling in the target human body detection frame.

[0113] Optionally, the human body screening module 404 includes:

[0114] The aspect ratio calculation unit is configured to calculate the aspect ratio of the human body detection frame with the falling posture type based on the coordinates of the human body detection frame;

[0115] The human body detection frame screening unit is configured to determine the human body detection frame with an aspect ratio less than a preset threshold in the human body detection frames with the falling posture type as the target human body detection frame.

[0116] Optionally, the falling state determination module 405 includes:

[0117] The vector calculation unit is configured to calculate the current vector of the human body in the target human body detection frame by using the three-dimensional coordinates of the human body key points of the human body in the target human body detection frame;

[0118] The record adding and deleting unit is configured to add the current vector to the vector record, and delete the vectors and times of the human bodies in the human body detection frames other than the target human body detection frame in the vector record, where the vector record is used to record the vectors and times of the human bodies in the target human body detection frame obtained after each frame extraction;

[0119] The falling determination unit is configured to determine whether the human body in the target human body detection frame is in a falling state based on the vectors and times of the human bodies in the vector record.

[0120] Optionally, the vector calculation unit includes:

[0121] The key point matrix construction subunit is configured to construct a key point matrix by using the three-dimensional coordinates of the human body key points of the human body in the target human body detection frame;

[0122] The normalization subunit is configured to perform normalization processing on the three-dimensional coordinates in the key point matrix to obtain a normalized key point matrix;

[0123] The matrix flattening subunit is configured to flatten the normalized key point matrix into a one-dimensional array to obtain the current vector of the human body in the target human body detection frame.

[0124] Optionally, the record adding and deleting unit includes:

[0125] A record judgment subunit, configured to judge whether there are a first vector and a second vector of the human body in the target human detection box in a pre-configured vector record, where the first vector is added to the vector record before the second vector;

[0126] A first addition subunit, configured to, if there are no first vector and second vector of the human body in the target human detection box, add the current vector as the first vector to the vector record;

[0127] A second addition subunit, configured to, if there is a first vector of the human body in the target human detection box but no second vector, add the current vector as the second vector to the vector record, and add the current time as the start time and end time to the vector record;

[0128] A third addition subunit, configured to, if there are a first vector and a second vector of the human body in the target human detection box, replace the second vector with the current vector in the vector record, and replace the end time with the current time.

[0129] Optionally, the falling determination unit includes:

[0130] A search subunit, configured to search for a first human body to be confirmed with a first vector and a second vector in the vector record;

[0131] A vector similarity calculation subunit, configured to calculate the similarity between the first vector and the second vector of each first human body to be confirmed;

[0132] A human body division subunit, configured to determine a second human body to be confirmed and a third human body to be confirmed from the first human bodies to be confirmed, where the second human body to be confirmed is a first human body to be confirmed with a similarity greater than or equal to a preset similarity threshold, and the third human body to be confirmed is a first human body to be confirmed with a similarity less than the similarity threshold;

[0133] A first time difference calculation subunit, configured to calculate the time difference between the end time and the start time of the vector of each second human body to be confirmed in the vector record, where the start time is the system time when the second vector is added to the vector record, and the end time is the system time when the second vector is updated;

[0134] A first falling determination subunit, configured to, when the time difference is greater than or equal to a first time threshold, determine that the second human body to be confirmed is a fallen human body without limb movement, and delete the vector and time of the fallen human body in the vector record;

[0135] The second falling determination subunit is configured to determine, for each third human body to be confirmed, whether the third human body to be confirmed has fallen by using the three-dimensional coordinates of the corresponding human key points of the first vector and the second vector of the third human body to be confirmed.

[0136] Optionally, the second falling determination subunit is specifically configured to:

[0137] Calculate the average human height by using the three-dimensional coordinates of the corresponding human key points of the first vector and the second vector of the third human body to be confirmed;

[0138] In the third human body to be confirmed, determine the fourth human body to be confirmed with an average human height less than the first height threshold, and determine the fifth human body to be confirmed with an average human height greater than or equal to the first height threshold;

[0139] For each fourth human body to be confirmed, calculate the time difference between the end time and the start time of the vector of the fourth human body to be confirmed in the vector record;

[0140] When the time difference is greater than or equal to the second time threshold, determine that the fourth human body to be confirmed is a fallen human body with limb movements, and delete the vector and time of the fallen human body in the vector record;

[0141] For each fifth human body to be confirmed, determine the three-dimensional coordinates of the two hip key points of the fifth human body to be confirmed;

[0142] When the three-dimensional coordinates of any one hip key point are less than the second height threshold, calculate the time difference between the end time and the start time of the vector of the fifth human body to be confirmed in the vector record;

[0143] When the time difference is greater than or equal to the third time threshold, determine that the fifth human body to be confirmed is a fallen human body and is in a sitting state with limb activities after falling, and delete the vector and time of the fallen human body in the vector record.

[0144] The personnel falling detection device provided by the embodiments of the present invention can execute the personnel falling detection method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0145] Embodiment 4

[0146] Figure 5FIG. 0 shows a schematic structural diagram of an electronic device 50 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0147] As Figure 5 shown, the electronic device 50 includes at least one processor 51, and a memory communicatively connected to the at least one processor 51, such as a read-only memory (ROM) 52, a random access memory (RAM) 53, etc. The memory stores a computer program executable by the at least one processor. The processor 51 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 52 or the computer program loaded from the storage unit 58 into the random access memory (RAM) 53. In the RAM 53, various programs and data required for the operation of the electronic device 50 can also be stored. The processor 51, the ROM 52, and the RAM 53 are connected to each other through a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0148] A plurality of components in the electronic device 50 are connected to the I / O interface 55, including: an input unit 56, such as a keyboard, a mouse, etc.; an output unit 57, such as various types of displays, speakers, etc.; a storage unit 58, such as a magnetic disk, an optical disk, etc.; and a communication unit 59, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 59 allows the electronic device 50 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0149] The processor 51 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 51 executes the various methods and processes described above, such as the person falling detection method.

[0150] In some embodiments, the person falling detection method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 50 via the ROM 52 and / or the communication unit 59. When the computer program is loaded into the RAM 53 and executed by the processor 51, one or more steps of the person falling detection method described above may be performed. Alternatively, in other embodiments, the processor 51 may be configured to execute the person falling detection method by any other suitable means (e.g., by means of firmware).

[0151] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0152] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer program, when executed by the processor, causes the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0153] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0154] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0155] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0156] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0157] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0158] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting a person's fall, characterized in that, Including: Frame extraction is performed on the video data collected by the depth camera to obtain a video frame, and the video frame includes depth values; The video frame is input into a pre-trained human pose estimation model to obtain the human key points, human detection box coordinates, and pose types of each detected human body; The three-dimensional coordinates of the human key points are determined according to the depth values; Based on the human detection box coordinates, the human detection boxes with the falling-down pose type are screened to obtain target human detection boxes; According to the three-dimensional coordinates of the human key points of the human body in the target human detection box, it is determined whether the human body in the target human detection box is in a falling-down state; If so, an alarm message for the human body in the target human detection box falling down is generated.

2. The method for detecting a person's fall according to claim 1, wherein Based on the human detection box coordinates, screening the human detection boxes with the falling-down pose type to obtain target human detection boxes, including: Calculating the aspect ratio of the human detection boxes with the falling-down pose type based on the human detection box coordinates; Determining the human detection boxes with the aspect ratio less than a preset threshold in the human detection boxes with the falling-down pose type as target human detection boxes.

3. The method for detecting a person's fall according to claim 1, characterized in that, According to the three-dimensional coordinates of the human key points of the human body in the target human detection box, determining whether the human body in the target human detection box is in a falling-down state, including: Calculating the current vector of the human body in the target human detection box by using the three-dimensional coordinates of the human key points of the human body in the target human detection box; Adding the current vector to the vector record, and deleting the vectors and times of the human bodies in the human detection boxes other than the target human detection box in the vector record, where the vector record is used to record the vectors and times of the human bodies in the target human detection box obtained after each frame extraction; Determining whether the human body in the target human detection box is in a falling-down state based on the vectors and times of the human body in the vector record.

4. The method for detecting a person's fall according to claim 3, wherein, Calculating the current vector of the human body in the target human detection box by using the three-dimensional coordinates of the human key points of the human body in the target human detection box, including: Constructing a key point matrix by using the three-dimensional coordinates of the human key points of the human body in the target human detection box; Performing normalization processing on the three-dimensional coordinates in the key point matrix to obtain a normalized key point matrix; Flattening the normalized key point matrix into a one-dimensional array to obtain the current vector of the human body in the target human detection box.

5. The method for detecting a person's fall according to claim 3, wherein Adding the current vector to the vector record, including: Judging whether there are a first vector and a second vector of the human body in the target human detection box in a pre-configured vector record, where the first vector is added to the vector record before the second vector; If the first vector and the second vector of the human body in the target human detection box do not exist, adding the current vector as the first vector to the vector record; If the first vector of the human body in the target human detection box exists and the second vector does not exist, adding the current vector as the second vector to the vector record, and adding the current time as the start time and end time to the vector record; If there are a first vector and a second vector of the human body in the target human body detection frame, replace the second vector with the current vector and replace the end time with the current time in the vector record.

6. The method for detecting a person's fall according to claim 3, characterized in that, Determining whether the human body in the target human body detection frame is in a fallen state based on the vector and time of the human body in the vector record includes: Finding a first human body to be confirmed with a first vector and a second vector in the vector record; Calculating the similarity between the first vector and the second vector of each first human body to be confirmed; Determining a second human body to be confirmed and a third human body to be confirmed from the first human bodies to be confirmed, where the second human body to be confirmed is a first human body to be confirmed with a similarity greater than or equal to a preset similarity threshold, and the third human body to be confirmed is a first human body to be confirmed with a similarity less than the similarity threshold; For each second human body to be confirmed, calculating the time difference between the end time and the start time of the vector of the second human body to be confirmed in the vector record, where the start time is the system time when the second vector is added to the vector record, and the end time is the system time when the second vector is updated; When the time difference is greater than or equal to a first time threshold, determining that the second human body to be confirmed is a fallen human body without limb movement, and deleting the vector and time of the fallen human body in the vector record; For each third human body to be confirmed, determining whether the third human body to be confirmed has fallen by using the three-dimensional coordinates of the human body key points corresponding to the first vector and the second vector of the third human body to be confirmed.

7. The method for detecting a person's fall according to claim 6, wherein, Determining whether the third human body to be confirmed has fallen by using the three-dimensional coordinates of the human body key points corresponding to the first vector and the second vector of the third human body to be confirmed includes: Calculating the average human body height by using the three-dimensional coordinates of the human body key points corresponding to the first vector and the second vector of the third human body to be confirmed; Determining a fourth human body to be confirmed with an average human body height less than a first height threshold and a fifth human body to be confirmed with an average human body height greater than or equal to the first height threshold from the third human bodies to be confirmed; For each fourth human body to be confirmed, calculating the time difference between the end time and the start time of the vector of the fourth human body to be confirmed in the vector record; When the time difference is greater than or equal to a second time threshold, determining that the fourth human body to be confirmed is a fallen human body with limb movement, and deleting the vector and time of the fallen human body in the vector record; For each fifth human body to be confirmed, determining the three-dimensional coordinates of the two hip bone key points of the fifth human body to be confirmed; When the three-dimensional coordinates of any one hip bone key point are less than a second height threshold, calculating the time difference between the end time and the start time of the vector of the fifth human body to be confirmed in the vector record; When the time difference is greater than or equal to a third time threshold, determining that the fifth human body to be confirmed is a fallen human body and is in a sitting state with limb activities after falling, and deleting the vector and time of the fallen human body in the vector record.

8. A personnel falling detection device, characterized in that, Including: A video frame extraction module for extracting frames from the video data collected by the depth camera to obtain a video frame, where the video frame includes depth values; A human body detection module, configured to input the video frame into a pre-trained human body pose estimation model to obtain the human body key points, the human body detection box coordinates, and the pose type of each detected human body; A key point three-dimensional coordinate determination module, configured to determine the three-dimensional coordinates of the human body key points according to the depth value; A human body screening module, configured to screen the human body detection boxes with the pose type of falling to the ground based on the human body detection box coordinates to obtain target human body detection boxes; A falling-to-the-ground state determination module, configured to determine whether the human body in the target human body detection box is in a falling-to-the-ground state according to the three-dimensional coordinates of the human body key points of the human body in the target human body detection box; If so, execute the falling-to-the-ground alarm module; A falling-to-the-ground alarm module, configured to generate an alarm message indicating that the human body in the target human body detection box has fallen to the ground.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the personnel falling-to-the-ground detection method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the personnel falling-to-the-ground detection method according to any one of claims 1-7 when executed by a processor.