Tumble risk assessment method and device for target object
By using an improved YOLOv11 model and continuous frame motion analysis, the accuracy and scene adaptability issues in fall detection for the elderly have been resolved, achieving efficient and accurate fall risk assessment, reducing false alarm rates, and supporting timely responses.
Patent Information
- Application Number
- CN202511080538.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies for fall detection in the elderly suffer from insufficient detection accuracy, poor adaptability to various scenarios, and high equipment costs, leading to frequent false alarms and low usage rates.
An improved YOLOv11 model is used to detect human key points in real-time monitoring video data. Combined with continuous frame motion analysis and a counter accumulation mechanism, the risk of falling is assessed by extracting weighted scores of high-level feature values and fall judgment conditions.
It improves the accuracy and efficiency of fall detection, reduces false alarm rates, provides broad coverage and non-contact fall risk assessment, and supports timely response measures.
Smart Images

Figure CN120913849A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence and computer vision, and more particularly, to a fall risk assessment method and device for a target object. BACKGROUND
[0002] In the related art, with the rapid growth of the proportion of the global aging population, the physical and mental health problems of the elderly have become the focus of attention of the society. For example, with the increase of the number of elderly people living alone, the demand for emergency assistance is gradually increasing. According to the statistics of the World Health Organization, falling is one of the main causes of accidental injury of the elderly, which can lead to fractures, brain damage and even more serious harm. Therefore, it is particularly urgent and important to study and implement an efficient and timely fall monitoring and emergency response mechanism to protect the safety of the elderly and reduce the burden of health care. SUMMARY
[0003] Embodiments of the present disclosure provide a fall risk assessment method and device for a target object, which aims to realize the detection of falling of the target object and improve the accuracy and efficiency of the detection.
[0004] In one general aspect, a fall risk assessment method for a target object is provided, which comprises: acquiring real-time monitoring video data associated with the target object, the real-time monitoring video data comprising a plurality of image data; inputting the plurality of image data into a pre-trained target detection model to obtain human key point coordinate data of the target object; performing target object state analysis processing on the human key point coordinate data to obtain state data of the target object; in response to the state data indicating that the target object has a high possibility of falling, determining a first vertical speed associated with the target object based on the speed change rate between consecutive frames in the plurality of image data; in response to the first vertical speed being greater than or equal to a first preset speed threshold, determining fall judgment data of the target object based on frame image data of a preset number of frames extracted at a preset time interval, wherein the fall judgment data comprises an included angle between the torso of the target object and the vertical direction, a height of the head of the target object from a reference surface, and a support area of the target object on the reference surface; and in response to the fall judgment data satisfying a preset condition, outputting a fall risk assessment result indicating that the target object has fallen.
[0005] Optionally, the fall risk assessment method can further comprise: for the plurality of consecutive image data, in response to the state data of the target object determined for a preset number of consecutive frames of image data all indicating that the target object has a high possibility of falling and the corresponding determined first vertical speed all being less than the first preset speed threshold, determining the fall judgment data of the target object based on the frame image data of the preset number of frames extracted at the preset time interval.
[0006] Optionally, the step of performing target object state analysis processing on the human key point coordinate data to obtain state data of the target object can include: for each frame of image data in the plurality of frames of image data, performing the following processing: performing data validity processing on the human key point coordinate data corresponding to each frame of image data to obtain processed valid key point coordinate data; in response to a classification result obtained by performing posture classification processing on the valid key point coordinate data indicating that the posture of the target object does not belong to any one of a standing posture and a sitting posture, performing high-level feature extraction on the valid key point coordinate data to obtain a weighted score of a high-level feature value; and determining the state data of the target object based on a comparison result of the weighted score of the high-level feature value and a preset score threshold.
[0007] Optionally, the step of determining the fall judgment data of the target object based on the frame image data of the preset number of frames extracted at the preset time interval can include: determining a second vertical velocity associated with the target object based on the frame image data of the preset number of frames extracted at the preset time interval; in response to the second vertical velocity being greater than or equal to a second preset speed threshold, determining that the fall judgment data of the target object includes an included angle between a torso of the target object and a vertical direction; and in response to the included angle between the torso of the target object and the vertical direction being greater than or equal to a preset angle threshold, determining that the fall judgment data of the target object includes a support area of the target object on a reference surface.
[0008] Optionally, the target detection model can be an improved YOLOv11 model, the improved YOLOv11 model including a backbone network, a neck network, and a head network, the backbone network including a multi-scale convolution attention decoding module, and the neck network including a dynamic up-sampling module.
[0009] Optionally, the target detection model can be trained in the following manner: obtaining fall video data for a training sample object, the fall video data including a plurality of frames of sample image data; pre-processing the plurality of frames of sample image data to obtain a processed sample training data set, a sample verification data set, and a sample test data set; inputting the sample training data set, the sample verification data set, and the sample test data set into the target detection model to obtain an average precision index of all categories associated with the plurality of frames of sample image data for each training cycle, and taking model parameters corresponding to a training cycle with the maximum value of the average precision index as optimal model parameters to obtain a target detection model corresponding to the optimal model parameters as the pre-trained target detection model.
[0010] Optionally, the weighted score of the high-level feature value can be calculated by the following formula: , wherein, This represents a weighted score representing high-level feature values. A marker indicating the angle between the torso of the target object and the vertical direction. A marker indicating the height of the target object's head from its reference plane. A marker indicating the area supported by the target object on the reference plane.
[0011] Alternatively, the angle between the target object's torso and the vertical direction, the height of the target object's head from the reference plane, and the supporting area of the target object on the reference plane can be calculated using the following formulas: , , , in, This indicates the angle between the target object's torso and the vertical direction. This indicates the height of the target object's head from the reference plane. This represents the area supported by the target object on the reference plane. The ordinate value represents the coordinates of the key points of the pelvis of the target object. The vertical coordinate represents the coordinates of the neck key points of the target object, while Δx and Δy represent the differences in the horizontal and vertical coordinates between the coordinates of the pelvic key points and the neck key points of the target object, respectively. This represents the reference plane's ordinate value. The ordinate value represents the coordinates of the key points on the head of the target object. This represents the preset pixel-to-meter conversion factor. Indicates key points of the ankle. Indicates the key points of the wrist. This represents the convex hull area function used to calculate the ankle and wrist keypoints.
[0012] In another general aspect, there is provided a device for fall risk assessment of a target object, the device comprising: a data acquisition module configured to acquire real-time monitoring video data associated with the target object, the real-time monitoring video data comprising a plurality of image data frames; a key point determination module configured to input the plurality of image data frames into a pre-trained target detection model to obtain human key point coordinate data of the target object; a state analysis module configured to perform target object state analysis processing on the human key point coordinate data to obtain state data of the target object; and a fall assessment module configured to: in response to the state data indicating that the target object has a high possibility of falling, determine a first vertical speed associated with the target object based on a speed change rate between consecutive frames in the plurality of image data frames; in response to the first vertical speed being greater than or equal to a first preset speed threshold, determine fall judgment data of the target object based on frame image data of a preset number of frames extracted at a preset time interval, wherein the fall judgment data comprises an included angle between a torso of the target object and a vertical direction, a height of a head of the target object from a reference surface on which the target object is located, and a support area of the target object on the reference surface; and in response to the fall judgment data satisfying a preset condition, output a fall risk assessment result indicating that the target object has fallen.
[0013] Optionally, the fall assessment module can be further configured to, for a plurality of consecutive image data frames, in response to the state data of the target object determined for a preset number of consecutive image data frames all indicating that the target object has a high possibility of falling and the first vertical speed determined for the preset number of consecutive image data frames all being less than the first preset speed threshold, determine the fall judgment data of the target object based on frame image data of a preset number of frames extracted at a preset time interval.
[0014] Optionally, the state analysis module performing target object state analysis processing on the human key point coordinate data to obtain state data of the target object can comprise: for each image data frame in the plurality of image data frames, performing the following processing: performing data validity processing on the human key point coordinate data corresponding to each image data frame to obtain processed valid key point coordinate data; in response to a classification result obtained by performing posture classification processing on the valid key point coordinate data indicating that a posture of the target object does not belong to any one of a standing posture and a sitting posture, performing high-level feature extraction on the valid key point coordinate data to obtain a weighted score of a high-level feature value; and determining the state data of the target object based on a comparison result of the weighted score of the high-level feature value and a preset score threshold.
[0015] Optionally, the operation of determining, by the fall assessment module, the fall judgment data of the target object based on the frame image data of the preset number of frames extracted at the preset time interval can include: determining, based on the frame image data of the preset number of frames extracted at the preset time interval, a second vertical speed associated with the target object; in response to the second vertical speed being greater than or equal to a second preset speed threshold, determining that an included angle between a torso of the target object and a vertical direction is included in the fall judgment data of the target object; and in response to the included angle between the torso of the target object and the vertical direction being greater than or equal to a preset angle threshold, determining that a support area of the target object on a reference surface is included in the fall judgment data of the target object.
[0016] Optionally, the target detection model can be an improved YOLOv11 model, the improved YOLOv11 model including a backbone network, a neck network, and a head network, the backbone network including a multi-scale convolution attention decoding module, and the neck network including a dynamic up-sampling module.
[0017] Optionally, the target detection model can be trained by: obtaining fall video data for a training sample object, the fall video data including a plurality of frames of sample image data; pre-processing the plurality of frames of sample image data to obtain a processed sample training data set, a sample verification data set, and a sample test data set; inputting the sample training data set, the sample verification data set, and the sample test data set into the target detection model to obtain an average precision index of all categories associated with the plurality of frames of sample image data for each training cycle, and taking model parameters corresponding to a training cycle with the maximum value of the corresponding average precision index in all training cycles as optimal model parameters to obtain a target detection model corresponding to the optimal model parameters as the pre-trained target detection model.
[0018] Optionally, the weighted score of the high-level feature value can be calculated by the following formula: , wherein, represents the weighted score of the high-level feature value, represents a label of the included angle between the torso of the target object and the vertical direction, represents a label of the height of the head of the target object from the reference surface, represents a label of the support area of the target object on the reference surface.
[0019] Optionally, the included angle between the torso of the target object and the vertical direction, the height of the head of the target object from the reference surface, and the support area of the target object on the reference surface can be calculated by the following formula: , , , wherein, represents an included angle between the torso of the target object and the vertical direction, represents a height of the head of the target object from the reference surface, represents a support area of the target object on the reference surface, represents a longitudinal coordinate value of the pelvic key point coordinate of the target object, represents a longitudinal coordinate value of the neck key point coordinate of the target object, and Δx and Δy respectively represent a horizontal coordinate difference value and a longitudinal coordinate difference value between the pelvic key point coordinate and the neck key point coordinate of the target object, represents a reference longitudinal coordinate value of the reference surface, represents a longitudinal coordinate value of the head key point coordinate of the target object, represents a preset pixel-to-meter conversion coefficient, represents an ankle key point, represents a wrist key point, represents a convex hull area function for calculating the ankle key point and the wrist key point.
[0020] In another general aspect, there is provided a computer program product including computer programs / instructions that, when executed by a processor, implement the method for fall risk assessment of a target object as described above.
[0021] In another general aspect, there is provided a computer-readable storage medium that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, enables the electronic device / server to perform the method for fall risk assessment of a target object as described above.
[0022] In another general aspect, there is provided a computing device including at least one processor; at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform the method for fall risk assessment of a target object as described above.
[0023] The method and device for fall risk assessment of a target object according to embodiments of the present disclosure achieve fall detection of a target object and improve detection accuracy and detection efficiency by proposing a target object fall detection strategy. In addition, by using an improved YOLOv11 model to monitor the fall of a target object in real time, more accurate feature information for judging whether to fall can be extracted. In addition, by using a continuous frame motion analysis and a counter accumulation mechanism, the accuracy of detection can be improved and the false positive rate can be reduced. BRIEF DESCRIPTION OF DRAWINGS
[0024] The above and other objects and features of the present disclosure will become more apparent from the following description made with reference to the accompanying drawings, in which: Figure 1 is a flowchart illustrating a fall risk assessment method for a target object according to an embodiment of the present disclosure; Figure 2 is a flowchart illustrating an example of a fall risk assessment method for a target object according to an embodiment of the present disclosure; Figure 3 is an example flowchart illustrating a key point state analysis process according to an embodiment of the present disclosure; Figure 4 is an example flowchart illustrating a dynamic verification process according to an embodiment of the present disclosure; Figure 5 is a schematic diagram of an overall network architecture of an improved YOLOv11 model according to an embodiment of the present disclosure; Figure 6 is a structural block diagram of a fall risk assessment apparatus for a target object according to an embodiment of the present disclosure; Figure 7 is a block diagram of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The following detailed description is provided to help the reader understand the method, device and / or system described herein. However, various changes, modifications and equivalents can be resorted to by those skilled in the art after understanding the disclosure of the present application. For example, the order of the operations described herein is merely an example, and is not limited to those set forth herein, but can be changed as will be apparent to those skilled in the art after understanding the disclosure of the present application, except for operations that must occur in a specific order. Also, the description of features known in the art can be omitted for the sake of clarity and conciseness.
[0026] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals refer to like elements throughout. The embodiments will be explained by referring to the drawings in order to explain the present disclosure.
[0027] In the related art, detection and protection technologies for falls of the elderly mainly include: a detection method based on biological feature monitoring, a detection method based on scene intelligent perception, etc.
[0028] Specifically, for the detection method based on biometric monitoring, focus on collecting physiological and motion characteristic data of the elderly (for example, through monitoring devices such as heart rate sensors, body motion sensors, etc.), real-time tracking of physiological signal fluctuations and motion posture changes, trying to trigger a fall alarm through abnormal data. However, the natural body shaking in the daily activities of the elderly, the normal fluctuations of physiological indicators, are easy to interfere with the accuracy of the data judgment, and finally lead to frequent false alarms. In addition, long-term wearing of such monitoring devices is easy to make the elderly have psychological resistance, and some elderly people refuse to cooperate due to resistance, which greatly reduces the actual use rate and monitoring effectiveness of the device.
[0029] For the detection method based on scene intelligent perception, the core is to build an intelligent scene perception system, such as deploying infrared sensors, sound sensors and other perception devices in the activity space of the elderly, and determining the fall by analyzing the human motion trajectory, speed mutation and other information in the space. However, this method has the following disadvantages: scene intelligent transformation requires a large amount of equipment deployment cost, and the later maintenance is complex; the perception range has natural limitations, it is difficult to cover all the activity range of the elderly without dead angle, in addition, there are other objects, multiple target activities and other interference factors in complex environment, which can easily reduce the perception accuracy and affect the reliability of the detection result.
[0030] In view of the defects and deficiencies in detection accuracy, scene adaptability and protection process integrity in the prior art and the problems in the above and other aspects, the present disclosure provides a fall risk assessment method and device for a target object, which can quickly and effectively determine the fall of the target object by monitoring the real-time state of the target object through video monitoring, and notify the relevant medical staff and provide the location information and personal situation of the fallen object, thereby realizing high detection efficiency in the fall condition processing aspect, and further providing timely response measures to achieve the purpose of effective prevention.
[0031] Reference will now be made to Figures 1 to 7 The fall risk assessment method and device for a target object according to the embodiments of the present disclosure are described in detail.
[0032] First, reference is made to Figures 1 to 5 The fall risk assessment method for a target object according to the embodiments of the present disclosure is described in detail.
[0033] Figure 1 is a flowchart illustrating a fall risk assessment method for a target object according to an embodiment of the present disclosure. Figure 2 is a flowchart illustrating an example of a fall risk assessment method for a target object according to an embodiment of the present disclosure. Figure 3 is an example flowchart illustrating a key point state analysis process according to an embodiment of the present disclosure. Figure 4 is an example flowchart illustrating a dynamic verification process according to an embodiment of the present disclosure.Figure 5 is a schematic diagram of the overall network architecture of the improved YOLOv11 model according to an embodiment of the present disclosure.
[0034] With reference to Figure 1 , according to an embodiment of the present disclosure, at step S101, real-time monitoring video data associated with a target object is acquired. Here, the real-time monitoring video data includes multiple frames of image data.
[0035] According to an embodiment of the present disclosure, at step S102, the multiple frames of image data are input to a pre-trained target detection model to obtain human key point coordinate data of the target object. Here, the human key point coordinate data can include coordinates of 14 human key points (i.e., neck, right shoulder, right elbow, right wrist, left shoulder, left elbow, left wrist, right hip, right knee, right ankle, left hip, left knee, left ankle, and pelvic center).
[0036] For example, the target detection model can be an improved YOLOv11 model, which includes a backbone network, a neck network, and a head network, the backbone network includes a multi-scale convolution attention decoding module, and the neck network includes a dynamic upsampling module.
[0037] Here, for example, with reference to Figure 5 , the overall network architecture of the improved YOLOv11 model includes a backbone network Backbone, a neck network Neck, and a head network Head. The backbone network Backbone includes a Conv (convolution block), a C3k2_EMCAD (a C3k2 module with added multi-scale convolution attention decoding), an SPPF (spatial pyramid fast pooling) module (including a CBS (convolution-batch normalization-Silu module), a Maxpool (max-pooling block), and a Concat (concatenation block)), and a C2PSA module composed of a Conv layer, a PSA Block (multi-scale split attention block), and a Concat operation, wherein the PSA Block inside includes an Attention (attention mechanism), an Add (addition), a Conv, and a shortcut (shortcut connection) operation. The neck network Neck includes a DySample (dynamic upsampling module), a Concat, a Conv, and a C3k2_EMCAD. The head network Head includes a Pose (pose prediction head).
[0038] By adding a double attention module to the original YOLOv11 and improving the loss function of the algorithm, the relevant feature information can be extracted more accurately, and the algorithm accuracy can be improved.
[0039] In addition, the target detection model can be trained by the following steps S21 to S23: At step S21, fall video data for a training sample object is obtained, the fall video data including a plurality of frames of sample image data.
[0040] At step S22, the plurality of frames of sample image data is preprocessed to obtain a processed sample training data set, a sample validation data set and a sample test data set.
[0041] At step S23, the sample training data set, the sample validation data set and the sample test data set are input into the target detection model to obtain an average precision indicator of all classes associated with the plurality of frames of sample image data for each training cycle, and the model parameters corresponding to the training cycle with the maximum value of the average precision indicator in all training cycles are taken as the optimal model parameters, and a target detection model corresponding to the optimal model parameters is obtained as a pre-trained target detection model.
[0042] As an example, in the case of the improved YOLOv11 model as the target detection model, the following indicators of the model are verified at the end of each round of training in the above training process: precision, recall, mAP50 (average precision of the model at an IoU threshold of 0.5) indicator, mAP50-95 (average precision of the model at an IoU threshold ranging from 0.5 to 0.95) indicator, and the model weight with the highest mAP50-95 indicator is selected as the final selected model parameter, and the corresponding fall risk assessment model is obtained.
[0043] Here, in the target detection task of the present disclosure, mAP (Mean Average Precision) represents the average precision mean of all classes, which is obtained by calculating the precision mean (AP, Average Precision) of each class and then averaging these AP values. mAP can measure the comprehensive performance of the model in terms of detection precision and recall for different object classes.
[0044] In addition, the calculation formulas of recall and precision are as follows: Recall=TP / (TP+FN) (1) Precision=TP / (TP+FP) (2) Where TP (True Position) represents true positive and represents correctly classified positive samples, FP (False Positives) represents false positive and represents incorrectly classified negative samples, and FN (False negatives) represents false negative and represents incorrectly classified positive samples.
[0045] In addition, it can be concluded from Table 1 below that the improved YOLOv11 model has improved in various indicators.
[0046] Table 1 Comparison table of target detection model evaluation indexes
[0047] The trained target detection model obtained by the above training method can overcome the problem of insufficient pose recognition accuracy of target objects in complex scenes, such as missing detection in cases of multiple people interacting, clothing blocking, and the like. Moreover, by using the improved YOLOv11 model, the performance of various indexes of the model can be improved.
[0048] According to an embodiment of the present disclosure, in step S103, the human key point coordinate data is subjected to target object state analysis processing to obtain state data of the target object.
[0049] As an example, step S103 can further include: for each frame of image data in the plurality of frames of image data, performing the following processing S1031 to S1033: In processing S1031, the human key point coordinate data corresponding to each frame of image data is subjected to data validity processing to obtain processed valid key point coordinate data.
[0050] In processing S1032, in response to a classification result indicating that the pose of the target object does not belong to any one of the standing pose and the sitting pose after the pose classification processing on the valid key point coordinate data, the valid key point coordinate data is subjected to high-level feature extraction to obtain a weighted score of the high-level feature value.
[0051] In processing S1033, based on a comparison result of the weighted score of the high-level feature value and a preset score threshold, the state data of the target object is determined.
[0052] Through the improvement of the target detection model based on the present disclosure in the extraction capability of fine-grained features related to the falling of the target object, the behaviors of daily lying, sitting, and the like can be accurately distinguished from the real falling actions, greatly reducing the false positive rate.
[0053] As an example, the weighted score of the high-level feature value can be calculated by the following formula (1): (3) wherein, represents the weighted score of the high-level feature value, represents a label of an included angle between a torso of the target object and a vertical direction, represents a label of a height of a head of the target object from a reference surface (e.g., a ground surface on which the target object is located, etc.) in which the target object is located, represents a label of a support area of the target object on the reference surface.
[0054] In addition, for example, the angle between the torso of the target object and the vertical direction, the height of the head of the target object from the reference surface, and the support area of the target object on the reference surface can be calculated by the following formulas (2) to (4), respectively: (4) (5) (6) wherein, represents the angle between the torso of the target object and the vertical direction, represents the height of the head of the target object from the reference surface, represents the support area of the target object on the reference surface, represents the longitudinal coordinate value of the pelvic key point coordinate of the target object, represents the longitudinal coordinate value of the neck key point coordinate of the target object, and Δx and Δy represent the horizontal coordinate difference and the longitudinal coordinate difference between the pelvic key point coordinate and the neck key point coordinate of the target object, respectively, represents the reference surface reference longitudinal coordinate value, represents the longitudinal coordinate value of the head key point coordinate of the target object, represents a preset pixel-to-meter conversion coefficient, represents an ankle key point, represents a wrist key point, represents a convex hull area function for calculating the ankle key point and the wrist key point.
[0055] In the present disclosure, by adopting the above-mentioned advanced features related to human body key points, the problem of insufficient pose recognition accuracy of the target object in a complex scene (e.g., missed detection) can be overcome, and the falling action can be accurately recognized, thereby reducing false positives.
[0056] According to an embodiment of the present disclosure, in step S104, in response to the state data indicating that the target object is likely to fall, a first vertical speed associated with the target object is determined based on the speed change rate between consecutive frames in the plurality of frames of image data.
[0057] According to an embodiment of the present disclosure, in step S105, in response to the first vertical speed being greater than or equal to a first preset speed threshold, falling judgment data of the target object is determined based on frame image data of a preset number of frames extracted at a preset time interval.
[0058] According to an embodiment of the present disclosure, if it is determined in step S105 that the first vertical speed is less than the first preset speed threshold, the value of the counter is increased by 1, and the operation in step S103 and the subsequent related operations are performed on the next frame of image data, until the following condition is met: "for a plurality of consecutive frames of image data, the state data of the target object determined in response to the image data of the consecutive preset number of frames all indicate that the target object has a high possibility of falling and the corresponding determined first vertical speed is all less than the first preset speed threshold", after the above condition is met, the processing of determining the fall judgment data of the target object based on the preset number of frames of image data extracted at the preset time interval is performed.
[0059] For example, in the case where the preset number in the above condition is 3, in response to the state data of the target object determined in response to the image data of the consecutive 3 frames all indicating that the target object has a high possibility of falling and the corresponding determined first vertical speed being all less than the first preset speed threshold, the processing of determining the fall judgment data of the target object based on the preset number of frames of image data extracted at the preset time interval can be performed.
[0060] Here, the fall judgment data includes the angle between the torso of the target object and the vertical direction, the height of the head of the target object from the reference surface, and the support area of the target object on the reference surface.
[0061] As an example, the step of determining the fall judgment data of the target object based on the preset number of frames of image data extracted at the preset time interval in step S105 can further include steps S1051 to S1053: In step S1051, a second vertical speed associated with the target object is determined based on the preset number of frames of image data extracted at the preset time interval.
[0062] In step S1052, in response to the second vertical speed being greater than or equal to a second preset speed threshold, the angle between the torso of the target object and the vertical direction is determined to be included in the fall judgment data of the target object.
[0063] In step S1053, in response to the angle between the torso of the target object and the vertical direction being greater than or equal to a preset angle threshold, the support area of the target object on the reference surface is determined to be included in the fall judgment data of the target object.
[0064] In the present disclosure, by means of continuous frame motion analysis and counter accumulation mechanism, the accuracy of detection is improved and the false positive rate is reduced.
[0065] According to an embodiment of the present disclosure, in step S106, in response to the fall judgment data meeting the preset condition, a fall risk assessment result indicating that the target object has fallen is output.
[0066] The fall risk assessment method of the present disclosure does not need to contact the body of the target object and has wide coverage, and by fusing the more optimal feature extraction and recognition strategy based on the improved YOLOv11 algorithm, an integrated method covering fall prevention, accurate detection and efficient disposal is constructed, providing more reliable technical support for the safety protection of the target object.
[0067] Hereinafter, with reference to Figures 2 to 5 The fall risk assessment method for a target object according to an embodiment of the present disclosure is described by way of example.
[0068] With reference to Figure 2 In step S201, real-time video stream input is obtained, for example, a segment of monitoring video data is extracted, and each frame of image data is obtained.
[0069] In step S202, the pre-trained improved YOLOv11 model is used to detect the real-time state of the target object, and the human body key point coordinates of the target object are output.
[0070] In step S203, key point state analysis is performed based on the human body key point coordinates. If the state of the target object is determined to be suspected of falling, go to step S205, otherwise, go to step S204, and output "normal behavior".
[0071] Here, as an example, with reference to Figure 3 , step S203 can further include steps S301 to S319.
[0072] In step S301, the single-frame human body key point coordinates are taken as the input of the key point state analysis.
[0073] In step S302, data validity test is performed on the input to determine whether the confidence of the data is greater than 0.8. If it is satisfied, go to step S305 to perform basic posture classification, otherwise go to step S303 to perform Kalman filtering after completion and then perform basic posture classification.
[0074] In S305, if it is determined that the target object is in a standing posture or a sitting posture, go to step S304 to output normal, otherwise go to step S306 to perform advanced feature extraction calculation, that is, calculate the trunk-vertical direction angle , the actual height H of the head from the reference surface, and the support area A, respectively.
[0075] If it is determined in S307 that ≥ 50°, then in S308 _flag = 1, otherwise in S309 _flag = 0.
[0076] If it is judged in S310 that H≤0.6m, then H_flag=1 is set in S311, otherwise H_flag=0 is set in S312.
[0077] If it is judged in S313 that A≤0.4㎡, then A_flag=1 is set in S314, otherwise A_flag=0 is set in S315.
[0078] In S316, the above three flags are weighted to obtain a total score S, and in S317, it is judged whether S is greater than or equal to 0.7, if yes, go to S319, determine as suspected fall, otherwise go to S318, determine as watchful.
[0079] Referring back to Figure 2 , in step S205, continuous frame motion analysis is performed.
[0080] Here, as an example, referring to Figure 4 , the continuous frame motion analysis can further include steps S401 to S410.
[0081] In step S401, dynamic verification is triggered. In step S402, 3 frame image data are loaded, their acquisition time interval Δt=1 / 30s, and in step S403, the vertical velocity v_z is calculated.
[0082] Here, the calculation formula of the vertical velocity is as follows: (7) wherein, represents the current frame pelvis y coordinate, represents the second frame pelvis y coordinate, Δt represents the frame time interval, and pixel2meter represents the pixel to meter conversion coefficient.
[0083] In S404, it is judged whether |v_z|≥1.2m / s is satisfied, if yes, in S405, actual angle verification is performed, otherwise go to S410 to discard the data. In S406, it is judged whether the actual angle satisfies ≥30°, if yes, in S407, support area verification is performed, otherwise go to S410 to discard the data. In S408, it is judged whether the support area A≤0.3㎡ is satisfied, if yes, go to S409 to trigger an alarm, otherwise discard the data.
[0084] Referring back to Figure 2 , in step S206, the speed change rate between each frame is detected, and the vertical velocity is calculated. If the vertical velocity is greater than or equal to 1.5m / s, go to step S209, otherwise go to step S207 to make the counter plus 1.
[0085] In step S208, it is determined whether the counter is equal to 3, i.e., three consecutive frames of abnormalities occur, and if so, it is determined that dynamic verification is performed in step S209.
[0086] In step S212, the trunk-vertical direction angle of the target object, the actual height H of the head from the reference surface, and the support area A are calculated, and it is determined whether the following conditions are satisfied: ≥ 30°, H ≤ 0.5 m, and A ≤ 0.3 m2. If all three conditions are met, it is determined that an alarm is triggered in step S211, and the fall area is marked in step S211. Otherwise, it is determined to be a false alarm in step S213.
[0087] Next, the fall risk assessment device 600 for a target object according to an embodiment of the present disclosure will be described in detail with reference to Figure 6
[0088] Figure 6 is a structural block diagram illustrating the fall risk assessment device 600 for a target object according to an embodiment of the present disclosure.
[0089] Referring to Figure 6 , the fall risk assessment device 600 for a target object according to an embodiment of the present disclosure can include a data acquisition module 610, a key point determination module 620, a state analysis module 630, and a fall assessment module 640.
[0090] According to an embodiment of the present disclosure, the data acquisition module 610 can perform: acquiring real-time monitoring video data associated with a target object. Here, the real-time monitoring video data includes multiple frames of image data.
[0091] According to an embodiment of the present disclosure, the key point determination module 620 can perform: inputting the multiple frames of image data into a pre-trained target detection model to obtain human key point coordinate data of the target object.
[0092] For example, the target detection model can be an improved YOLOv11 model. Here, the improved YOLOv11 model includes a backbone network, a neck network, and a head network, the backbone network includes a multi-scale convolution attention decoding module added on the basis of the original YOLOv11 model, and the neck network includes a dynamic up-sampling module to replace the up-sampling model in the neck network of the original YOLOv11 model.
[0093] In addition, the target detection model can be trained through the following processes (1) to (3): In process (1), fall video data for a training sample object is acquired.
[0094] Here, the fall video data includes multiple frames of sample image data.
[0095] In processing (2), the multi-frame sample image data is preprocessed (for example, picture preprocessing is performed by a labelimg tool) to obtain a processed sample training data set, a sample verification data set, and a sample test data set.
[0096] In processing (3), the sample training data set, the sample verification data set, and the sample test data set are input into the target detection model to obtain an average precision index of all categories associated with the multi-frame sample image data for each training cycle, and the model parameters corresponding to the training cycle with the maximum value of the average precision index in all training cycles are taken as optimal model parameters, so as to obtain a target detection model corresponding to the optimal model parameters as a pre-trained target detection model.
[0097] According to an embodiment of the present disclosure, the state analysis module 630 can perform target object state analysis processing on the human key point coordinate data to obtain state data of the target object.
[0098] As an example, the state analysis module 630 can perform operations 631) to 633) on each frame of image data in the multi-frame image data: In operation 631), data validity processing is performed on the human key point coordinate data corresponding to each frame of image data to obtain processed valid key point coordinate data.
[0099] In operation 632), in response to a classification result of the posture classification processing on the valid key point coordinate data indicating that the posture of the target object does not belong to any of the standing posture and the sitting posture, advanced feature extraction is performed on the valid key point coordinate data to obtain a weighted score of the advanced feature value.
[0100] For example, the weighted score of the advanced feature value can be calculated by formula (3) described above.
[0101] In operation 633), the state data of the target object is determined based on a comparison result of the weighted score of the advanced feature value and a preset score threshold.
[0102] According to an embodiment of the present disclosure, the fall assessment module 640 can perform operations 641) to 643): In operation 641), in response to the state data indicating that the target object has a high fall probability, a first vertical speed associated with the target object is determined based on a speed change rate between consecutive frames in the multi-frame image data.
[0103] In operation 642), in response to the first vertical speed being greater than or equal to a first preset speed threshold, fall judgment data of the target object is determined based on frame image data of a preset number of frames extracted at a preset time interval.
[0104] In addition, the fall assessment module 640 can further perform: determining, for the continuous multiple frames of image data, fall judgment data of the target object based on the preset number of frames of frame image data extracted at the preset time interval, in response to the state data of the target object determined for the continuous preset number of frames of image data all indicating that the target object has a high possibility of falling and the corresponding determined first vertical speed all being less than the first preset speed threshold.
[0105] Here, the fall judgment data includes an angle between a torso of the target object and a vertical direction, a height of a head of the target object from a reference surface on which the target object is located, and a support area of the target object on the reference surface on which the target object is located.
[0106] As an example, the operation of determining, by the fall assessment module, the fall judgment data of the target object based on the preset number of frames of frame image data extracted at the preset time interval in operation 642) can further include operations 6421) to 6423): In operation 6421), a second vertical speed associated with the target object is determined based on the preset number of frames of frame image data extracted at the preset time interval.
[0107] In operation 6422), in response to the second vertical speed being greater than or equal to a second preset speed threshold, an angle between a torso of the target object and a vertical direction is determined to be included in the fall judgment data of the target object. For example, the angle between the torso of the target object and the vertical direction can be calculated by the above-described equation (4).
[0108] In operation 6423), in response to the angle between the torso of the target object and the vertical direction being greater than or equal to a preset angle threshold, a support area of the target object on a reference surface on which the target object is located is determined to be included in the fall judgment data of the target object. For example, the support area of the target object on the reference surface on which the target object is located can be calculated by the above-described equation (6).
[0109] In operation 643), in response to the fall judgment data satisfying a preset condition, a fall risk assessment result indicating that the target object has fallen is output.
[0110] It should be noted that the operations performed with respect to each of the above-described structural blocks can be similar to those described with reference to Figure 1 and will not be described again here.
[0111] Figure 7 is a block diagram illustrating a computing device 700 according to an embodiment of the present disclosure.
[0112] With reference to Figure 7According to embodiments of the disclosure, the computing device 700 can include a processor 710 and a memory 720. The processor 710 can include, but is not limited to, a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a system on chip (SoC), a microprocessor, an application specific integrated circuit (ASIC), etc. The memory 720 can store computer-executable instructions to be executed by the processor 710. The memory 720 includes a high-speed random access memory and / or a non-volatile computer-readable storage medium. When the processor 710 executes the computer-executable instructions stored in the memory 720, the fall risk assessment method for a target object as described above can be implemented.
[0113] The fall risk assessment method for a target object according to embodiments of the disclosure can be written as computer programs / instructions to form a computer program product and stored on a computer-readable storage medium. When the computer programs / instructions are executed by a processor, the fall risk assessment method for a target object as described above can be implemented. When the instructions in the computer-readable storage medium are executed by the processor of an electronic device / server, the electronic device / server is enabled to perform the fall risk assessment method for a target object as described above. Examples of the computer-readable storage medium include a read-only memory (ROM), a random access programmable read-only memory (PROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash memory, a non-volatile memory, a CD-ROM, a CD-R, a CD+R, a CD-RW, a CD+RW, a DVD-ROM, a DVD-R, a DVD+R, a DVD-RW, a DVD+RW, a DVD-RAM, a BD-ROM, a BD-R, a BD-R LTH, a BD-RE, a Blu-ray or an optical disk memory, a hard disk drive (HDD), a solid state drive (SSD), a card memory (such as a multimedia card, a secure digital (SD) card, or an extreme digital (XD) card), a magnetic tape, a floppy disk, a magneto-optical data storage device, an optical data storage device, a hard disk, a solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or a computer so that the processor or the computer can execute the computer program. In one example, the computer program and any associated data, data files, and data structures are distributed over a networked computer system so that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0114] According to the fall risk assessment method and device for a target object, the falling detection strategy for the target object is proposed, so that the falling detection of the target object is realized, and the detection accuracy and efficiency are improved.
[0115] On the other hand, according to the fall risk assessment method and device for a target object, the improved YOLOv11 model is used to monitor the falling of the target object in real time, so that the feature information related to the falling of the target object can be extracted more accurately.
[0116] On the other hand, according to the fall risk assessment method and device for a target object, the continuous frame motion analysis and the counter accumulation mechanism are used, so that the detection accuracy can be improved and the false positive rate can be reduced.
[0117] In addition, the fall risk assessment method for a target object can be mainly applied to the evaluation of the falling of the elderly. However, the present disclosure is not limited thereto, and the fall risk assessment method for a target object can also be applied to the falling detection and evaluation of other specific groups of people other than the elderly.
[0118] In addition, the algorithm design of the fall risk assessment method for a target object according to the present disclosure is reasonable and has strong practicability, and the implementation cost is low.
[0119] Although some embodiments of the present disclosure have been disclosed and described, those skilled in the art should understand that modifications and variations can be made to these embodiments without departing from the concept and spirit of the present disclosure, which is defined by the claims and their equivalents.
Claims
1. A fall risk assessment method for a target subject, characterized by, The fall risk assessment method comprises the following steps: obtaining real-time monitoring video data associated with a target object, the real-time monitoring video data comprising a plurality of image data frames; inputting the plurality of image data frames into a pre-trained target detection model to obtain human key point coordinate data of the target object; performing target object state analysis processing on the human key point coordinate data to obtain state data of the target object; in response to the state data indicating that the target object has a high possibility of falling, determining a first vertical speed associated with the target object based on a speed change rate between consecutive frames in the plurality of image data frames; in response to the first vertical speed being greater than or equal to a first preset speed threshold, determining fall judgment data of the target object based on frame image data of a preset number of frames extracted at a preset time interval, wherein the fall judgment data comprises an included angle between the torso of the target object and the vertical direction, a height of the head of the target object from a reference surface, and a support area of the target object on the reference surface; in response to the fall judgment data satisfying a preset condition, outputting a fall risk assessment result indicating that the target object has fallen.
2. The fall risk assessment method according to claim 1, characterized in that, The fall risk assessment method further comprises: for a plurality of consecutive image data frames, in response to the state data of the target object determined for a preset number of consecutive image data frames all indicating that the target object has a high possibility of falling and the first vertical speed determined for the image data all being less than the first preset speed threshold, determining the fall judgment data of the target object based on frame image data of a preset number of frames extracted at a preset time interval.
3. The fall risk assessment method according to claim 1, characterized in that, The step of performing target object state analysis processing on the human key point coordinate data to obtain state data of the target object comprises: for each image data frame in the plurality of image data frames, performing the following processing: performing data validity processing on the human key point coordinate data corresponding to each image data frame to obtain processed valid key point coordinate data; in response to a classification result obtained by performing posture classification processing on the valid key point coordinate data indicating that the posture of the target object does not belong to any one of a standing posture and a sitting posture, performing high-level feature extraction on the valid key point coordinate data to obtain a weighted score of a high-level feature value; determining the state data of the target object based on a comparison result of the weighted score of the high-level feature value and a preset score threshold.
4. The fall risk assessment method according to claim 1, characterized in that, The step of determining the fall judgment data of the target object based on frame image data of a preset number of frames extracted at a preset time interval comprises: determining a second vertical speed associated with the target object based on frame image data of a preset number of frames extracted at a preset time interval; in response to the second vertical speed being greater than or equal to a second preset speed threshold, determining that the included angle between the torso of the target object and the vertical direction is included in the fall judgment data of the target object; in response to the included angle between the torso of the target object and the vertical direction being greater than or equal to a preset angle threshold, determining that the support area of the target object on the reference surface is included in the fall judgment data of the target object.
5. The fall risk assessment method according to claim 1, characterized in that, The target detection model is trained in the following manner: obtaining fall video data for a training sample object, the fall video data comprising a plurality of sample image data frames; Preprocess the plurality of frame sample image data to obtain a processed sample training data set, a sample verification data set, and a sample test data set; input the sample training data set, the sample verification data set, and the sample test data set into the target detection model to obtain an average precision index of all categories associated with the plurality of frame sample image data for each training cycle, and take the model parameters corresponding to the training cycle with the maximum value of the corresponding average precision index in all training cycles as the optimal model parameters, to obtain a target detection model corresponding to the optimal model parameters as the pre-trained target detection model, The target detection model is an improved YOLOv11 model, and the improved YOLOv11 model includes a backbone network, a neck network, and a head network, the backbone network includes a multi-scale convolution attention decoding module, and the neck network includes a dynamic up-sampling module.
6. The fall risk assessment method according to claim 3, characterized in that, The weighted score of the high-level feature value is calculated by the following formula: , wherein, a weighted score representing a high-level feature value, a marker representing an angle of a torso of a target object with a vertical direction, a marker representing a height of a head of a target object from a reference surface, a marker representing a support area of a target object on a reference surface.
7. The fall risk assessment method according to claim 1, characterized in that, The angle between the torso of the target object and the vertical direction, the height of the head of the target object from the reference surface, and the support area of the target object on the reference surface are calculated by the following formula: , , , wherein, represents the included angle between the torso of the target object and the vertical direction, represents the height of the head of the target object from the reference surface, represents the support area of the target object on the reference surface, represents the longitudinal coordinate value of the pelvic key point coordinate of the target object, represents the longitudinal coordinate value of the neck key point coordinate of the target object, and Δx and Δy respectively represent the horizontal coordinate difference and the longitudinal coordinate difference between the pelvic key point coordinate and the neck key point coordinate of the target object, represents the reference longitudinal coordinate value of the reference surface, represents the longitudinal coordinate value of the head key point coordinate of the target object, represents a preset pixel-to-meter conversion coefficient, represents an ankle key point, represents a wrist key point, represents a convex hull area function for calculating the ankle key point and the wrist key point.
8. A fall risk assessment apparatus for a target subject, characterized by, The fall risk assessment device includes: The data acquisition module is configured to acquire real-time monitoring video data associated with the target object, the real-time monitoring video data including a plurality of frame image data; The key point determination module is configured to input the plurality of frame image data into a pre-trained target detection model to obtain human key point coordinate data of the target object; The state analysis module is configured to perform target object state analysis processing on the human key point coordinate data to obtain state data of the target object; The fall assessment module is configured to: In response to the state data indicating that the target object has a high fall probability, determine a first vertical velocity associated with the target object based on the speed change rate between consecutive frames in the plurality of frame image data; In response to the first vertical velocity being greater than or equal to a first preset speed threshold, determine fall judgment data of the target object based on frame image data of a preset number of frames extracted at a preset time interval, wherein the fall judgment data includes the angle between the torso of the target object and the vertical direction, the height of the head of the target object from the reference surface, and the support area of the target object on the reference surface; In response to the fall judgment data satisfying a preset condition, output a fall risk assessment result indicating that the target object has fallen.
9. A computer program product, characterised in that, The computer program product includes computer programs / instructions that, when executed by a processor, implement the fall risk assessment method for a target object as claimed in any one of claims 1 to 7.
10. A computing device, comprising: The computing device includes at least one processor and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, cause the at least one processor to perform the fall risk assessment method for a target object as claimed in any one of claims 1 to 7.