Method, device and computer program for detecting personnel falls

Through the combination of the improved YOLOv11-pose model and the T-STGCN model, the problems of misjudgment and missed detection of human falls in complex occlusion environments are solved, accurate detection of fall behavior is achieved, and safety management capabilities in industrial scenarios are improved.

CN120048005BActive Publication Date: 2025-07-11HEBEI BAISHA TOBACCO

Patent Information

Application Number
CN202510533832.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-11
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In complex occlusion environments, the existing human fall detection technology has the problems of high misjudgment rate and high missed detection rate. Especially in industrial scenarios, traditional methods are difficult to effectively distinguish between real fall status and high-intensity operation, and the equipment is highly invasive and the multi-node data coordination efficiency is low.

Method used

The improved YOLOv11-pose model is used to integrate the weighted dynamic fusion function module to extract bone key points, and combined with the T-STGCN model, the Transformer encoder is used to perform behavior detection. By dynamically capturing the movement trajectory and posture changes of the human body's active area in complex spaces, the dependence relationship between long-distance time states is established to achieve accurate detection of fall behavior.

Benefits of technology

It improves detection accuracy and robustness in complex occlusion environments, reduces misjudgment and missed detection, and provides reliable security management support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048005B_ABST
    Figure CN120048005B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device and computer program for detecting human falls, which relates to the technical field of image recognition and processing. The method includes: obtaining a real-time video data stream of a complex occlusion video scene; inputting the real-time video data stream into a skeletal key point extraction model to obtain extraction result data; inputting the extraction result data into a behavior detection model to output a detection result of a fall behavior; wherein, the skeletal key point extraction model is constructed based on an improved YOLOv11-pose model, and the improved YOLOv11-pose model incorporates a weighted dynamic fusion function WDFM module; the behavior detection model is constructed based on the T-STGCN model; the T-STGCN model incorporates a Transformer encoder on the basis of the STGCN model. The present invention can solve the problems of easy loss of key features and difficult modeling of temporal features in complex occlusion scenarios, and achieve accurate detection of fall behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition processing, and particularly to a method, device and computer program for detecting human falls. Background Art

[0002] In key fields such as industrial production, medical health monitoring and public safety, real-time and accurate human fall detection technology is of crucial practical significance for reducing the risk of accidental injuries and improving the efficiency of safety emergency response. Especially in complex industrial environments such as mechanical workshops and warehousing and logistics centers, problems such as dynamic occlusion (such as operation equipment occlusion, multi-person interaction occlusion) and uneven illumination are widespread, severely restricting the reliability of traditional detection methods, and at the same time highlighting the urgent need for highly robust detection algorithms in complex environments.

[0003] With the rapid development of artificial intelligence technology, related research methods have gradually penetrated deeply into the field of fall detection. Currently, the mainstream research methods are mainly divided into two categories: one is the threshold judgment method based on wearable sensors. Such methods usually rely on sensors such as accelerometers and gyroscopes to collect human motion data, and judge the fall state by setting thresholds. Although this method can avoid the problem of visual occlusion to a certain extent, its inherent defects such as strong device invasiveness and low multi-node data collaboration efficiency, as well as practical application problems such as device update and maintenance and wearing forgetfulness, limit the wide deployment of wearable devices in complex industrial scenarios. The other is the computer vision algorithm based on a monocular camera. Such methods usually use deep learning models such as OpenPose and YOLO to extract human skeleton key points and pose features (such as torso inclination angle, joint speed, etc.), and then realize the recognition and classification of fall behaviors. However, the above traditional methods based on computer vision still face significant limitations in complex occlusion environments: sensor-dependent solutions are difficult to effectively distinguish real fall states from high-intensity operation action states, and are prone to misjudgment; while traditional vision algorithms are prone to losing key feature information when part of the target human body is occluded, resulting in an increase in the false detection rate and even the risk of missed detection, seriously affecting the detection accuracy and reliability. Summary of the Invention

[0004] Embodiments of the present invention provide a method, device and computer program for detecting human falls to solve the problem of accurately detecting human fall behaviors in complex occlusion environments.

[0005] In a first aspect, embodiments of the present invention provide a method for detecting human falls, including:

[0006] Obtain a real-time video data stream of a complex occlusion video scene;

[0007] Input the real-time video data stream into a skeleton key point extraction model to obtain extraction result data;

[0008] Input the extracted result data into the behavior detection model to output the detection result of the fall behavior;

[0009] Among them, the skeletal key point extraction model is constructed based on the improved YOLOv11-pose model, and the improved YOLOv11-pose model incorporates a Weighted Dynamic Fusion Function (WDFM) module; the behavior detection model is constructed based on the T-STGCN (Transformer enhanced Spatial-Temporal Graph Convolutional Network) model; the T-STGCN model incorporates a Transformer encoder on the basis of the STGCN model.

[0010] In a possible implementation manner, inputting the extracted result data into the behavior detection model to output the detection result of the fall behavior includes:

[0011] Input the extracted result data into the behavior detection model to output a fall label and the corresponding first confidence level, a normal label and the corresponding second confidence level;

[0012] Determine the detection result of the fall behavior according to the label corresponding to the larger value among the first confidence level and the second confidence level; among them, the fall label corresponds to a fall; the normal label corresponds to no fall.

[0013] In a possible implementation manner, the inputting the real-time video data stream into the skeletal key point extraction model includes:

[0014] Extract frames from the real-time video data stream according to a set interval and / or key image features to obtain frame extraction data;

[0015] Input the frame extraction data into the skeletal key point extraction model.

[0016] In a possible implementation manner, the extracting frames from the real-time video data stream according to a set interval and / or key image features includes:

[0017] Extract frames from the real-time video data stream according to a set interval; where the set interval is a set frame rate interval or a set time interval; or,

[0018] Extract frames from the real-time video data stream according to the target object in the image.

[0019] Second aspect, an embodiment of the present invention provides a training method for a skeletal key point extraction model. The YOLOv11-pose model includes a Backbone module, a Neck module, and a Head module; a concat function module is embedded in the Neck module of the YOLOv11-pose model; the concat function module in the Neck module of the improved YOLOv11-pose model is modified to a WDFM module; the training method includes:

[0020] Obtain the original video data containing complex occlusion scenarios and extract frames from the original video data;

[0021] Input the frame-extracted video data into the YOLOv11-pose model to extract a preset number of skeletal key points, generating an initial key point data set; wherein, each video corresponds to a set of continuous skeletal key point sequences; each frame of video data includes some or all of the corresponding preset number of skeletal key points; the preset skeletal key points include 17 skeletal key points;

[0022] Add black occlusion color blocks to the original video data and perform data augmentation processing to generate an enhanced video data set; wherein, the data augmentation processing includes one or more of scaling, translation, and brightness adjustment;

[0023] Synchronously perform the same scaling or translation transformation on the key point data in the initial key point data set corresponding to the enhanced video data set to generate a matching key point data set;

[0024] Train the improved YOLOv11-pose model based on the enhanced video data set and the matching key point data set to obtain a skeletal key point extraction model.

[0025] In a possible implementation manner, the generating of the initial key point data set includes:

[0026] Determine the confidence of the skeletal key points extracted by the YOLOv11-pose model;

[0027] Select the key points whose confidence meets the set conditions to construct the initial key point data set.

[0028] Third aspect, an embodiment of the present invention provides a training method for a behavior detection model. The STGCN model includes a BN module, a plurality of ST-Conv modules, and an Out module. On the basis of the STGCN model, a Transformer encoder is added to each ST-Conv module in the T-STGCN model; the training method includes:

[0029] Obtain the original video data containing complex occlusion scenarios and extract frames from the original video data;

[0030] Input the frame-extracted video data into the YOLOv11-pose model to extract multiple preset skeletal key points, and generate an initial key point data set; wherein, each video corresponds to a set of continuous skeletal key point sequences; each frame of video data includes some or all of the corresponding preset multiple skeletal key points.

[0031] Perform the same translation, scaling, or flipping transformation on the key point data in the initial key point data set to obtain an augmented training set.

[0032] Randomly reset the set number of key points corresponding to each frame of video data in the augmented training set to (0, 0) to obtain a reset data set.

[0033] Train the T-STGCN model based on the reset data set to obtain a behavior detection model.

[0034] In a possible implementation, the set number is 2 to 5; the number of corresponding preset skeletal key points is greater than the set number.

[0035] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method in any possible implementation of the above aspects.

[0036] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the method in any possible implementation of the above aspects.

[0037] In a sixth aspect, an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method in any possible implementation of the above aspects.

[0038] In the embodiments of the present invention, in this embodiment, by constructing a skeletal key point extraction model based on the improved YOLOv11-pose model and using the incorporated weighted dynamic fusion function module, it can effectively avoid the feature confusion and noise interference caused by the traditional simple splicing or element addition feature fusion method, and improve the extraction accuracy of human skeletal key points under complex occlusion. At the same time, the behavior detection model based on the T-STGCN model incorporates a Transformer encoder, and with its multi-head self-attention mechanism, it can explicitly establish the dependency relationship between long-distance time states, and more effectively capture the long-range temporal sequence information and global context information of the fall behavior in the time dimension, thus solving the problems of easy loss of key features and difficult modeling of temporal features in complex occlusion scenarios, and realizing the accurate detection of fall behavior. Description of the Drawings

[0039] Figure 1 It is the implementation flowchart of the personnel fall detection method provided by the embodiments of the present invention;

[0040] Figure 2 It is the implementation flowchart of the training method of the bone key point extraction model provided by the embodiments of the present invention;

[0041] Figure 3 It is the implementation flowchart of the training method of the behavior detection model provided by the embodiments of the present invention;

[0042] Figure 4 It is the overall network structure diagram of the personnel fall detection method provided by the embodiments of the present invention;

[0043] Figure 5 It is the schematic diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0044] To effectively address the problem of accurate detection of human fall behavior in complex occlusion environments, the present application innovatively proposes a two-stage adaptive fall detection algorithm for complex occlusion environments. This algorithm cleverly integrates an improved YOLOv11-pose object detection network and a redesigned T-STGCN behavior recognition network, and can dynamically capture the motion trajectories and posture changes of the human activity area in complex spaces. Specifically, the model first uses the improved YOLOv11-pose network to efficiently generate the human target detection area and the high-precision human bone key point framework at one time. Subsequently, the extracted key point framework sequence is input into the T-STGCN network, and in the time domain, a deep discriminant analysis of the fall behavior is carried out, thus effectively avoiding the problem of loss of key information caused by the uncertainty of the occlusion position, and significantly improving the robustness and detection accuracy of the algorithm in complex occlusion environments.

[0045] Next, the embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0046] Figure 1 It is the implementation flowchart of the personnel fall detection method provided by the embodiments of the present invention. As Figure 1 shown, it includes the following steps:

[0047] S101, obtain the real-time video data stream of the complex occlusion video scene.

[0048] The execution subject of each embodiment of the present application can be a device with data processing functions such as a server, a processor, and a microprocessor. In the actual implementation process, the specific implementation manner of the execution subject can be selected according to actual needs, and this embodiment does not make special restrictions on this, as long as it is a device with data processing functions.

[0049] Among them, in industrial scenarios, complex occlusion video scenarios are areas such as factory workshops, warehousing and logistics areas, and equipment operation areas, with dynamic or static occlusions such as mechanical equipment occlusion, personnel cross-occlusion, and cargo stacking occlusion.

[0050] The real-time video data stream of complex occlusion video scenarios is the video recorded by cameras in an indoor environment. The cameras cover key operation areas (such as around machine tools, conveyor belt channels, and aisles between shelves), support fixed perspectives or 360° rotation through a pan-tilt head to ensure no dead-angle monitoring. There are some natural occlusions in these videos, such as desks, chairs, machines, shelves, etc., and contain the pose data of staff, such as various daily activities like sitting, squatting, and walking.

[0051] S102, input the real-time video data stream into the skeletal key point extraction model to obtain the extraction result data; among them, the skeletal key point extraction model is constructed based on the improved YOLOv11-pose model, and the improved YOLOv11-pose model incorporates the WDFM module.

[0052] During the detection process, the skeletal key points include main skeletal parts such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. Through the coordinate combination of the skeletal key points, the spatial position relationship of each joint of the human body is formed. For example, features such as the torso inclination angle (the angle between the shoulder-hip connection line and the vertical direction) and joint speed (the displacement of key points between adjacent frames / time interval) are calculated, directly reflecting the dynamic changes of human actions.

[0053] In the specific implementation process, the skeletal key points include 17 human skeletal key points. The 17 human skeletal key points are respectively the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. In the specific implementation process, according to the specific application scenario, some of the 17 human skeletal key points can be collected. For example: for the working area of employment-disabled persons with lower limb disabilities, in order to improve the extraction efficiency, only the skeletal key points related to the upper limbs such as the head, neck, shoulders, elbows, and wrists can be extracted; for the working area of employment-disabled persons with upper limb disabilities, in order to improve the extraction efficiency, only the main skeletal key points such as the head, neck, shoulders, hips, knees, and ankles can be extracted.

[0054] In complex occlusion scenarios, some key points may be occluded (such as the arm being blocked by a machine), but the unoccluded key points (such as the legs and torso) can still provide key pose information. The improved YOLOv11-pose model incorporating the WDFM module enhances the detection robustness of partially occluded key points, ensuring that even if only 10 - 15 effective key points are extracted, a complete pose skeleton can still be generated for subsequent analysis.

[0055] Among them, the improved YOLOv11-pose model incorporates the WDFM module, which makes up for the confusion caused by traditional feature fusion of simple vector splicing and the noise information in the differential information caused by vector addition operations.

[0056] S103, input the extraction result data into the behavior detection model, and output the detection result of the falling behavior; among them, the behavior detection model is constructed based on the T-STGCN model; the T-STGCN model incorporates a Transformer encoder on the basis of the STGCN model.

[0057] During the behavior detection process, falling behaviors and normal behaviors will be detected. Among them, falling behaviors include standing upright → suddenly falling, losing balance and falling while walking, falling from a height, etc. The core features are a sudden drop in the body's center of gravity, a sudden change in joint angles (such as the knee flexion angle suddenly dropping from 180° to below 90°), and limb extension or curling (such as the arm abducting to support the ground).

[0058] Normal behaviors include standing, walking, squatting, sitting, bending over to work, etc. The features are a gentle change in joint angles and a stable body center of gravity. For example, when squatting, the knee joint angle gradually decreases and the center of gravity drops evenly.

[0059] In this embodiment, in view of the deficiencies of the STGCN network in dealing with complex occlusions and long-term temporal dependency modeling, the T-STGCN model is proposed. This model integrates a Transformer encoder and uses the multi-head self-attention mechanism of the Transformer to explicitly construct the dependencies between distant temporal states, thereby more effectively capturing the long-range temporal information and global context information of falling behaviors in the time dimension. Specifically, a multi-head attention mechanism is introduced into the graph convolution module, enabling the model to adaptively learn and effectively associate the human skeleton key points in the occluded part during the training process, realizing context reasoning and feature repair in the occluded area, and further improving the behavior recognition accuracy of the model in complex occlusion scenarios.

[0060] The behavior detection model determines the behavior of a person by synthesizing spatial features and temporal features.

[0061] Spatial features are the relative positions of key points within a single frame (such as the ratio of shoulder width to hip width) and joint angles (calculated from three-point coordinates, such as the angle formed by the shoulder-elbow-wrist), which are used to judge the human body posture, such as whether it is in an upright, bent or fallen state.

[0062] Temporal features are the displacement speed (such as the moving distance of the head coordinates within 2 frames / time interval) and acceleration (rate of change of speed) of key points in consecutive frames, which are used to identify the intensity of actions (such as the head acceleration significantly exceeding the threshold when falling).

[0063] When some key points are occluded, the T-STGCN model uses the multi-head attention mechanism to infer the potential features of the occluded area by leveraging the motion trends of adjacent key points and historical frame information, thus achieving feature repair. For example, if the wrist is blocked by a device, by using the motion trends of adjacent key points such as the elbow and shoulder and historical frame information, if it is inferred that the elbow is moving rapidly downward, it is speculated that the wrist may touch the ground.

[0064] In this embodiment, by constructing a skeletal key point extraction model based on the improved YOLOv11-pose model and using the incorporated weighted dynamic fusion function module, it is possible to effectively avoid feature confusion and noise interference caused by traditional simple splicing or element-wise addition feature fusion methods, and improve the extraction accuracy of human skeletal key points under complex occlusions. At the same time, the behavior detection model based on the T-STGCN model incorporates a Transformer encoder. With its multi-head self-attention mechanism, it can explicitly establish dependencies between distant temporal states, more effectively capture the long-range temporal information and global context information of the fall behavior in the time dimension, thus solving the problems of easy loss of key features and difficult modeling of temporal features in complex occlusion scenarios, and achieving accurate detection of fall behavior, providing reliable technical support for safety management in complex environments such as factories.

[0065] In a possible implementation, the extraction result data is input into the behavior detection model, and the detection result of the fall behavior is output, including:

[0066] The extraction result data is input into the behavior detection model, and a fall label and the corresponding first confidence level, a normal label and the corresponding second confidence level are output;

[0067] The detection result of the fall behavior is determined according to the label corresponding to the larger value among the first confidence level and the second confidence level; where the fall label corresponds to a fall; the normal label corresponds to no fall.

[0068] During the fall behavior detection process, if the detection result is directly output based on a single label, the reliability of the detection result may be insufficient due to the uncertainty of the model prediction.

[0069] In a specific embodiment, let the fall label be , and the normal label be

[0070] The fall label indicates that there is a person fall behavior in the current video segment, and the confidence level , and the higher the value, the more certain the model's prediction of "fall"; the normal label indicates that the person's behavior in the current video segment is normal activities such as standing, walking, squatting, etc., and the confidence level , and .

[0071] Direct comparison and of the numerical values, and select the label with a higher confidence as the final detection result:

[0072] If , it is determined as "fall", and the alarm mechanism is triggered;

[0073] If , it is determined as "normal", and the alarm mechanism is not triggered.

[0074] When determining the detection result of the fall behavior based on the label corresponding to the larger value between the first confidence level and the second confidence level, a flexible decision is made based on the "certainty" of the two labels by the model. For example, when , , although the model tends to "fall", the confidence level difference is small, and there may be a risk of misjudgment. At this time, the manual review process can be further started to prompt the management staff to verify in time to prevent unnecessary alarms caused by model misjudgment; and when , , the high confidence level difference indicates that the model's judgment of "fall" is more reliable, thus reducing misjudgment under noise interference, and there is no need for manual review intervention. Among them, a flexible decision is implemented by relying on the fall label and the normal label output by the behavior detection model and their corresponding confidence levels. This method not only reduces the dependence on a single label compared with directly determining the detection result based on the label and triggering the alarm mechanism, but also pays more attention to the "certainty degree" of the model's prediction of the two labels, thereby improving the accuracy of detection and reducing misjudgment.

[0075] In this embodiment, by outputting the fall label and the corresponding first confidence level, the normal label and the corresponding second confidence level through the behavior detection model, and determining the detection result based on the label corresponding to the larger value of the two, the prediction credibility of the model for different labels can be fully utilized. When the prediction confidence level of the model for a certain label is significantly higher than that of another label, it is determined that the label is the final detection result, effectively reducing the risk of misjudgment caused by fuzzy model prediction or noise interference, improving the reliability and accuracy of the detection result, and ensuring more robust judgment of whether a person falls in a complex occlusion scenario.

[0076] In a possible implementation manner, inputting the real-time video data stream into the skeleton key point extraction model includes:

[0077] Performing frame extraction on the real-time video data stream according to the set interval and / or the key features of the image to obtain the frame-extracted data;

[0078] Inputting the frame-extracted data into the skeleton key point extraction model.

[0079] When processing the real-time video data stream in the complex occlusion scenario of the factory, directly extracting the skeletal key points for all video frames will cause waste of computing resources due to the large amount of data, and affect the real-time performance of detection. In a possible implementation, key features of the image are used to extract frames from the real-time video data stream, which can effectively avoid extracting skeletal key points for irrelevant frames (such as frames without target objects or frames with unchanged content), thereby saving computing resources and improving the real-time performance of detection.

[0080] Optionally, frames are extracted from the real-time video data stream according to the target object in the image. Among them, the target object is a worker or an industrial robot.

[0081] In addition, when processing the real-time video data stream in the complex occlusion scenario of the factory, random frame extraction may lead to a decrease in detection accuracy due to missing frames containing key changes in human actions. In a possible implementation, frames are extracted from the real-time video data stream at a set interval. Optionally, the set interval is a set frame rate interval or a set time interval. Among them, the set frame rate interval is determined according to the basic frame rate of the original video. The set time interval ensures that the frame extraction is evenly distributed in the time series.

[0082] In the processing of real-time video data stream, a single frame extraction method is difficult to meet the requirements of complex scenarios. Only extracting frames according to the set interval may extract a large number of invalid frames when the target object does not appear or is in a static state, wasting computing resources; only extracting frames according to the target object in the image, if the motion state of the target object is not considered, key action details may be lost due to insufficient frame extraction frequency when the target is moving violently. In other possible implementations, a frame extraction strategy combining the set interval and key features of the image is adopted. Optionally, frames are extracted from the real-time video data stream according to the set interval and / or key features of the image, including: extracting frames from the real-time video data stream according to the set interval and the target object in the image; among them, the set interval is a set frame rate interval or a set time interval.

[0083] First, a frame extraction rule is preset according to the regular change frequency of the video content, such as at a fixed frame rate interval (such as extracting 1 frame every 2 frames from an original video of 30fps to form frame extraction data of 10fps) or a fixed time interval (such as extracting 1 frame every 3 seconds), to ensure that the frame extraction is evenly distributed in the time series and avoid computational redundancy caused by too dense intervals or broken action features caused by too sparse intervals.

[0084] Secondly, the image recognition algorithm is used to detect in real time whether there is a human target object in the video frame. When the target object is detected, the frame extraction operation is triggered. If the target object is in a moving state (such as walking, bending down, etc.), the frame extraction frequency is further dynamically adjusted according to its moving speed and direction. When the moving speed of the target object is fast or the direction changes suddenly (such as the rapid dumping during a fall), the frame extraction interval is automatically shortened (such as from 10 fps to 20 fps) to ensure that the detailed changes of key actions are captured; when the target object is stationary or moving smoothly, the preset frame extraction interval is restored to reduce the processing of invalid frames.

[0085] For example, in the area of the factory conveyor belt, when the worker is in a stationary operation state, the system extracts 1 frame every 2 seconds for processing; when the worker starts walking and the target object movement is detected and approaches the machine tool, the frame extraction frequency is automatically adjusted to 2 frames per second to capture the key point changes of the walking posture; if the worker suddenly falls during walking (the moving speed and direction change suddenly), the frame extraction frequency is temporarily increased to 5 frames per second to ensure that the key features such as the sudden change of joint angles and the displacement of the center of gravity at the moment of falling are completely captured. This frame extraction strategy not only ensures the continuity of the timing characteristics through the set interval, but also realizes the adaptive response to the dynamic scene through the key feature detection, reduces the computational load while avoiding missing key frames, provides efficient and complete input data for the bone key point extraction model, and ensures the real-time performance and accuracy of subsequent fall detection. Through this frame extraction strategy, the system can effectively balance the utilization of computing resources and the detection accuracy in complex occlusion scenarios, avoid missed detection or misjudgment caused by improper frame processing strategies, and provide a reliable front-end data processing solution for industrial safety monitoring.

[0086] In this implementation method, the real-time video data stream is frame-extracted according to the set interval and the key features of the image. First, frame extraction is performed through the set interval, which can reduce the number of video frames to be processed on the basis of ensuring uniform distribution of the video frame time series, and improve the detection efficiency. Secondly, frame extraction is combined with the key features of the image, which can focus on the video frames containing key information and avoid the invalid processing of irrelevant frames. The combination of the two takes into account both the real-time performance of the detection and the effective capture of key frames, thus improving the accuracy of extracting bone key points and detecting fall behaviors.

[0087] In this embodiment, two frame extraction methods are provided. On the one hand, by setting the frame rate interval or time interval for frame extraction, video frames can be extracted in a fixed pattern, ensuring the stability and temporal uniformity of frame extraction, which is suitable for scenarios where the video content changes regularly. On the other hand, frame extraction is triggered when a target object is detected in the image, and the frame extraction frequency is dynamically adjusted according to the movement speed and direction of the target object. Key attention can be paid when the target object appears. Especially when the target moves fast and the direction changes greatly (such as during a fall), the frame extraction frequency is increased to ensure that the detailed changes of key actions are captured and missed detections are avoided. The two methods are flexibly combined, which not only improves the pertinence and efficiency of frame extraction but also enhances the adaptability to complex dynamic occlusion scenarios.

[0088] Figure 2 It is the implementation flowchart of the training method of the skeleton key point extraction model provided by the embodiment of the present invention. As Figure 2 shown, the training method includes the following steps:

[0089] S201, obtain the original video data containing complex occlusion scenarios, and perform frame extraction on the original video data.

[0090] First, introduce the differences between the YOLOv11-pose model and the improved YOLOv11-pose model. The YOLOv11-pose model includes a Backbone module, a Neck module, and a Head module; a concat function module is embedded in the Neck module of the YOLOv11-pose model; the concat function module in the Neck module of the improved YOLOv11-pose model is modified to a WDFM module.

[0091] When training the skeleton key point extraction model for complex occlusion scenarios, it is necessary to construct an adapted training dataset for the original video data. Specifically, starting from the original video data containing natural occlusions (such as factory equipment, cargo stacks) and simulated fall actions, the continuous video stream is converted into a single-frame image sequence through frame extraction operations.

[0092] S202, input the frame-extracted video data into the YOLOv11-pose model to extract a preset number of skeleton key points, and generate an initial key point dataset; among them, each video corresponds to a set of continuous skeleton key point sequences; each frame of video data includes some or all of the corresponding preset skeleton key points; the preset skeleton key points include 17 skeleton key points.

[0093] When constructing the initial key point dataset, the frame-extracted video data is input into the YOLOv11-pose model. This model presets 17 skeletal key points based on the human body structure (including key joint points such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles), and outputs the two-dimensional coordinates and confidence levels of the corresponding key points for each frame of video image. Due to the fact that in complex occlusion scenarios, the human body may be partially occluded by equipment, goods, or other people, or some joint points may be out of the camera's field of view due to the pose angle, usually only some or all of the preset 17 key points can be extracted from each frame of video data. For example, when a worker faces the camera, the front head and shoulder key points can be effectively extracted, while the key points on the back (such as the thoracic vertebra point) may not be detected; if the worker's arm is blocked by a machine tool, the wrist and elbow key points may be lost, and only the key points above the shoulders and the legs are retained.

[0094] The continuous sequence of skeletal key points generated for each video is essentially the arrangement of the key point coordinates of each frame in chronological order, forming a dynamic human body pose trajectory. Even if some key points are missing in a single frame (such as only 10 valid key points are detected), the combination of multiple consecutive frames can still reflect the overall trend of the human body movement. For example, the continuous displacement of the leg key points during walking can infer the gait, and the angle change of the trunk key points during bending can identify the action type.

[0095] S203, Add black occlusion color blocks to the original video data and perform data augmentation processing to generate an enhanced video dataset; among them, the data augmentation processing includes one or more of scaling, translation, and brightness adjustment.

[0096] During the training process of the skeletal key point extraction model, the natural occlusion scenarios in the original video data are limited, and if the traditional data augmentation method does not synchronously process the key point data, it will cause the video frames and the key point data to be mismatched, affecting the model training effect. To enhance the model's adaptability to occlusion scenarios, the original video data is subjected to data augmentation processing. By artificially adding black occlusion color blocks in the video frames to simulate equipment occlusion, and at the same time performing operations such as scaling, translation, and brightness adjustment to expand data diversity.

[0097] S204, Synchronously perform the same scaling or translation transformation on the key point data in the initial key point dataset corresponding to the enhanced video dataset to generate a matching key point dataset.

[0098] To ensure the consistency between video frames and key point data, the same scaling or translation transformation is synchronously performed on the key point data corresponding to the enhanced video dataset in the initial key point dataset, avoiding coordinate misalignment caused by video enhancement, and generating a matching key point dataset that does not require re-annotation. This process is automatically implemented by an algorithm. For example, when a certain frame of video is processed by translating it 50 pixels to the right, the x coordinate of all key points in this frame is synchronously increased by 50 to ensure that the bone position is strictly aligned with the image content.

[0099] S205. Train the improved YOLOv11-pose model based on the enhanced video dataset and the matching key point dataset to obtain a bone key point extraction model.

[0100] In the embodiment of this application, training the improved YOLOv11-pose model based on the enhanced video dataset and the matching key point dataset can enhance the robustness of YOLOv11-pose for key point detection in the case of occlusion.

[0101] In terms of model improvement, aiming at the confusion problem caused by simply splicing feature maps in the Neck module of the traditional YOLOv11-pose model, the concat function module is replaced with the WDFM module. This module generates dynamic weights through learnable parameters The specific weight formula is as follows:

[0102]

[0103] Among them, σ represents the Sigmoid activation function, whose role is to constrain the weight value within the interval [0,1] to ensure the effectiveness and stability of the weight value. In the initial state of the model, the parameter , corresponding to the weights w1 = w2 = 0.5, to avoid the gradient anomaly problem caused by excessive weight initialization deviation in the initial stage of training and ensure the smoothness of model training.

[0104] Different from the traditional method of directly splicing or adding feature maps, WDFM adaptively balances the rich features of channel splicing and the detail retention of element addition through the following weighted fusion formula.

[0105]

[0106] Among them, represents the two input feature maps, , , respectively represent the number of channels, height, and width of the feature map; represents the channel splicing operation; represents the element-wise addition operation; It represents a convolution operation. Secondly, the WDFM module can be easily embedded into the existing feature fusion network without significantly adjusting the network structure. And compared with the traditional attention mechanism, the WDFM module can achieve efficient feature selection and fusion with extremely small number of parameters.

[0107] For example, when processing a human body image occluded by a device, a higher splicing weight is adopted for the low-level feature map containing contour information to preserve the complete structure, and a higher addition weight is adopted for the high-level feature map containing texture details to highlight local features, avoiding information loss or noise introduction caused by a single fusion method.

[0108] In this embodiment, by obtaining the original video data containing complex occlusion scenarios, extracting frames and extracting skeletal key points to generate an initial key point dataset, then adding black occlusion color blocks to the original video data and performing data augmentation such as scaling, translation, brightness and contrast adjustment, and at the same time, performing the same scaling or translation transformation on the corresponding key point data in the initial key point dataset synchronously to generate a matching key point dataset. This synchronous processing method ensures the consistency between the enhanced video frames and the key point data, avoids the cumbersome work of manual re-annotation, and at the same time, by simulating diverse occlusion scenarios and data augmentation, expands the diversity of training data, enabling the improved YOLOv11-pose model to learn the key point features in different occlusion situations during the training process, effectively improving the detection robustness of the model for human body skeletal key points in complex occlusion scenarios, enabling it to more accurately extract partially or fully occluded skeletal key points, and providing reliable input data for subsequent behavior detection.

[0109] In a possible implementation manner, generating the initial key point dataset includes:

[0110] Determining the confidence of the skeletal key points extracted by the YOLOv11-pose model;

[0111] Selecting the key points whose confidence meets the set conditions to construct the initial key point dataset.

[0112] When generating the initial key point dataset, the skeletal key points extracted by the YOLOv11-pose model may have low confidence due to factors such as occlusion and lighting. If these unreliable key points are directly included in the dataset, it will introduce noise and affect the model training effect. The annotation errors are corrected through the manual verification step, and the key points with confidence higher than the threshold (such as 0.5) are screened to ensure the reliability of the initial dataset. For the key points that are not detected, the coordinates are default set to (0,0) and the confidence is set to 0, which not only preserves the data integrity but also provides a basis for subsequent data augmentation (such as simulating occlusion).

[0113] In this embodiment, by determining the confidence of the skeletal key points extracted by the YOLOv11-pose model and selecting the key points whose confidence meets the set conditions to construct the initial key point dataset, low-confidence noisy key points can be filtered out, and the key point information with higher reliability can be retained. This operation avoids interference with the training process caused by incorrect or unreliable key point data, ensures the quality of the initial key point dataset, so that the improved YOLOv11-pose model trained based on this dataset can focus more on learning the features of high-confidence key points, further improving the accuracy and reliability of the model in extracting skeletal key points in complex occlusion scenarios, and providing more accurate basic data for the fall detection system.

[0114] Figure 3 It is the implementation flowchart of the training method of the behavior detection model provided by the embodiment of the present invention. As Figure 3 shown, the training method includes the following steps:

[0115] S301, obtain the original video data containing complex occlusion scenarios and perform frame extraction on the original video data.

[0116] First, introduce the differences between the STGCN model and the T-STGCN model. The STGCN model includes a BN module, multiple ST-Conv modules, and an Out module. On the basis of the STGCN model, the T-STGCN model adds a Transformer encoder to each ST-Conv module.

[0117] When training a skeletal key point extraction model for complex occlusion scenarios, an adapted training dataset needs to be constructed for the original video data. Specifically, starting from the original video data containing natural occlusions (such as factory equipment, cargo stacks) and simulated fall actions, the continuous video stream is converted into a single-frame image sequence through frame extraction.

[0118] S302, input the frame-extracted video data into the YOLOv11-pose model to extract a preset number of skeletal key points, and generate an initial key point dataset; among them, each video corresponds to a set of continuous skeletal key point sequences; each frame of video data includes some or all of the corresponding preset skeletal key points.

[0119] Input the video data after frame extraction into the YOLOv11-pose model, and use this model to extract 17 preset skeletal key points to generate an initial key point dataset. In this dataset, each video corresponds to a set of continuous skeletal key point sequences, and each frame of video data contains some or all of the 17 preset skeletal key points. For example, in a certain frame, the worker's arm is blocked by the machine, so in this frame, only some of the key points other than the key points related to the arm may be obtained, while in another frame without occlusion, all 17 key points can be obtained.

[0120] S303. Perform the same translation, scaling, or flipping transformation on the key point data in the initial key point dataset to obtain an augmented training set.

[0121] Perform augmentation processing on the key point data in the initial key point dataset. By performing the same translation, scaling, or flipping transformation on the key points corresponding to each frame of video data, the diversity of training data is increased, enabling the model to learn the key point features in different poses and positions. For example, translate all the key points of the entire human body to the right by a certain number of pixels, or scale all the key points by a certain ratio to simulate human body images at different distances.

[0122] S304. Randomly reset a set number of key points corresponding to each frame of video data in the augmented training set to (0, 0) to obtain a reset dataset.

[0123] Randomly reset a set number of key points corresponding to each frame of video data in the augmented training set to (0, 0) to simulate the situation where some skeletal key points are occluded in the actual scenario and obtain a reset dataset. For example, randomly select 2 leg key points and 1 hand key point in a certain frame and set their coordinates to (0, 0), indicating that these key points are occluded.

[0124] S305. Train the T-STGCN model based on the reset dataset to obtain a behavior detection model.

[0125] In the training of the behavior detection model, when dealing with complex occlusion scenarios, the STGCN model has insufficient ability to model the temporal features in the case of missing or occluded key points, and it is easy to cause incorrect recognition of fall behaviors due to the loss of some key point information.

[0126] In terms of model improvement, the T-STGCN model is optimized based on the STGCN model. Add a Transformer encoder to each ST-ConvBlock module of the STGCN model to enhance the context connection of the model, enabling the model to better learn the mutual correlation between image actions. It can be expressed by the following formula:

[0127]

[0128]

[0129] Among them, represents the output of the previous layer, represents the input of the current layer. At the same time, for the enhancement of the TGC model, the multi-head attention mechanism is added, and the formula is as follows:

[0130]

[0131] Among them, represents the multi-head attention mechanism, enabling the model to adaptively learn and effectively associate the occluded human skeleton key points during the training process, achieving context reasoning and feature repair in the occluded area, and further improving the behavior recognition accuracy of the model in complex occlusion scenarios. The Transformer encoder uses the multi-head self-attention mechanism to explicitly establish the dependencies between distant temporal states.

[0132] When the input resets the dataset for training, the model can better capture the long-range temporal information of the fall behavior in the time dimension. Even if some key points are occluded, it can infer the action state of the current frame based on the unoccluded key point information in the previous and subsequent frames. At the same time, the multi-head attention mechanism is introduced into the graph convolution module, and through the formula, context reasoning and feature repair in the occluded area are realized. For example, during the training process, when the knee key point of a person in a certain frame is occluded, the model can use the multi-head attention mechanism to learn the association between these key points based on other unoccluded key points (such as the hip and ankle) around and the information of the previous and subsequent frames, and then infer the possible position and state of the occluded knee, achieving feature repair.

[0133] In this embodiment, the augmented training set is obtained by performing translation, scaling, or flipping transformations on the key-point data in the initial key-point dataset, which expands the diversity of training data and improves the generalization ability of the model. Further, a set number of key points corresponding to each frame of video data in the augmented training set are randomly reset to (0, 0) to simulate the situation where some skeletal key points are occluded in the actual scenario, generating a reset dataset. Training the T-STGCN model based on this reset dataset enables the model to learn how to utilize the temporal features and context information of the unoccluded key points to make fall behavior judgments under the condition of missing key points. At the same time, the T-STGCN model adds a Transformer encoder in each ST-Conv module, enhancing the ability to model long-range temporal dependencies, thus effectively solving the problem of difficult feature inference caused by missing key points in complex occlusion scenarios, improving the model's context reasoning and feature repair capabilities for occluded regions, and ultimately enhancing the recognition accuracy of fall behavior temporal features in scenarios with missing key points.

[0134] In a possible implementation, the set number is 2 to 5; the number of corresponding preset skeletal key points is greater than the set number.

[0135] During the training process of the behavior detection model, if too many key points are reset (the set number is too large), it may cause the model to be unable to effectively learn fall features due to missing key information; if the set number is too small, it is difficult to fully simulate the situation of missing key points in the actual occlusion scenario and cannot achieve the purpose of improving the model's robustness.

[0136] In this embodiment, the set number is limited to 2 to 5, and the number of preset skeletal key points is greater than this set number, ensuring that there are still enough unoccluded key points to provide effective feature information for the model after resetting some key points. This setting not only simulates the common situation where some key points are occluded in the actual complex occlusion scenario (such as 2 to 5 key points are lost due to occlusion), but also avoids the model being unable to capture basic action features due to excessive key point resetting, enabling the model to focus on learning to utilize the temporal correlation and context information of unoccluded key points for fall behavior reasoning within a reasonable range of missing key points. Thus, on the premise of ensuring the learnability of the model, it effectively improves its adaptability to the common situation of missing key points in the actual occlusion scenario and enhances the robustness of fall behavior recognition in complex occlusion environments.

[0137] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0138] Figure 4It is a schematic diagram of the overall network structure of the personnel fall detection method provided by the embodiments of the present invention. As Figure 4 shown, the system input end receives the real-time video data stream of the complex occlusion video scene. The data stream first passes through the improved YOLOv11-pose model to extract the human skeleton key point framework, and then the extracted human skeleton key point framework sequence is input into the T-STGCN model to perform in-depth discriminant analysis on the fall behavior in the time domain. Finally, the system output end outputs whether there is a falling action in the video stream. Among them, the WDFM module is embedded in the Neck module of the improved YOLOv11-pose model, and the Transformer encoder is added to each ST-Conv module in the T-STGCN model.

[0139] Figure 5 It is a schematic diagram of the electronic device provided by the embodiments of the present invention. As Figure 5 shown, the electronic device 5 of this embodiment includes: a processor 50 and a memory 51. The memory 51 stores a computer program 52. When the processor 50 executes the computer program 52, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 50 executes the computer program 52, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0140] Exemplarily, the computer program 52 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 51 and executed by the processor 50 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 52 in the electronic device 5.

[0141] The electronic device 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 merely examples of the electronic device 5 do not constitute a limitation on the electronic device 5, and may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device 5 may also include input and output devices, network access devices, buses, etc.

[0142] The processor 50 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0143] The memory 51 may be an internal storage unit of the electronic device 5, such as the hard disk or memory of the electronic device 5. The memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 5. Further, the memory 51 may also include both the internal storage unit and the external storage device of the electronic device 5. The memory 51 is used to store the computer program 52 and other programs and data required by the electronic device 5. The memory 51 may also be used to temporarily store the data that has been output or is to be output.

[0144] For the convenience and simplicity of description, only the above division of each functional module / unit is used as an example. In practical applications, the above functions may be assigned to different functional modules / units according to needs. The above modules / units may be implemented in the form of hardware, or may be implemented in the form of software, or may be implemented in the form of a combination of hardware and software.

[0145] The embodiment of the present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.

[0146] The embodiment of the present invention also provides a computer program product including a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.

[0147] Among them, the computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0148] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Without special instructions and logical conflicts, the terms and / or descriptions between different embodiments are consistent and can be cross-referenced. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0149] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the respective embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for detecting human falls, characterized in that, Including: Obtain the real-time video data stream of the complex occlusion video scene; Input the real-time video data stream into the skeletal key point extraction model to obtain the extraction result data; Input the extraction result data into the behavior detection model to output the detection result of the fall behavior; Among them, the bone key point extraction model is constructed based on the improved YOLOv11-pose model; the improved YOLOv11-pose model incorporates a weighted dynamic fusion function WDFM module, including: through learnable parameters to generate dynamic weights; the weight formula is as follows: Where, σ represents the Sigmoid activation function, and its function is to constrain the weight value within the interval [0, 1]; Perform feature fusion through the following weighted fusion formula: Among them, represents two input feature maps, , , respectively represent the number of channels, height, and width of the feature map; represents the channel concatenation operation; represents the element-wise addition operation; represents the convolution operation; The behavior detection model is constructed based on the T-STGCN model; the T-STGCN model incorporates a Transformer encoder on the basis of the STGCN model, including: Among them, represents the output of the previous layer, represents the input of the current layer; is a simple graph convolution; for TGC enhancement of the model, a multi-head attention mechanism is added, and the formula is as follows: Among them, represents the multi-head attention mechanism, represents the activation function, represents the convolution operation.

2. The personnel fall detection method according to claim 1, characterized in that Input the extraction result data into the behavior detection model to output the detection result of the fall behavior, including: Input the extraction result data into the behavior detection model to output the fall label and the corresponding first confidence level, the normal label and the corresponding second confidence level; Determine the detection result of the fall behavior according to the label corresponding to the larger value among the first confidence level and the second confidence level; where, the fall label corresponds to a fall; the normal label corresponds to no fall.

3. The method for detecting a person's fall according to claim 1, characterized in that, The inputting the real-time video data stream into the skeletal key point extraction model includes: Perform frame extraction on the real-time video data stream according to the set interval and / or the key features of the image to obtain the frame-extracted data; Input the frame-extracted data into the skeletal key point extraction model.

4. The personnel fall detection method according to claim 3, wherein The performing frame extraction on the real-time video data stream according to the set interval and / or the key features of the image includes: Perform frame extraction on the real-time video data stream according to the set interval; where, the set interval is the set frame rate interval or the set time interval; or, Perform frame extraction on the real-time video data stream according to the target object in the image.

5. The personnel fall detection method according to claim 1, wherein, The YOLOv11-pose model includes a Backbone module, a Neck module, and a Head module; a concat function module is embedded in the Neck module of the YOLOv11-pose model; the concat function module in the Neck module of the improved YOLOv11-pose model is modified to a WDFM module; The training process of the skeletal key point extraction model includes: Obtain the original video data containing complex occlusion scenes and perform frame extraction on the original video data; Input the frame-extracted video data into the YOLOv11-pose model to extract a preset number of skeletal key points to generate an initial key point data set; where, each video corresponds to a set of continuous skeletal key point sequences; each frame of video data includes some or all of the corresponding preset number of skeletal key points; Add black occlusion color blocks to the original video data and perform data augmentation processing to generate an augmented video data set; where, the data augmentation processing includes one or more of scaling, translation, and brightness adjustment; Perform the same scaling or translation transformation on the key point data in the initial key point data set corresponding to the augmented video data set to generate a matching key point data set; Train the improved YOLOv11-pose model based on the augmented video data set and the matching key point data set to obtain the skeletal key point extraction model.

6. The training method of the skeletal key point extraction model according to claim 5, characterized in that The generation of the initial key point dataset includes: Determine the confidence of the skeletal key points extracted by the YOLOv11-pose model; Select the key points whose confidence meets the set conditions to construct the initial key point dataset.

7. The method for detecting human fall according to claim 1, characterized in that, The STGCN model includes a BN module, multiple ST-Conv modules, and an Out module. Based on the STGCN model, the T-STGCN model adds a Transformer encoder to each ST-Conv module; The training process of the behavior detection model includes: Obtain the original video data containing complex occlusion scenarios and extract frames from the original video data; Input the frame-extracted video data into the YOLOv11-pose model to extract a preset number of skeletal key points, generating an initial key point dataset; wherein, each video corresponds to a set of continuous skeletal key point sequences; each frame of video data includes some or all of the corresponding preset skeletal key points; Perform the same translation, scaling, or flipping transformation on the key point data in the initial key point dataset to obtain an augmented training set; Randomly reset the set number of key points corresponding to each frame of video data in the augmented training set to (0, 0) to obtain a reset dataset; Train the T-STGCN model based on the reset dataset to obtain a behavior detection model.

8. The training method of the behavior detection model according to claim 7, wherein The set number is 2 to 5; the number of corresponding preset skeletal key points is greater than the set number.

9. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method according to any one of claims 1 to 8.

10. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Fall detection model training improvement method and fall detection method

    CN117671794A

  • Human body posture estimation method based on dynamic graph convolutional network

    CN119445672A

Cited By

  • Systems and methods for detecting fall events

    US12620232B2