Personnel fall detection method, model training method, equipment and computer program

Through the improved YOLOv11-pose model and T-STGCN model, combined with the weighted dynamic fusion function and Transformer encoder, the accuracy problem of human fall behavior detection in complex occlusion environments is solved, and efficient and robust fall behavior detection is achieved.

CN120048005AActive Publication Date: 2025-05-27HEBEI BAISHA TOBACCO

Patent Information

Application Number
CN202510533832.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In complex occlusion environments, it is difficult for the existing technology to achieve accurate detection of human fall behavior. Traditional methods have low detection reliability under dynamic occlusion and uneven light problems, and high misjudgment and missed detection rates.

Method used

The improved YOLOv11-pose model is used to combine the weighted dynamic fusion function module for bone key points extraction, and the T-STGCN model is integrated into the Transformer encoder for behavior detection. The dependence between long-distance time states is established through the multi-head self-attention mechanism, and the long-range timing information and global context information of fall behavior are captured.

Benefits of technology

It significantly improves the accuracy of extracting key points of human bones and the accuracy of behavior detection in complex occlusion environments, reduces the rate of misjudgment and missed detection, and improves the robustness and real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048005A_ABST
    Figure CN120048005A_ABST
Patent Text Reader

Abstract

The invention provides a personnel falling detection method, a model training method, equipment and a computer program, and relates to the technical field of image recognition processing. The method comprises the following steps: acquiring a real-time video data stream of a complex shielding video scene; inputting the real-time video data stream into the skeleton key point extraction model to obtain extraction result data; inputting the extraction result data into a behavior detection model, and outputting a detection result of the falling behavior; wherein the skeleton key point extraction model is constructed on the basis of an improved YOLOv11-pose model, and the improved YOLOv11-pose model is fused into a weighted dynamic fusion function (WDFM) module; the behavior detection model is constructed based on a T-STGCN model; according to the T-STGCN model, a Transform encoder is integrated on the basis of an STGCN model. According to the method, the problems that key features are easy to lose and time sequence features are difficult to model in a complex shielding scene can be solved, and accurate detection of a falling behavior is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition processing technology, and in particular to a person fall detection method, a model training method, a device and a computer program. Background Art

[0002] In key areas such as industrial production, medical health monitoring and public safety, real-time and accurate human fall detection technology is of vital practical significance for reducing the risk of accidental injuries and improving the efficiency of safety emergency response. Especially in complex industrial environments such as mechanical workshops, warehousing and logistics centers, dynamic occlusion (such as occlusion by operating equipment and occlusion by multi-person interaction) and uneven lighting are common, which seriously restricts the reliability of traditional detection methods and highlights the urgent need for highly robust detection algorithms in complex environments.

[0003] With the rapid development of artificial intelligence technology, related research methods have gradually penetrated into the field of fall detection. At present, the mainstream research methods are mainly divided into two categories: one is the threshold judgment method based on wearable sensors. This method usually relies on sensors such as accelerometers and gyroscopes to collect human motion data and judge the fall status by setting thresholds. Although this method can avoid the problem of visual occlusion to a certain extent, its inherent defects such as strong equipment invasiveness and low efficiency of multi-node data coordination, as well as practical application difficulties such as equipment update and maintenance, and forgetting to wear, limit the widespread deployment of wearable devices in complex industrial scenarios. The second is a computer vision algorithm based on a monocular camera. This type of method usually uses deep learning models such as OpenPose and YOLO to extract key points of human bones and posture features (such as trunk inclination, joint speed, etc.), and then realizes the recognition and classification of fall behavior. However, the above-mentioned traditional computer vision-based methods still face significant limitations in complex occlusion environments: sensor-dependent solutions are difficult to effectively distinguish between real fall states and high-intensity work action states, which can easily lead to misjudgment; and traditional visual algorithms are prone to losing key feature information when the target human body is partially occluded, resulting in an increased false detection rate and even the risk of missed detection, which seriously affects detection accuracy and reliability. Summary of the invention

[0004] The embodiments of the present invention provide a person fall detection method, a model training method, a device and a computer program to solve the problem of accurate detection of human fall behavior in a complex occlusion environment.

[0005] In a first aspect, an embodiment of the present invention provides a method for detecting a person falling, comprising: Obtain real-time video data streams of complex occlusion video scenes; Inputting the real-time video data stream into a skeleton key point extraction model to obtain extraction result data; Inputting the extracted result data into a behavior detection model, and outputting a detection result of the falling behavior; Among them, the skeleton key point extraction model is constructed based on the improved YOLOv11-pose model, and the improved YOLOv11-pose model is integrated with the weighted dynamic fusion function (WDFM) module; the behavior detection model is constructed based on the T-STGCN (Transformer enhanced Spatial-Temporal GraphConvolutional Network) model; the T-STGCN model is integrated with the Transformer encoder based on the STGCN model.

[0006] In a possible implementation, the extraction result data is input into a behavior detection model, and the detection result of the fall behavior is output, including: Input the extracted result data into a behavior detection model, and output a fall label and a corresponding first confidence, a normal label and a corresponding second confidence; The detection result of the fall behavior is determined according to the label corresponding to the larger value of the first confidence level and the second confidence level; wherein the fall label corresponds to a fall; and the normal label corresponds to no fall.

[0007] In a possible implementation, inputting the real-time video data stream into a skeleton key point extraction model includes: Extract frames from the real-time video data stream according to a set interval and / or key image features to obtain extracted frame data; The frame data is input into the skeleton key point extraction model.

[0008] In a possible implementation, the extracting frames of the real-time video data stream according to the set interval and / or the key image features includes: Extract frames from the real-time video data stream according to a set interval; wherein the set interval is a set frame rate interval or a set time interval; or, Extract frames from real-time video data stream based on target objects in the image.

[0009] In a second aspect, an embodiment of the present invention provides a training method for a skeleton key point extraction model, wherein the YOLOv11-pose model includes a Backbone module, a Neck module and a Head module; a concat function module is embedded in the Neck module of the YOLOv11-pose model; the concat function module in the Neck module of the improved YOLOv11-pose model is modified to a WDFM module; the training method includes: Obtaining original video data containing complex occlusion scenes, and extracting frames from the original video data; The frame-extracted video data is input into the YOLOv11-pose model to extract multiple preset skeleton key points and generate an initial key point data set; wherein each video segment corresponds to a set of continuous skeleton key point sequences; each frame of video data includes part or all of the corresponding preset multiple skeleton key points; the preset skeleton key points include 17 skeleton key points; Adding black mask blocks to the original video data and performing data enhancement processing to generate an enhanced video data set; wherein the data enhancement processing includes one or more of scaling, translation, and brightness adjustment; Synchronously performing the same scaling or translation transformation on the key point data in the initial key point data set corresponding to the enhanced video data set to generate a matching key point data set; The improved YOLOv11-pose model is trained based on the enhanced video data set and the matching key point data set to obtain a skeleton key point extraction model.

[0010] In a possible implementation, generating an initial key point data set includes: Determine the confidence of the skeleton key points extracted by the YOLOv11-pose model; The key points whose confidence meets the set conditions are selected to construct the initial key point dataset.

[0011] In a third aspect, an embodiment of the present invention provides a training method for a behavior detection model, wherein the STGCN model includes a BN module, multiple ST-Conv modules and an Out module, and the T-STGCN model is based on the STGCN model, and a Transformer encoder is added to each ST-Conv module; the training method includes: Acquire the original video data containing the complex occlusion scene and extract frames from the original video data; The frame-extracted video data is input into the YOLOv11-pose model to extract multiple preset skeleton key points and generate an initial key point data set; wherein each video segment corresponds to a set of continuous skeleton key point sequences; each frame of video data includes part or all of the corresponding preset multiple skeleton key points; Performing the same translation, scaling or flipping transformation on the key point data in the initial key point data set to obtain an expanded training set; Randomly resetting a set number of key points corresponding to each frame of video data in the expanded training set to (0, 0) to obtain a reset data set; The T-STGCN model is trained based on the reset data set to obtain a behavior detection model.

[0012] In a possible implementation, the set number is 2-5; the number corresponding to the preset skeleton key points is greater than the set number.

[0013] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method in any possible implementation manner in the above aspects is implemented.

[0014] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method in any possible implementation manner in the above aspects is implemented.

[0015] In a sixth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements a method in any possible implementation manner in the above aspects.

[0016] In the embodiment of the present invention, in this embodiment, by constructing a skeleton key point extraction model based on the improved YOLOv11-pose model and using the integrated weighted dynamic fusion function module, it is possible to effectively avoid the feature confusion and noise interference caused by the traditional feature fusion method of simple splicing or element addition, and improve the extraction accuracy of human skeleton key points under complex occlusion. At the same time, the behavior detection model based on the T-STGCN model is integrated into the Transformer encoder, and with the help of its multi-head self-attention mechanism, the dependency relationship between long-distance time states can be explicitly established, and the long-range temporal information and global context information of the fall behavior in the time dimension can be more effectively captured, thereby solving the problem that key features are easily lost and temporal features are difficult to model in complex occlusion scenes, and accurate detection of fall behavior is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flow chart of an implementation method of a person fall detection method provided by an embodiment of the present invention; Figure 2 It is a flow chart of the implementation of the training method of the skeleton key point extraction model provided by the embodiment of the present invention; Figure 3 is a flow chart of an implementation of a method for training a behavior detection model provided by an embodiment of the present invention; Figure 4 is a schematic diagram of the overall network structure of the method for detecting a person falling down provided by an embodiment of the present invention; Figure 5 is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to effectively deal with the problem of accurate detection of human fall behavior in complex occluded environments, this application innovatively proposes a two-stage adaptive fall detection algorithm for complex occluded environments. The algorithm cleverly combines the improved YOLOv11-pose target detection network with the redesigned T-STGCN behavior recognition network, and can dynamically capture the motion trajectory and posture changes of human activity areas in complex spaces. Specifically, the model first adopts the improved YOLOv11-pose network, which enables it to efficiently generate human target detection areas and high-precision human skeleton key point frameworks at one time. Subsequently, the extracted key point frame sequence is input into the T-STGCN network, and the fall behavior is deeply discriminated and analyzed in the time domain, thereby effectively avoiding the problem of key information loss caused by the uncertainty of the occlusion position, and significantly improving the robustness and detection accuracy of the algorithm in complex occluded environments.

[0019] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] Figure 1 is a flow chart of the implementation of the method for detecting a person falling down provided by an embodiment of the present invention, such as Figure 1 As shown, the following steps are included: S101, obtaining a real-time video data stream of a complex occlusion video scene.

[0021] The execution subjects of each embodiment of the present application can be servers, processors, microprocessors and other devices with data processing functions. In the actual implementation process, the specific implementation method of the execution subject can be selected according to actual needs. This embodiment does not impose any special restrictions on this, as long as it is a device with data processing functions.

[0022] Among them, in industrial scenes, complex occlusion video scenes include factory workshops, warehousing and logistics areas, equipment operation areas and other areas, where there are dynamic or static occlusions such as mechanical equipment occlusion, personnel crossing occlusion, and cargo stacking occlusion.

[0023] Real-time video data streams of complex occlusion video scenes are recorded by cameras in indoor environments. The cameras cover key operating areas (such as machine tool perimeters, conveyor belt channels, and aisles between shelves), and support fixed viewing angles or 360° rotation through the pan / tilt to ensure blind spot monitoring. These videos contain some natural occlusions, such as tables, chairs, machines, shelves, etc., and contain posture data of workers, such as sitting, squatting, walking, and other daily activities.

[0024] S102, inputting the real-time video data stream into the skeleton key point extraction model to obtain extraction result data; wherein, the skeleton key point extraction model is constructed based on the improved YOLOv11-pose model, and the improved YOLOv11-pose model is integrated into the WDFM module.

[0025] During the detection process, the key bone points include the head, neck, shoulder, elbow, wrist, hip, knee, ankle and other major bone parts. The spatial position relationship of each joint of the human body is formed by the coordinate combination of the key bone points. For example, the trunk inclination angle (the angle between the shoulder-hip line and the vertical direction) and joint velocity (displacement / time interval of key points in adjacent frames) are calculated to directly reflect the dynamic changes of human body movements.

[0026] In the specific implementation process, the skeleton key points include 17 human skeleton key points. The 17 human skeleton key points are nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle. In the specific implementation process, according to the specific application scenario, some of the 17 human skeleton key points can be collected. For example, for the working area of ​​employees with lower limb disabilities, in order to improve the extraction efficiency, only the upper limb related skeleton key points such as head, neck, shoulders, elbows, wrists, hips can be extracted; for the working area of ​​employees with upper limb disabilities, in order to improve the extraction efficiency, only the main skeleton key points such as head, neck, shoulders, hips, knees, ankles can be extracted.

[0027] In complex occlusion scenarios, some key points may be occluded (such as arms blocked by machines), but unoccluded key points (such as legs and torso) can still provide key pose information. The improved YOLOv11-pose model incorporates the WDFM module to enhance the robustness of detection of partially occluded key points, ensuring that even if only 10-15 valid key points are extracted, a complete pose skeleton can still be generated for subsequent analysis.

[0028] Among them, the improved YOLOv11-pose model is integrated into the WDFM module to make up for the confusion caused by the traditional feature fusion of simple vector splicing and the noise information in the difference information caused by the vector addition operation.

[0029] S103, input the extracted result data into the behavior detection model, and output the detection result of the fall behavior; wherein, the behavior detection model is constructed based on the T-STGCN model; the T-STGCN model incorporates the Transformer encoder based on the STGCN model.

[0030] During the behavior detection process, falling behaviors and normal behaviors will be detected. Among them, falling behaviors include standing up → suddenly falling, falling down due to imbalance while walking, falling from a height, etc. The core features are sudden drop in the body's center of gravity, sudden change in joint angles (such as knee flexion angle suddenly dropping from 180° to below 90°), and limb extension or curling (such as arms abducted to support the ground).

[0031] Normal behaviors include standing, walking, squatting, sitting, bending over, etc., and are characterized by smooth changes in joint angles and a stable center of gravity. For example, when squatting, the knee joint angle gradually decreases and the center of gravity drops at a uniform rate.

[0032] In this embodiment, a T-STGCN model is proposed to address the shortcomings of the STGCN network in dealing with complex occlusions and long-term temporal dependency modeling. This model integrates the Transformer encoder and uses the Transformer's multi-head self-attention mechanism to explicitly construct dependencies between long-distance temporal states, thereby more effectively capturing the long-range temporal information and global contextual information of the falling behavior in the time dimension. Specifically, a multi-head attention mechanism is introduced in the graph convolution module, so that the model can adaptively learn and effectively associate the key points of the human skeleton in the occluded part during training, realize contextual reasoning and feature repair of the occluded area, and further improve the model's behavior recognition accuracy in complex occlusion scenarios.

[0033] The behavior detection model combines spatial and temporal features to determine the behavior of people.

[0034] Spatial features refer to the relative positions of key points in a single frame (such as the ratio of shoulder width to hip width) and joint angles (calculated through three-point coordinates, such as the angle formed by the shoulder-elbow-wrist), which are used to determine the posture of the human body, such as whether it is standing upright, bending over or falling down.

[0035] Temporal features refer to the displacement speed of key points in consecutive frames (such as the moving distance / time interval of the head coordinates within 2 frames) and acceleration (rate of change of velocity), which are used to identify the intensity of the action (such as the head acceleration significantly exceeding the threshold during a fall).

[0036] When some key points are occluded, the T-STGCN model uses a multi-head attention mechanism and the motion trends and historical frame information of adjacent key points to infer the potential features of the occluded area and achieve feature restoration. For example, if the wrist is blocked by the equipment, the motion trends and historical frame information of adjacent key points such as the elbow and shoulder can be used to infer that the elbow moves downward quickly, and it is speculated that the wrist may touch the ground.

[0037] In this embodiment, by constructing a skeleton key point extraction model based on the improved YOLOv11-pose model and using the integrated weighted dynamic fusion function module, it is possible to effectively avoid feature confusion and noise interference caused by the traditional feature fusion method of simple splicing or element addition, and improve the extraction accuracy of human skeleton key points under complex occlusion. At the same time, the behavior detection model based on the T-STGCN model is integrated into the Transformer encoder. With the help of its multi-head self-attention mechanism, the dependency relationship between long-distance time states can be explicitly established, and the long-range time series information and global context information of the fall behavior in the time dimension can be more effectively captured, thereby solving the problem of easy loss of key features and difficult modeling of time series features in complex occlusion scenes, realizing accurate detection of fall behavior, and providing reliable technical support for safety management in complex environments such as factories.

[0038] In a possible implementation, the extracted result data is input into a behavior detection model, and the detection result of the fall behavior is output, including: Input the extracted result data into the behavior detection model, and output the fall label and the corresponding first confidence, the normal label and the corresponding second confidence; The detection result of the fall behavior is determined according to the label corresponding to the larger value of the first confidence and the second confidence; wherein the fall label corresponds to falling; and the normal label corresponds to not falling.

[0039] In the process of fall behavior detection, if the detection results are directly output based on a single label, the uncertainty of the model prediction may lead to insufficient reliability of the detection results.

[0040] In a specific embodiment, the fall label is , the normal label is Fall Tags Indicates that there is a person falling in the current video clip, and the confidence level , the higher the value, the more certain the model is about the prediction of "fall"; normal label Indicates that the person's behavior in the current video clip is normal, such as standing, walking, squatting, etc. ,and .

[0041] Direct comparison and The value of , selects the label with higher confidence as the final detection result: like , determined as a "fall", triggering the alarm mechanism; like , judged as "normal", and no alarm mechanism is triggered.

[0042] When determining the detection result of the fall behavior according to the label corresponding to the larger value of the first confidence and the second confidence, a flexible decision is made on the "confidence" of the two labels based on the model. , When the model tends to "fall", the confidence difference is small, and there may be a risk of misjudgment. At this time, the manual review process can be further initiated to prompt the management personnel to verify in time to prevent unnecessary alarms due to model misjudgment; and when , , a high confidence difference indicates that the model's judgment of "fall" is more reliable, thereby reducing misjudgments under noise interference, and no manual review intervention is required. Among them, flexible decision-making is implemented based on the fall label and normal label output by the behavior detection model and their corresponding confidence. Compared with directly determining the detection result based on the label and triggering the alarm mechanism, this method not only reduces the dependence on a single label, but also pays more attention to the model's "confidence" in predicting the two labels, thereby improving the accuracy of detection and reducing misjudgments.

[0043] In this embodiment, the behavior detection model outputs the fall label and the corresponding first confidence, the normal label and the corresponding second confidence, and determines the detection result according to the label corresponding to the larger value of the two. This can make full use of the model's prediction credibility for different labels. When the model's prediction confidence for a certain label is significantly higher than that for another label, the label is determined to be the final detection result, which effectively reduces the risk of misjudgment caused by model prediction ambiguity or noise interference, improves the reliability and accuracy of the detection results, and ensures a more robust judgment of whether a person has fallen in complex occlusion scenarios.

[0044] In a possible implementation, the real-time video data stream is input into a skeleton key point extraction model, including: Extract frames from a real-time video data stream according to a set interval and / or key image features to obtain extracted frame data; Input the extracted frame data into the skeleton key point extraction model.

[0045] When processing real-time video data streams of complex occlusion scenes in factories, directly extracting skeleton key points from all video frames will waste computing resources due to the large amount of data and affect the real-time detection. In one possible implementation, the real-time video data stream is extracted using image key features, which can effectively avoid extracting skeleton key points from irrelevant frames (such as frames that do not contain target objects or whose content has not changed), thereby saving computing resources and improving the real-time detection.

[0046] Optionally, the real-time video data stream is frame extracted according to a target object in the image, wherein the target object is a worker or an industrial robot.

[0047] In addition, when processing real-time video data streams of complex occlusion scenes in factories, irregular frame extraction may cause a decrease in detection accuracy due to the omission of frames containing key changes in personnel movements. In a possible implementation, a set interval is used to extract frames from the real-time video data stream. Optionally, the set interval is a set frame rate interval or a set time interval. The set frame rate interval is determined according to the basic frame rate of the original video. The set time interval ensures that the extracted frames are evenly distributed in the time series.

[0048] In the processing of real-time video data stream, a single frame extraction method is difficult to adapt to the needs of complex scenes. If the frame is extracted only according to the set interval, a large number of invalid frames may be extracted when the target object does not appear or is in a static state, wasting computing resources; if the frame is extracted only according to the target object in the image, if the motion state of the target object is not considered, key action details may be lost due to insufficient frame extraction frequency when the target moves violently. In other possible implementations, a frame extraction strategy combining a set interval with key image features is adopted. Optionally, the real-time video data stream is extracted according to the set interval and / or the key image features, including: extracting the real-time video data stream according to the set interval and the target object in the image; wherein the set interval is a set frame rate interval or a set time interval.

[0049] First, preset frame extraction rules according to the regular change frequency of the video content, such as fixed frame rate intervals (such as extracting 1 frame every 2 frames from the original 30fps video to form 10fps frame extraction data) or fixed time intervals (such as extracting 1 frame every 3 seconds) to ensure that the frame extraction is evenly distributed in the time series to avoid computational redundancy due to too dense intervals or interruption of motion features due to too sparse intervals.

[0050] Secondly, the image recognition algorithm is used to detect in real time whether there is a human target object in the video frame. When the target object is detected, the frame extraction operation is triggered. If the target object is in motion (such as walking, bending, etc.), the frame extraction frequency is further dynamically adjusted according to its movement speed and direction. When the target object moves faster or changes direction suddenly (such as falling quickly when falling), the frame extraction interval is automatically shortened (such as from 10fps to 20fps) to ensure that the details of key actions are captured; when the target object is stationary or moving slowly, the preset frame extraction interval is restored to reduce invalid frame processing.

[0051] For example, in the conveyor belt area of ​​the factory, when the worker is in a stationary operating state, the system extracts 1 frame every 2 seconds for processing; when the worker starts walking and detects the movement of the target object and approaches the machine tool, the frame extraction frequency is automatically adjusted to 2 frames per second to capture the key point changes of his walking posture; if the worker suddenly falls during walking (sudden change in movement speed and direction), the frame extraction frequency is temporarily increased to 5 frames per second to ensure that key features such as joint angle mutation and center of gravity displacement at the moment of falling are fully captured. This frame extraction strategy not only ensures the continuity of timing features by setting intervals, but also realizes adaptive response to dynamic scenes through key feature detection, reduces the computational load while avoiding key frame omission, and provides efficient and complete input data for the skeleton key point extraction model, ensuring the real-time and accuracy of subsequent fall detection. Through this frame extraction strategy, the system can effectively balance the utilization of computing resources and detection accuracy in complex occlusion scenes, avoid missed detection or misjudgment caused by improper frame processing strategy, and provide a reliable front-end data processing solution for industrial safety monitoring.

[0052] In this implementation, the real-time video data stream is framed according to the set interval and key image features. First, by setting the interval frame, the number of video frames to be processed can be reduced while ensuring the uniform distribution of the video frame time series, thereby improving the detection efficiency. Secondly, by combining the key image feature frame extraction, it is possible to focus on the video frames containing key information and avoid invalid processing of irrelevant frames. The combination of the two takes into account the real-time detection and ensures the effective capture of key frames, thereby improving the accuracy of extracting skeletal key points and detecting falling behaviors.

[0053] In this embodiment, two frame extraction methods are provided. On the one hand, by setting the frame rate interval or time interval frame extraction, video frames can be extracted in a fixed pattern to ensure the stability and timing uniformity of frame extraction, which is suitable for scenes with relatively regular changes in video content. On the other hand, frame extraction is triggered when a target object is detected in the image, and the frame extraction frequency is dynamically adjusted according to the movement speed and direction of the target object. When the target object appears, it can be focused on, especially when the target moves fast and changes direction greatly (such as falling process), the frame extraction frequency is increased to ensure that the details of key actions are captured and to avoid missed detection. The two methods are flexibly combined, which not only improves the pertinence and efficiency of frame extraction, but also enhances the adaptability to complex dynamic occlusion scenes.

[0054] Figure 2 is a flow chart of the implementation of the training method of the skeleton key point extraction model provided by the embodiment of the present invention, such as Figure 2 As shown, the training method includes the following steps: S201, obtaining original video data containing a complex occlusion scene, and extracting frames from the original video data.

[0055] First, the difference between the YOLOv11-pose model and the improved YOLOv11-pose model is introduced. The YOLOv11-pose model includes the Backbone module, the Neck module, and the Head module; the concat function module is embedded in the Neck module of the YOLOv11-pose model; the concat function module in the Neck module of the improved YOLOv11-pose model is modified to the WDFM module.

[0056] When training a skeleton key point extraction model for complex occlusion scenes, it is necessary to build an adapted training dataset for the original video data. Specifically, starting from the original video data containing natural occlusions (such as factory equipment, cargo stacking) and simulated falling actions, the continuous video stream is converted into a single-frame image sequence through frame extraction.

[0057] S202, input the frame-extracted video data into the YOLOv11-pose model to extract multiple preset skeletal key points and generate an initial key point data set; wherein each video segment corresponds to a set of continuous skeletal key point sequences; each frame of video data includes part or all of the corresponding preset multiple skeletal key points; the preset skeletal key points include 17 skeletal key points.

[0058] When constructing the initial key point data set, the extracted video data is input into the YOLOv11-pose model, which presets 17 skeletal key points (including key joints such as head, neck, shoulder, elbow, wrist, hip, knee, ankle, etc.) based on the human body structure, and outputs the two-dimensional coordinates and confidence of the corresponding key points for each frame of video image. Since the human body may be partially blocked by equipment, goods or other personnel in complex occlusion scenes, or some joints may be beyond the camera's field of view due to posture angles, each frame of video data can usually only extract part or all of the preset 17 key points. For example, when a worker faces the camera, the key points of the head and shoulders on the front can be effectively extracted, while the key points on the back (such as the thoracic vertebrae) may not be detected; if the worker's arms are blocked by the machine tool, the key points of the wrist and elbow may be lost, and only the key points above the shoulder and the legs are retained.

[0059] The continuous sequence of skeletal key points generated for each video is actually the arrangement of the key point coordinates of each frame in chronological order, forming a dynamic human posture trajectory. Even if some key points are missing in a single frame (for example, only 10 valid key points are detected), the combination of multiple consecutive frames can still reflect the overall trend of human movements. For example, the continuous displacement of the key points of the legs when walking can infer the gait, and the angle change of the key points of the torso when bending can identify the type of action.

[0060] S203, adding black mask blocks to the original video data, and performing data enhancement processing to generate an enhanced video data set; wherein the data enhancement processing includes one or more of scaling, translation, and brightness adjustment.

[0061] In the training process of the skeleton key point extraction model, there are limited natural occlusion scenes in the original video data, and if the traditional data enhancement method does not synchronously process the key point data, it will cause the video frame and the key point data to not match, affecting the model training effect. In order to enhance the model's adaptability to occlusion scenes, the original video data is enhanced by artificially adding black occlusion blocks to the video frame to simulate device occlusion, and at the same time performing operations such as scaling, translation, and brightness adjustment to expand data diversity.

[0062] S204, synchronously performing the same scaling or translation transformation on the key point data in the initial key point data set corresponding to the enhanced video data set to generate a matching key point data set.

[0063] To ensure the consistency between the video frames and the key point data, the same scaling or translation transformation is synchronously performed on the key point data in the initial key point data set that corresponds to the enhanced video data set to avoid coordinate misalignment caused by video enhancement and generate a matching key point data set that does not require re-labeling. This process is automatically implemented through the algorithm. For example, when a frame of video is translated to the right by 50 pixels, the x-coordinates of all key points in the frame are synchronously increased by 50 to ensure that the skeleton position is strictly aligned with the image content.

[0064] S205, training an improved YOLOv11-pose model based on the enhanced video dataset and the matching key point dataset to obtain a skeleton key point extraction model.

[0065] The embodiment of the present application trains an improved YOLOv11-pose model based on an enhanced video dataset and a matching key point dataset, which can enhance the robustness of YOLOv11-pose for key point detection in the case of occlusion.

[0066] In terms of model improvement, the concat function module is replaced with the WDFM module to solve the confusion problem caused by the simple concatenation of feature maps in the Neck module of the traditional YOLOv11-pose model. Generate dynamic weights. The specific weight formula is as follows:

[0067] Among them, σ represents the Sigmoid activation function, which constrains the weight value in the interval [0,1] to ensure the effectiveness and stability of the weight value. , corresponding to the weight w1=w2=0.5, to avoid the gradient abnormality problem caused by excessive weight initialization deviation in the early stage of training and ensure the stability of model training.

[0068] Different from the traditional method of directly concatenating or adding feature maps, WDFM adaptively balances the rich features of channel concatenation and the detail retention of element addition through the following weighted fusion formula.

[0069]

[0070] in, Represents two feature maps of the input, , , Represent the number of channels, height and width of the feature map respectively; Indicates channel splicing operation; Represents element-wise addition operation; Represents the convolution operation. Secondly, the WDFM module can be easily embedded into the existing feature fusion network without major adjustments to the network structure. Compared with the traditional attention mechanism, the WDFM module can achieve efficient feature selection and fusion with a very small number of parameters.

[0071] For example, when processing human body images occluded by the device, a higher splicing weight is used for the low-level feature map containing contour information to retain the complete structure, and a higher addition weight is used for the high-level feature map containing texture details to highlight local features, avoiding information loss or noise introduction caused by a single fusion method.

[0072] In this embodiment, by acquiring the original video data containing complex occlusion scenes, extracting frames and extracting skeleton key points to generate an initial key point data set, then adding black occlusion blocks to the original video data and performing data enhancements such as scaling, translation, and light and dark adjustment, and synchronously performing the same scaling or translation transformation on the corresponding key point data in the initial key point data set, a matching key point data set is generated. This synchronous processing method ensures the consistency of the enhanced video frame and the key point data, avoids the tedious work of manual re-labeling, and at the same time, by simulating a variety of occlusion scenes and data enhancement, the diversity of training data is expanded, so that the improved YOLOv11-pose model can learn the key point features under different occlusion conditions during the training process, effectively improving the model's detection robustness of human skeleton key points in complex occlusion scenes, so that it can more accurately extract some or all of the occluded skeleton key points, and provide reliable input data for subsequent behavior detection.

[0073] In a possible implementation, generating an initial key point data set includes: Determine the confidence of the skeleton key points extracted by the YOLOv11-pose model; The key points whose confidence meets the set conditions are selected to construct the initial key point dataset.

[0074] When generating the initial key point data set, the skeleton key points extracted by the YOLOv11-pose model may have low confidence due to factors such as occlusion and lighting. If these unreliable key points are directly included in the data set, noise will be introduced, affecting the model training effect. Correct the annotation errors through manual verification steps, and filter the key points with confidence higher than the threshold (such as 0.5) to ensure the reliability of the initial data set. For undetected key points, the coordinates are set to (0,0) by default, and the confidence is set to 0, which not only preserves the integrity of the data, but also provides a basis for subsequent data enhancement (such as simulated occlusion).

[0075] In this embodiment, by determining the confidence of the skeleton key points extracted by the YOLOv11-pose model and selecting the key points whose confidence meets the set conditions to construct the initial key point data set, it is possible to filter out the noise key points with low confidence and retain the key point information with higher reliability. This operation avoids interference with the training process due to erroneous or unreliable key point data, ensures the quality of the initial key point data set, so that the improved YOLOv11-pose model trained based on the data set can focus more on learning the features of high-confidence key points, further improve the accuracy and reliability of the model in extracting skeleton key points in complex occlusion scenes, and provide more accurate basic data for the fall detection system.

[0076] Figure 3 is a flow chart of the implementation of the training method of the behavior detection model provided by the embodiment of the present invention, such as Figure 3 As shown, the training method includes the following steps: S301, obtaining existing original video data containing a complex occlusion scene, and extracting frames from the original video data.

[0077] First, we introduce the difference between the STGCN model and the T-STGCN model. The STGCN model includes a BN module, multiple ST-Conv modules, and an Out module. The T-STGCN model is based on the STGCN model, and adds a Transformer encoder to each ST-Conv module.

[0078] When training a skeleton key point extraction model for complex occlusion scenes, it is necessary to build an adapted training dataset for the original video data. Specifically, starting from the original video data containing natural occlusions (such as factory equipment, cargo stacking) and simulated falling actions, the continuous video stream is converted into a single-frame image sequence through frame extraction.

[0079] S302, input the frame-extracted video data into the YOLOv11-pose model to extract multiple preset skeletal key points and generate an initial key point data set; wherein each video segment corresponds to a set of continuous skeletal key point sequences; each frame of video data includes part or all of the corresponding preset multiple skeletal key points.

[0080] The extracted video data is input into the YOLOv11-pose model, and the model is used to extract the preset 17 skeletal key points to generate an initial key point data set. In this data set, each video corresponds to a set of continuous skeletal key point sequences, and each frame of video data contains part or all of the preset 17 skeletal key points. For example, in a certain frame, if the worker's arm is blocked by the machine, then this frame may only obtain some key points other than the key points related to the arm, while in another frame, if there is no blockage, all 17 key points can be obtained.

[0081] S303, performing the same translation, scaling or flipping transformation on the key point data in the initial key point data set to obtain an expanded training set.

[0082] The key point data in the initial key point data set is expanded. By performing the same translation, scaling or flipping transformation on the key points corresponding to each frame of video data, the diversity of training data is increased, allowing the model to learn the key point features in different postures and positions. For example, the key points of the entire human body are translated to the right by a certain number of pixels, or all key points are scaled according to a certain ratio to simulate human images at different distances.

[0083] S304 , randomly resetting a set number of key points corresponding to each frame of video data in the expanded training set to (0, 0) to obtain a reset data set.

[0084] Randomly reset a set number of key points corresponding to each frame of video data in the expanded training set to (0,0) to simulate the situation where some skeleton key points are blocked in the actual scene, and obtain the reset data set. For example, randomly select 2 leg key points and 1 hand key point in a frame, and set their coordinates to (0,0), indicating that these key points are blocked.

[0085] S305, training the T-STGCN model based on the reset data set to obtain a behavior detection model.

[0086] In the training of the behavior detection model, the STGCN model has insufficient ability to model the temporal features when key points are missing or occluded when dealing with complex occlusion scenes, which can easily lead to errors in the recognition of falling behaviors due to the loss of some key point information.

[0087] In terms of model improvement, the T-STGCN model is optimized based on the STGCN model. A Transformer encoder is added to each ST-ConvBlock module of the STGCN model to enhance the contextual connection of the model, which enables the model to better learn the correlation between image actions. It can be expressed as follows:

[0088]

[0089] in, represents the output of the previous layer, Represents the input of the current layer. TGC The model is enhanced by adding a multi-head attention mechanism. The formula is as follows:

[0090] in, The multi-head self-attention mechanism enables the model to adaptively learn and effectively associate the key points of the occluded human skeleton during training, realizes contextual reasoning and feature restoration of the occluded area, and further improves the model's behavior recognition accuracy in complex occlusion scenarios. The Transformer encoder uses a multi-head self-attention mechanism to explicitly establish dependencies between long-distance temporal states.

[0091] When the reset data set is input for training, the model can better capture the long-range temporal information of the falling behavior in the time dimension. Even if some key points are blocked, the action state of the current frame can be inferred based on the key point information of the previous and next frames. At the same time, the multi-head attention mechanism is introduced in the graph convolution module. The formula realizes contextual reasoning and feature repair of occluded areas. For example, during training, when the key point of a character's knee is occluded in a certain frame, the model can use the multi-head attention mechanism to learn the relationship between these key points based on other unoccluded key points around (such as hips and ankles) and the information of previous and next frames, and then infer the possible position and state of the occluded knee to achieve feature repair.

[0092] In this embodiment, an expanded training set is obtained by translating, scaling or flipping the key point data in the initial key point data set, which expands the diversity of the training data and improves the generalization ability of the model; further, a set number of key points corresponding to each frame of video data in the expanded training set are randomly reset to (0,0) to simulate the situation where some skeletal key points are occluded in the actual scene, and a reset data set is generated. Based on the reset data set, the T-STGCN model is trained, which enables the model to learn how to use the temporal features and contextual information of the unoccluded key points to judge the fall behavior under the condition of missing key points. At the same time, the T-STGCN model adds a Transformer encoder in each ST-Conv module, which enhances the modeling ability of long-range temporal dependencies, thereby effectively solving the problem of feature reasoning difficulties caused by missing key points in complex occlusion scenes, and improving the model's contextual reasoning and feature repair capabilities for occluded areas, and ultimately improving the recognition accuracy of the temporal features of fall behavior in key point missing scenes.

[0093] In a possible implementation, the number is set to 2-5; the number corresponding to the preset skeleton key points is greater than the set number.

[0094] During the training process of the behavior detection model, if too many key points are reset (the number is set too large), the model may not be able to effectively learn the fall characteristics due to the lack of key information; if the number is set too small, it will be difficult to fully simulate the situation of missing key points in the actual occlusion scene, and the purpose of improving the robustness of the model cannot be achieved.

[0095] In this embodiment, the set number is limited to 2 to 5, and the preset number of skeleton key points is greater than the set number, ensuring that after resetting some key points, there are still enough unobstructed key points to provide effective feature information for the model. This setting not only simulates the situation where some key points are occluded (such as 2 to 5 key points are lost due to occlusion) that is common in actual complex occlusion scenes, but also avoids the inability of the model to capture basic action features due to excessive key point resets, so that the model can focus on learning to use the temporal association and contextual information of unobstructed key points to reason about fall behavior within a reasonable range of key point loss, thereby effectively improving its adaptability to common key point loss situations in actual occlusion scenes while ensuring the learnability of the model, and enhancing the robustness of identifying fall behaviors in complex occlusion environments.

[0096] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0097] Figure 4Schematic diagram of the overall network structure of the method for detecting a person falling down provided by an embodiment of the present invention. Figure 4 As shown in the figure, the input of the system receives the real-time video data stream of the complex occluded video scene. The data stream first passes through the improved YOLOv11-pose model to extract the human skeleton key point frame, and then the extracted human skeleton key point frame sequence is input into the T-STGCN model to perform deep discrimination analysis on the fall behavior in the time domain. Finally, the output of the system outputs whether there is a fall action in the video stream. Among them, the Neck module in the improved YOLOv11-pose model is embedded in the WDFM module, and the Transformer encoder is added to each ST-Conv module in the T-STGCN model.

[0098] Figure 5 Schematic diagram of an electronic device provided by an embodiment of the present invention. Figure 5 As shown, the electronic device 5 of this embodiment includes: a processor 50 and a memory 51. The memory 51 stores a computer program 52. When the processor 50 executes the computer program 52, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 50 executes the computer program 52, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0099] Exemplarily, the computer program 52 may be divided into one or more modules / units, which are stored in the memory 51 and executed by the processor 50 to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program 52 in the electronic device 5.

[0100] The electronic device 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will appreciate that Figure 5 It is only an example of the electronic device 5 and does not constitute a limitation of the electronic device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device 5 may also include input and output devices, network access devices, buses, etc.

[0101] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0102] The memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. The memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the electronic device 5. The memory 51 is used to store the computer program 52 and other programs and data required by the electronic device 5. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0103] For the convenience and simplicity of description, only the division of the above functional modules / units is used as an example for illustration. In actual applications, the above functions can be assigned to different functional modules / units as needed. The above modules / units can be implemented in the form of hardware, software, or a combination of hardware and software.

[0104] The embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.

[0105] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.

[0106] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0107] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. If there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form a new embodiment according to their internal logical relationship.

[0108] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for detecting a person falling, characterized in that: include: Obtain real-time video data streams of complex occlusion video scenes; Inputting the real-time video data stream into a skeleton key point extraction model to obtain extraction result data; Inputting the extracted result data into a behavior detection model, and outputting a detection result of the falling behavior; Among them, the skeleton key point extraction model is constructed based on the improved YOLOv11-pose model, and the improved YOLOv11-pose model incorporates the weighted dynamic fusion function WDFM module; the behavior detection model is constructed based on the T-STGCN model; the T-STGCN model incorporates the Transformer encoder based on the STGCN model.

2. The method for detecting a person falling down according to claim 1, characterized in that: The extracted result data is input into the behavior detection model, and the detection result of the fall behavior is output, including: Input the extracted result data into a behavior detection model, and output a fall label and a corresponding first confidence, a normal label and a corresponding second confidence; The detection result of the fall behavior is determined according to the label corresponding to the larger value of the first confidence level and the second confidence level; wherein the fall label corresponds to a fall; and the normal label corresponds to no fall.

3. The method for detecting a person falling down according to claim 1, characterized in that: The step of inputting the real-time video data stream into a skeleton key point extraction model comprises: Extract frames from the real-time video data stream according to a set interval and / or key image features to obtain extracted frame data; The frame data is input into the skeleton key point extraction model.

4. The method for detecting a person falling down according to claim 3, characterized in that: The step of extracting frames from the real-time video data stream according to a set interval and / or key image features includes: Extract frames from the real-time video data stream according to a set interval; wherein the set interval is a set frame rate interval or a set time interval; or, Extract frames from real-time video data stream based on target objects in the image.

5. A training method for a skeleton key point extraction model, characterized in that: The YOLOv11-pose model includes a Backbone module, a Neck module and a Head module; a concat function module is embedded in the Neck module of the YOLOv11-pose model; the concat function module in the Neck module of the improved YOLOv11-pose model is modified to a WDFM module; the training method includes: Obtaining original video data containing complex occlusion scenes, and extracting frames from the original video data; The frame-extracted video data is input into the YOLOv11-pose model to extract multiple preset skeleton key points and generate an initial key point data set; wherein each video segment corresponds to a set of continuous skeleton key point sequences; each frame of video data includes part or all of the corresponding preset multiple skeleton key points; Adding black mask blocks to the original video data and performing data enhancement processing to generate an enhanced video data set; wherein the data enhancement processing includes one or more of scaling, translation, and brightness adjustment; Synchronously performing the same scaling or translation transformation on the key point data in the initial key point data set corresponding to the enhanced video data set to generate a matching key point data set; The improved YOLOv11-pose model is trained based on the enhanced video data set and the matching key point data set to obtain a skeleton key point extraction model.

6. The training method of the skeleton key point extraction model according to claim 5, characterized in that: The generating of the initial key point data set comprises: Determine the confidence of the skeleton key points extracted by the YOLOv11-pose model; The key points whose confidence meets the set conditions are selected to construct the initial key point dataset.

7. A method for training a behavior detection model, characterized in that: The STGCN model includes a BN module, multiple ST-Conv modules and an Out module. The T-STGCN model is based on the STGCN model, and a Transformer encoder is added to each ST-Conv module. The training method includes: Acquire the original video data containing the complex occlusion scene and extract frames from the original video data; The frame-extracted video data is input into the YOLOv11-pose model to extract multiple preset skeleton key points and generate an initial key point data set; wherein each video segment corresponds to a set of continuous skeleton key point sequences; each frame of video data includes part or all of the corresponding preset multiple skeleton key points; Performing the same translation, scaling or flipping transformation on the key point data in the initial key point data set to obtain an expanded training set; Randomly resetting a set number of key points corresponding to each frame of video data in the expanded training set to (0, 0) to obtain a reset data set; The T-STGCN model is trained based on the reset data set to obtain a behavior detection model.

8. The method for training a behavior detection model according to claim 7, characterized in that: The set number is 2 to 5; the number corresponding to the preset bone key points is greater than the set number.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Multi-angle tumble high-risk identification method and system based on skeleton key points

    CN113496216A

  • Fall detection method and device, electronic equipment and storage medium

    CN114694243A

  • Fall detection model training improvement method and fall detection method

    CN117671794A

  • Human body tumble prediction method and device, electronic equipment and storage medium

    CN118648892A

  • ICU patient body motion video identification method based on human skeleton key point detection

    CN118781654A

Cited By

  • Single-person abnormal behavior identification method and system based on multi-modal skeleton feature fusion

    CN120997901A

  • Unplanned operation risk early warning method based on image recognition and AI algorithm

    CN121280771A

  • Production line personnel behavior identification method and system

    CN121305679A

  • Dynamic human body falling-to-ground detection method based on future human body time sequence posture prediction

    CN121938052A

  • Pedestrian tumble detection method and system based on attitude estimation

    CN121963299A