Lightweight fall recognition method and system based on human skeleton data and three-branch space-time diagram convolutional network
Through lightweight human skeleton key point detection and three-branch spatio-temporal graph convolution network, the accuracy and computing resource problems of existing fall detection technologies are solved, and efficient fall recognition on embedded devices is achieved.
Patent Information
- Application Number
- CN202510581785.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
AI Technical Summary
The existing fall detection technology has problems such as low detection accuracy, limited scope of application, large calculation volume and high resource demand, making it difficult to realize real-time monitoring on embedded terminal devices.
The lightweight human skeleton key point detection network and three-branch spatio-temporal graph convolution network are adopted to optimize bone key point detection through deep separation convolution and self-attention mechanism, and combine lightweight convolution and multi-scale feature extraction of three-branch spatio-temporal graph convolution network to reduce the computational complexity and improve the recognition accuracy.
It realizes efficient and accurate fall recognition on embedded devices, reduces the demand for computing resources, and is suitable for real-time monitoring of terminal devices.
Smart Images

Figure CN120452066A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human posture estimation and action recognition, and relates to a lightweight fall recognition method and system based on human skeleton data and a three-branch spatiotemporal graph convolutional network. Background Art
[0002] As the global population accelerates toward an aging trend, the safety and health of the elderly has become a pressing social issue. According to authoritative statistics, the incidence of falls among the elderly is increasing year by year. The resulting serious consequences, such as fractures and head injuries, not only severely impair the elderly's ability to care for themselves and their quality of life, but also place a heavy burden on their families and the healthcare system. In home-based elderly care scenarios, achieving real-time fall monitoring and timely warnings has become a key technical challenge in ensuring the safety and independence of the elderly.
[0003] Traditional fall detection technology mainly relies on two types of solutions: wearable devices and environmental sensors. The former collects kinematic parameters for behavioral analysis by wearing inertial measurement units such as accelerometers and gyroscopes on the human body. However, it has inherent defects such as discomfort, limited battery life, and high false alarm rate, resulting in poor compliance in long-term use. The latter is based on environmental sensing devices such as pressure sensors and infrared arrays. Although it does not require active cooperation from the user, it is limited by problems such as fixed installation location and narrow monitoring range, making it difficult to adapt to complex and changing daily activity scenarios. What is more noteworthy is that neither of the above two types of technologies can effectively distinguish the subtle differences between normal sitting and lying and accidental falls, resulting in detection accuracy in actual applications that is difficult to meet clinical needs.
[0004] In recent years, computer vision-based behavior recognition technology has provided a new technical path for fall detection. By analyzing the characteristics of human motion in video streams, the visual system can capture complete action sequence information in a non-contact manner, and in theory has higher environmental adaptability and recognition reliability. However, existing visual algorithms generally face two major technical bottlenecks: First, traditional methods rely on raw RGB image processing, which not only poses a risk of privacy leakage, but also causes insufficient stability in feature extraction due to factors such as background interference and lighting changes; second, mainstream deep learning models use dense convolution operations to process high-dimensional visual data. The number of model parameters and computational complexity remain high, making it difficult to meet the real-time requirements and resource constraints of embedded terminal devices. Although some studies have attempted to reduce data dimensions through skeleton extraction, the existing skeleton key point detection network still has problems such as bloated models and inference delays, making it impossible to achieve lightweight deployment.
[0005] At the level of action recognition algorithms, although the spatiotemporal modeling method based on graph convolutional networks (GCN) can effectively characterize the topological relationship and temporal evolution of human joints, the traditional single-branch network architecture has difficulty capturing the multi-scale characteristics of falling movements. Specifically, the falling process involves complex features such as sudden changes in the center of gravity and expansion of the limb contact surface, requiring the algorithm to have both local joint motion tracking and global posture evolution analysis capabilities. Existing single-stream networks often achieve feature abstraction by stacking convolutional layers. This cascade structure easily leads to the loss of shallow fine-grained features, and the spatial receptive field of conventional convolution operations is fixed, making it difficult to adaptively capture key features of different fall stages. In addition, the intensive computational characteristics of the standard convolutional layer in the spatiotemporal graph convolution module further aggravate the computational load of the model, restricting its application prospects in edge computing devices.
[0006] In summary, existing technologies for fall action detection generally have problems such as low detection accuracy, limited scope of application, and high intrusiveness to users, making them difficult to be widely used in practical environments. At the same time, visual algorithms require a large amount of computation and high requirements for computing resources, making them difficult to apply to terminal embedded devices. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a lightweight fall recognition method and system based on human skeleton data and a three-branch spatiotemporal graph convolutional network, to solve the problems of high computational cost and slow operation speed of existing vision-based fall detection solutions, and to achieve a good balance between detection performance and detection speed.
[0008] To achieve the above objectives, the present invention provides a lightweight fall recognition method based on human skeleton data and a three-branch spatiotemporal graph convolutional network, the method comprising:
[0009] First, a lightweight human skeleton key point detection network and a three-branch spatiotemporal graph convolutional network are constructed respectively; wherein, the lightweight human skeleton key point detection network includes a backbone network, a neck network and a head network. The convolutions in the backbone network and the neck network both adopt depthwise separable convolution modules, and the detection head in the head network adopts a self-attention mechanism detection head; the three-branch spatiotemporal graph convolutional network includes multiple cascaded three-branch spatiotemporal graph convolution modules;
[0010] Then, the lightweight human skeleton key point detection network is trained and optimized to establish a human skeleton key point detection model, and the human skeleton key point detection model is used to perform real-time detection and extraction of human skeleton key point data; the three-branch spatiotemporal graph convolutional network is trained and optimized to establish a fall action recognition model, and the fall action recognition result is obtained through the fall action recognition model;
[0011] Finally, the real-time collected human motion image is input into the human skeleton key point detection model to extract the human skeleton key point data, and the human skeleton key point data is input into the falling action recognition model to determine whether the detected human body falls.
[0012] Furthermore, the lightweight human skeleton keypoint detection network is based on the YOLOv8-Pose network. The original convolution modules in the YOLOv8-Pose network backbone and neck network are replaced with depthwise separable convolution modules, and the original detection head in the YOLOv8-Pose network head network is replaced with a self-attention detection head. The self-attention detection head includes a self-attention module and a lightweight convolution module.
[0013] Furthermore, the three-branch spatiotemporal graph convolution network includes ten cascaded three-branch spatiotemporal graph convolution modules, each of which divides the input feature map into three feature subsets through 1×1 convolution, each feature subset has the same spatial size as the input feature map, but the number of channels is one-third of the input feature map; the first feature subset is directly output, the second feature subset is first passed through spatial convolution and then output, and the third feature subset is first passed through temporal convolution and then output, and the outputs of the three feature subsets are residually connected to enhance the multi-scale feature extraction capability; wherein, the second feature subset is first added to the first feature subset and then spatially convolved; the third feature subset is first added to the spatial convolution result and then temporally convolved.
[0014] Moreover, in each of the three-branch spatiotemporal graph convolution modules, the spatial convolution and temporal convolution both use lightweight convolution to reduce the number of parameters and computational complexity; the lightweight convolution process includes: first, converting the number of channels of the input feature map from C1 to C2 / 2 through standard convolution to generate a first feature vector of C2 / 2; then, convolving the first feature vector using depthwise separable convolution to obtain a second feature vector of C2 / 2; then, concatenating the first feature vector and the second feature vector; finally, fusing the features generated by standard convolution with the features generated by depthwise separable convolution through a Shuffle operation, and finally outputting the feature map of the C2 channel.
[0015] Furthermore, when training the lightweight human skeleton key point detection network, WIoU is used as the loss function to optimize the model parameters and accelerate the convergence speed;
[0016] When training the three-branch spatiotemporal graph convolutional network, the momentum stochastic gradient descent method is used to update the weights, Nesterov acceleration is used to speed up the convergence of gradient descent, and the cross entropy function is used as the loss function to optimize the model parameters.
[0017] On the other hand, the present invention provides a lightweight fall recognition system based on human skeleton data and a three-branch spatiotemporal graph convolutional network, which includes a data acquisition module, a human skeleton key point detection module and a fall action recognition module.
[0018] The data acquisition module is used to collect human motion images and perform preprocessing; the human skeleton key point detection module detects and extracts human skeleton key point data through a lightweight human skeleton key point detection network; the fall action recognition module processes the human skeleton key point data through a three-branch spatiotemporal graph convolutional network to obtain a fall action recognition result.
[0019] Among them, the lightweight human skeleton key point detection network includes a backbone network, a neck network and a head network. The convolutions in the backbone network and the neck network both adopt depth-separable convolution modules, and the detection head in the head network adopts a self-attention mechanism detection head; the three-branch spatiotemporal graph convolution network includes multiple cascaded three-branch spatiotemporal graph convolution modules.
[0020] Furthermore, the lightweight human skeleton key point detection network is based on the YOLOv8-Pose network, replacing the original convolution modules in the YOLOv8-Pose network backbone network and neck network with depthwise separable convolution modules, and replacing the original detection head in the YOLOv8-Pose network head network with a self-attention detection head.
[0021] Furthermore, the three-branch spatiotemporal graph convolutional network includes ten cascaded three-branch spatiotemporal graph convolutional modules, each of which divides the input feature map into three feature subsets through 1×1 convolution, each feature subset having the same spatial size as the input feature map but one-third of the number of channels of the input feature map;
[0022] The first feature subset is directly output, the second feature subset is first output after spatial convolution, and the third feature subset is first output after temporal convolution. The outputs of the three feature subsets are residually connected to enhance the multi-scale feature extraction capability;
[0023] The second feature subset is first added to the first feature subset and then spatial convolution is performed; the third feature subset is first added to the spatial convolution result and then temporal convolution is performed.
[0024] In each of the three-branch spatiotemporal graph convolution modules, the spatial convolution and temporal convolution both use lightweight convolution to reduce the number of parameters and computational complexity; the lightweight convolution process includes: first, converting the number of channels of the input feature map from C1 to C2 / 2 through standard convolution to generate a first eigenvector of C2 / 2; then, convolving the first eigenvector using depthwise separable convolution to obtain a second eigenvector of C2 / 2; then, concatenating the first eigenvector and the second eigenvector; finally, fusing the features generated by standard convolution with the features generated by depthwise separable convolution through a shuffle operation, and finally outputting a feature map of the C2 channel.
[0025] The beneficial effects of the present invention are as follows: the present invention proposes a fall recognition method based on human skeleton data. First, the depthwise separable convolution DWConv is introduced into the human skeleton key point detection network Light-YOLOv8-Pose to replace the ordinary convolution of the backbone network and the neck network, thereby reducing the number of model parameters and optimizing the detection speed; in the head network of Light-YOLOv8-Pose, the detection head is redesigned by combining lightweight convolution and self-attention mechanism to expand the receptive field of the model and enhance the feature extraction capability; in addition, WIoU is used as the loss function during the model training process to improve the convergence ability of the model.
[0026] For the human skeleton key point data extracted by the human skeleton key point detection network, the present invention proposes a three-branch spatiotemporal graph convolutional network TriBLight-STGCN to process and identify whether there is a falling action. TriBLight-STGCN improves the multi-scale feature extraction capability of the model through the cascaded three-branch spatiotemporal graph convolution module, thereby improving the recognition accuracy of the falling action; at the same time, the lightweight convolution GSConv is introduced to replace the ordinary convolution operations of spatial convolution and temporal convolution in the three-branch spatiotemporal graph convolution module, so as to further reduce the network parameters, so that the model has lightweight characteristics and high recognition accuracy, thereby reducing the computational amount of algorithm tasks and reducing computing resource requirements, which is suitable for scenarios with limited computing resources such as terminal embedded devices.
[0027] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0029] Figure 1 is a structural block diagram of the method of the present invention;
[0030] Figure 2 This is a schematic diagram of the Light-YOLOv8-Pose structure of the lightweight human key point detection network;
[0031] Figure 3 Schematic diagram of the self-attention detection head Pose_SA structure;
[0032] Figure 4 Schematic diagram of the three-branch spatiotemporal graph convolutional network TriBLight-STGCN;
[0033] Figure 5 Schematic diagram of the TriBLight-ST structure;
[0034] Figure 6 Schematic diagram of GSConv structure;
[0035] Figure 7 Schematic diagram of confusion matrix;
[0036] Figure 8 Schematic diagram of the recognition results of the self-built dataset. DETAILED DESCRIPTION
[0037] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0038] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0039] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0040] This embodiment provides a lightweight fall recognition method based on human skeleton data and a three-branch spatiotemporal graph convolutional network. The method includes the following steps:
[0041] S1. Build a lightweight human key point detection network Light-YOLOv8-Pose, which includes the depthwise separable convolution DWConv, the self-attention mechanism detection head Pose_SA, and the loss function WIoU.
[0042] like Figure 2 As shown in the figure, the constructed lightweight human key point detection network Light-YOLOv8-Pose includes a backbone network, a neck network, and a head network. In the backbone network and the neck network, the depthwise separable convolution DWConv is used to replace the ordinary convolution module. In the head network, the self-attention detection head Pose_SA is used to replace the original detection head. Finally, the optimized loss function WIoU is used to accelerate the convergence of the model. The principles of each module are as follows:
[0043] 1. Depthwise Separable Convolution DWConv
[0044] In standard convolution, filtering and combining the inputs are done in a single step, while depthwise separable convolution splits this process into two levels: depthwise convolution uses a single filter for each input channel, and pointwise convolution uses a 1×1 convolution to combine the outputs of the depthwise convolution. This decomposition significantly reduces computational complexity and model size. Compared to ordinary convolution, the reduction in computational complexity of DWConv is shown in the following formula:
[0045]
[0046] Among them, D K Represents the spatial size of the convolution kernel, M represents the number of channels of the input feature map, and D F represents the spatial size of the input feature map, and N represents the number of channels of the output feature map.
[0047] 2. Self-attention detection head Pose_SA
[0048] In the traditional YOLOv8-Pose network structure, the design of the detection head is relatively complex, and its parameters account for almost half of the total parameters of YOLOv8-Pose. However, since the present invention targets humans and only needs to identify one category of targets and key points, there is no need to use a complex detection head to ensure the completion of multi-category detection tasks. Therefore, the detection head of YOLOv8-Pose is redesigned by combining the self-attention module and the lightweight convolution module, abandoning the original method of realizing detection through multiple convolution operations, and obtaining the following Figure 3 The self-attention detection head shown.
[0049] The self-attention detection head, Pose_SA, is used to convert abstract features extracted by the network into specific detection results (such as target box coordinates, key points, category probabilities, and confidence levels). It includes a self-attention module and a lightweight convolution module. The self-attention module forms a residual-like structure based on the multi-head attention mechanism to preserve uncompressed residual information. This self-attention module can fully utilize a larger receptive field while maintaining low computational and memory overhead, enhancing feature extraction capabilities. The lightweight convolution module directly connects the input feature map to the output of the scaled convolution, effectively maintaining the input resolution. This lightweight convolution module significantly reduces information loss with minimal parameter increase, ensuring efficiency and accuracy.
[0050] 3. Optimize the loss function WIoU
[0051] The optimized loss function WIoU defines the abnormality of the anchor box as L IoU and The ratio between them is shown as follows:
[0052]
[0053] Among them, L IoU 、 They represent the base IoU loss of the sample and the loss adjusted by the dynamic focus mechanism, respectively.
[0054] A lower abnormality indicates that the anchor box has higher quality, allowing it to be assigned a smaller gradient gain, thereby focusing the bounding box regression on anchor boxes of normal quality. Conversely, anchor boxes with higher abnormality are assigned smaller gradient gains, effectively preventing low-quality samples from generating large and harmful gradients. Based on this concept, the WIoUv3 loss function is expressed as follows:
[0055] L WIoUv3 =rL WIoUv1 ,
[0056] Among them, when β = α, r = 1. When the abnormality of the anchor box reaches β = C (C is a constant value), the anchor box will obtain the maximum gradient gain. The quality classification criteria of anchor boxes are dynamic, allowing WIoUv3 to adaptively determine the most appropriate gradient gain allocation strategy for the current situation.
[0057] S2. By training the Light-YOLOv8-Pose network to obtain the optimal weights, a human key point detection model is established, and this model is used to achieve real-time human skeleton key point data detection and extraction.
[0058] The training process for the constructed Light-YOLOv8-Pose network is as follows: obtain training data and perform data preprocessing, including random scaling, random flipping, mixup and mosaic enhancement to improve the generalization ability of the model, and then divide the preprocessed dataset into training set and test set in an 8:2 ratio. The model training is completed based on the dataset to obtain the optimal model. The optimal model is used to process the image to be detected to obtain the key point data of the human skeleton.
[0059] In this example, the constructed Light-YOLOv8-Pose network was trained using the COCO Keypoints 2017 public dataset. The training environment was Python 3.8.19, Pytorch 1.13.1, CUDA 11.6 (NVIDIA GeForce RTX 3060 Laptop GPU, 8188 MiB), 300 training rounds, and a batch size of 8. Finally, the optimal model was obtained. Compared with the training results of the original YOLOv8-Pose model under the same training environment and parameter settings, as shown in Table 1, Light-YOLOv8-Pose reduced the number of parameters and computational complexity (GFLOPs) by 0.71 and 2.22, respectively, compared to the original model, while improving the average detection accuracy (mAP) by 2.2%.
[0060] Table 1
[0061]
[0062] S3. Construct a three-branch spatiotemporal graph convolutional network TriBLight-STGCN, which includes a three-branch spatiotemporal graph convolution module TriBLight-ST and a lightweight convolution GSConv.
[0063] like Figure 4The following figure shows the structure of the constructed three-branch spatiotemporal graph convolutional network TriBLight-STGCN, which includes 10 three-branch spatiotemporal graph convolution modules TriBLight-ST. The ordinary convolution in the spatial convolution and temporal convolution in each TriBLight-ST is replaced by the lightweight convolution GSConv. The principles of each module are as follows:
[0064] 1. Three-branch spatiotemporal graph convolution module TriBLight-ST
[0065] The three-branch spatiotemporal graph convolution module TriBLight-ST is different from the original spatiotemporal graph convolution module in that spatial convolution and temporal convolution are connected sequentially. TriBLight-ST divides the input into three branches. The first branch directly outputs, the second branch connects to the spatial convolution, and the third branch connects to the temporal convolution. Residual connections are adopted between the three branches to enhance the multi-scale feature extraction capability.
[0066] like Figure 5 As shown in the figure, the three-branch spatiotemporal graph convolution module TriBLight-ST divides the input feature map into s feature subsets, i.e. s branches, after 1×1 convolution, which are denoted as x i , where i∈{1,2,...,s}. Each feature subset x i It has the same spatial size as the input feature map, but the number of channels is 1 / s of the input feature map. Except for x1, each x i Each corresponds to a 3×3 convolution operation, denoted as K i (), its output is y i Represents the feature subset x i Enter K i () before, it will be compared with the output K of the previous convolutional layer i-1 () is added. In order to increase s and reduce the number of parameters, x i The 3×3 convolution is omitted, so y i It can be expressed as the following formula:
[0067]
[0068] Among them, each 3×3 convolution operation K i () will select all feature subsets x that satisfy j≤i j Receive feature information at j After 3×3 convolution, the output receptive field is larger than x j Bigger.
[0069] Since the correlation between a single multi-branch residual module and the overall network structure is weak, and its multi-scale representation capability works independently with aggregated feature models such as GCN and TCN, the multi-branch residual module can be fused with other modules to further improve network performance. Based on this, the TriBLight-ST module proposed in this embodiment transforms the spatial-temporal convolutional module (ST) of the original sequential structure into a structure similar to Res2Net. Specifically, the number of feature subsets in Res2Net is adjusted to 3, and one of the 3×3 convolutions is replaced by spatial convolution (GCN), and the other 3×3 convolution is replaced by temporal convolution (TCN).
[0070] 2. GSConv
[0071] The spatial convolution in the TriBLight-ST module and the ordinary convolution in the temporal convolution are replaced by GSConv to achieve network lightweight. Figure 6 As shown in the figure, the GSConv (Grouped Spatial Convolution) module combines standard convolution, depth-wise separable convolution (DWConv), concatenation (Concat), and shuffle (Shuffle) operations. In the specific process, GSConv first converts the number of channels of the input feature map from C1 to C2 / 2 through standard convolution to generate a C2 / 2 feature vector; then, depth-wise separable convolution (DWConv) is used to obtain another C2 / 2 feature vector. Then, the two feature vectors are concatenated through the Concat module, and finally the features generated by the standard convolution and the features generated by the depth-wise separable convolution are fused through the Shuffle operation, and finally the feature map of the C2 channel is output. By replacing the standard convolution with GSConv in the graph convolution, the number of parameters and computational complexity of the model can be effectively reduced. GSConv is expressed as follows:
[0072] X out =f shuffle (cat(f conv (X in ),f desc (f conv (X in ))))
[0073] S4. By training the TriBLight-STGCN network to obtain the optimal weight, a fall action recognition model is established, and the fall action recognition results are obtained by inputting the extracted human skeleton key points into this model.
[0074] In this embodiment, the three-branch spatiotemporal graph convolutional network TriBLight-STGCN constructed in step S3 is trained using the NTU RGB+D public dataset. The relevant training parameters are set as follows:
[0075] (1) The continuous frames are extracted from the human skeleton key point coordinates through the Light-YOLOv8-Pose model, converted into npy format data and input into TriBLight-STGCN. The number of key points is 17, the partitioning strategy selects the spatial configuration partitioning strategy, the dropout value is set to 0.5, and the training round epoch is set to 80.
[0076] (2) Considering the limited computing power of the computer equipment used for training, a smaller training batch size is set, with the input batch size set to 32 and the verification data processing batch size set to 32.
[0077] (3) Momentum Stochastic Gradient Descent (Momentum SGD) is used to update weights. Nesterov Accelerated Gradient (NAG) is used to speed up the convergence of gradient descent and reduce possible oscillations.
[0078] (4) The cross entropy function (Cross Entropy Loss) is used as the loss function, and its expression is as follows:
[0079]
[0080] Where x represents the input feature vector and class represents the label. The cross entropy function value is an indicator that can reflect the performance of the model.
[0081] Table 2 shows the parameter size comparison between the model proposed in this invention and the original model. It can be seen that the TriBLight-STGCN proposed in this invention reduces the parameter size by 9.69%.
[0082] Table 2
[0083] Model parameter number
[0084] ST-GCN 3098832
[0085] TriBLight-STGCN 2798313
[0086] Table 3 shows the experimental results of the proposed model and the original model on the NTU RGB+D dataset. For 60 categories of behaviors in the public dataset, the TriBLight-STGCN model achieved Top-1 and Top-5 accuracy of 83.17% and 95.29% under the cross-object criterion (X-Sub), respectively, an improvement of 2.82% and 3.24% compared to ST-GCN. The Top-1 and Top-5 accuracy under the cross-view criterion (X-View) were 88.99% and 98.22%, respectively, an improvement of 2.61% and 1.84% compared to ST-GCN.
[0087] Table 3
[0088]
[0089] In step S4, the fall recognition method combining Light-YOLOv8-Pose and TriBLight-STGCN is used to conduct experiments on a self-built dataset, which includes five normal behaviors: sitting down, standing up, walking, sitting, and squatting, as well as one abnormal behavior: falling, for a total of six types of actions. Figure 7 Shown is the confusion matrix of the recognition results on the self-built dataset. Figure 8 The figure shows the recognition results. It can be seen that the lightweight fall recognition method based on human skeleton data and three-branch spatiotemporal graph convolutional network proposed in the present invention has good detection effect.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A lightweight fall recognition method based on human skeleton data and a three-branch spatiotemporal graph convolutional network, characterized in that: The method includes: First, a lightweight human skeleton key point detection network and a three-branch spatiotemporal graph convolutional network are constructed respectively; wherein, the lightweight human skeleton key point detection network includes a backbone network, a neck network and a head network. The convolutions in the backbone network and the neck network both adopt depthwise separable convolution modules, and the detection head in the head network adopts a self-attention mechanism detection head; the three-branch spatiotemporal graph convolutional network includes multiple cascaded three-branch spatiotemporal graph convolution modules; Then, the lightweight human skeleton key point detection network is trained and optimized to establish a human skeleton key point detection model, and the human skeleton key point detection model is used to perform real-time detection and extraction of human skeleton key point data; the three-branch spatiotemporal graph convolutional network is trained and optimized to establish a fall action recognition model, and the fall action recognition result is obtained through the fall action recognition model; Finally, the real-time collected human motion image is input into the human skeleton key point detection model to extract the human skeleton key point data, and the human skeleton key point data is input into the falling action recognition model to determine whether the detected human body falls.
2. The method according to claim 1, characterized in that The lightweight human skeleton key point detection network is based on the YOLOv8-Pose network. The original convolution modules in the YOLOv8-Pose network backbone network and neck network are replaced with depthwise separable convolution modules, and the original detection head in the YOLOv8-Pose network head network is replaced with a self-attention detection head.
3. The method according to claim 2, characterized in that The self-attention detection head includes a self-attention module and a lightweight convolution module.
4. The method according to claim 1, wherein The three-branch spatiotemporal graph convolution network includes ten cascaded three-branch spatiotemporal graph convolution modules, each of which divides the input feature map into three feature subsets through 1×1 convolution, each feature subset has the same spatial size as the input feature map, but the number of channels is one-third of the input feature map; the first feature subset is directly output, the second feature subset is first spatially convolved and then output, and the third feature subset is first temporally convolved and then output, and the outputs of the three feature subsets are residually connected to enhance the multi-scale feature extraction capability; wherein, the second feature subset is first added to the first feature subset and then spatially convolved; the third feature subset is first added to the spatial convolution result and then temporally convolved.
5. The method according to claim 4, characterized in that In each of the three-branch spatiotemporal graph convolution modules, the spatial convolution and temporal convolution both use lightweight convolution to reduce the number of parameters and computational complexity; the lightweight convolution process includes: first, converting the number of channels of the input feature map from C1 to C2 / 2 through standard convolution to generate a first eigenvector of C2 / 2; then, convolving the first eigenvector using depthwise separable convolution to obtain a second eigenvector of C2 / 2; then, concatenating the first eigenvector and the second eigenvector; finally, fusing the features generated by standard convolution with the features generated by depthwise separable convolution through a shuffle operation, and finally outputting a feature map of the C2 channel.
6. The method according to claim 1, characterized in that When training the lightweight human skeleton key point detection network, WIoU is used as the loss function to optimize the model parameters and accelerate the convergence speed; When training the three-branch spatiotemporal graph convolutional network, the momentum stochastic gradient descent method is used to update the weights, Nesterov acceleration is used to speed up the convergence of gradient descent, and the cross entropy function is used as the loss function to optimize the model parameters.
7. A lightweight fall recognition system based on human skeleton data and a three-branch spatiotemporal graph convolutional network, characterized in that: The system includes a data acquisition module, a human skeleton key point detection module and a fall action recognition module; the data acquisition module is used to collect human motion images and perform preprocessing; the human skeleton key point detection module detects and extracts human skeleton key point data through a lightweight human skeleton key point detection network; the fall action recognition module processes the human skeleton key point data through a three-branch spatiotemporal graph convolutional network to obtain a fall action recognition result; The lightweight human skeleton key point detection network includes a backbone network, a neck network and a head network. The convolutions in the backbone network and the neck network both adopt a depth-separable convolution module, and the detection head in the head network adopts a self-attention mechanism detection head; the three-branch spatiotemporal graph convolution network includes multiple cascaded three-branch spatiotemporal graph convolution modules.
8. The system according to claim 7, characterized in that The lightweight human skeleton key point detection network is based on the YOLOv8-Pose network. The original convolution modules in the YOLOv8-Pose network backbone network and neck network are replaced with depthwise separable convolution modules, and the original detection head in the YOLOv8-Pose network head network is replaced with a self-attention detection head.
9. The system according to claim 7, wherein: The three-branch spatiotemporal graph convolutional network includes ten cascaded three-branch spatiotemporal graph convolutional modules, each of which divides the input feature map into three feature subsets through 1×1 convolution, each feature subset has the same spatial size as the input feature map, but the number of channels is one-third of the input feature map; The first feature subset is directly output, the second feature subset is first output after spatial convolution, and the third feature subset is first output after temporal convolution. The outputs of the three feature subsets are residually connected to enhance the multi-scale feature extraction capability; The second feature subset is first added to the first feature subset and then spatial convolution is performed; the third feature subset is first added to the spatial convolution result and then temporal convolution is performed.
10. The method according to claim 9, characterized in that In each of the three-branch spatiotemporal graph convolution modules, the spatial convolution and temporal convolution both use lightweight convolution to reduce the number of parameters and computational complexity; the lightweight convolution process includes: first, converting the number of channels of the input feature map from C1 to C2 / 2 through standard convolution to generate a first eigenvector of C2 / 2; then, convolving the first eigenvector using depthwise separable convolution to obtain a second eigenvector of C2 / 2; then, concatenating the first eigenvector and the second eigenvector; finally, fusing the features generated by standard convolution with the features generated by depthwise separable convolution through a shuffle operation, and finally outputting a feature map of the C2 channel.