Lightweight fall detection method and system based on improved YOLOv8
By improving the YOLOv8 model, GSConv, VoV-GSCSP and Detect_LADH modules were introduced, and inference acceleration was performed, which solved the problem of real-time demand for YOLOv8 in resource-constrained environments, realized efficient and lightweight fall detection, and expanded the potential for technical application.
Patent Information
- Application Number
- CN202510075999.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-23
AI Technical Summary
The existing YOLOv8 algorithms still need to be further lightened to meet real-time requirements in resource-constrained environments, such as mobile devices or embedded systems.
By introducing the GSConv module, VoV-GSCSP module and Detect_LADH lightweight detection head, the YOLOv8 model was improved to obtain the VGL-YOLO lightweight fall detection model, and optimized operations such as pruning, quantization, and distillation, combined with TensorRT for inference acceleration, and deployed to the edge computing terminal Jetson TX2 NX.
It realizes the lightweight model while maintaining high detection accuracy and performance, solves the problems of slow real-time response in edge computing devices and limited resource environment in traditional fall detection methods, and expands the application potential of fall detection technology in the Internet of Things, mobile devices and low-power scenarios.
Smart Images

Figure CN120032292A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a lightweight fall detection method and system based on improved YOLOv8. Background Art
[0002] As the aging population increases, the demand for smart elderly care services is also increasing. According to statistics, more than half of the elderly people visit the hospital due to falls, which may cause minor injuries or even fatal injuries. Fall detection systems can alleviate this problem and reduce the time it takes for patients who fall to receive assistance.
[0003] In recent years, with the development of deep learning, the YOLO series of algorithms have been widely studied and applied in the field of fall detection. The YOLOv8 algorithm shows great potential in the field of fall detection with its fast processing speed and high detection accuracy. However, in resource-constrained environments, such as mobile devices or embedded systems, the YOLOv8 algorithm still needs to be further lightweight to meet real-time requirements. Summary of the invention
[0004] The present invention aims to solve the problems existing in the above-mentioned prior art and provide a lightweight fall detection method and system based on improved YOLOv8.
[0005] The technical solutions adopted in the present invention are:
[0006] A lightweight fall detection method and system based on improved YOLOv8, comprising the following steps:
[0007] S1: Build a fall detection dataset;
[0008] S2: Improve the YOLOv8 model by introducing the GSConv module, VoV-GSCSP module and Detect_LADH lightweight detection head to obtain the VGL-YOLO lightweight fall detection model;
[0009] S3: Accelerate the reasoning of the lightweight detection model and embed it into the edge computing terminal Jetson TX2NX for deployment;
[0010] S4: Collect data through the camera and transmit it to the computing terminal Jetson TX2NX for processing;
[0011] S5: The processed data is displayed in real time through a visualization platform.
[0012] Furthermore, in step S1, the fall detection dataset includes an open source dataset of fall-related image data and image data obtained by photographing the experimenter with a camera; the image size is unified, and LabelImg is used to manually annotate the posture of the person in the image, and the minimum enclosing rectangle of the target is used as the real frame during annotation.
[0013] Furthermore, in step S2, the method for improving the YOLOv8 model is:
[0014] The GSConv module is introduced into the YOLOv8 model to replace the original Conv convolution module. The GSConv module first generates a feature map with a channel number of C2 / 2 by performing a convolution operation on the Conv module, and then uses the DWConv module to perform a convolution operation on each channel generated by the Conv module independently, outputting a feature map with a channel number of C2 / 2. The Concat module is used to splice the output results of the Conv module and the DWConv module. Finally, the information generated by the Conv module is infiltrated into the feature map of the DWConv through the shuffle layer operation. It adopts a uniform mixing strategy to exchange local feature information on different channels, so that the GSConv module can keep the computational efficiency of the DWConv module while being as close as possible to the feature expression ability of the Conv module.
[0015] The DWConv module consists of a 3×3 deep convolution and a 1×1 point-by-point convolution. Deep convolution requires only one filter per channel, unlike traditional convolution where each channel needs to be convolved with all filters. Point-by-point convolution linearly combines the feature maps of each channel generated by deep convolution to generate the final output feature map, which greatly reduces the number of model parameters and speeds up model training and reasoning.
[0016] In the YOLOv8 model, a cross-stage partial network module VoV-GSCSP is introduced to replace the original C2f module. The VoV-GSCSP module consists of two branches, one branch consists of a Conv module, a GSBottleneck module and a Conv convolution module, and the other branch is a Conv module. The two branches are connected by a Concat module. The GSBottleneck module is the core of VoV-GSCSP, which consists of two GSConv modules and one Conv module. The two GSConv modules are used to enhance the nonlinear expression and information reuse of features, while the middle Conv module is used to further process features and enhance the learning ability of the model. Each channel is responsible for processing different features and tasks to enhance the feature extraction ability and detection accuracy of the model.
[0017] The YOLOv8 model introduces a lightweight detection head Detect_LADH module to replace the original detection head Detect module. The Detect_LADH module consists of three different channels. The first channel consists of two 1×1 Conv modules, the second channel consists of three 3×3 DWConv modules and one 1×1 Conv module, and the third channel consists of two 1×1 Conv modules.
[0018] Furthermore, the GSConv module first downsamples the input feature map through a common convolution Conv to generate a feature map with a channel number of C2 / 2, and the formula is:
[0019] x′=W cv1 *x+b cv1
[0020] Where x′ is the feature map with C2 / 2 channels after downsampling of the ordinary convolution, x is the input feature map, cv1 is the ordinary convolution, and W cv1 is the convolution kernel, b cv1 is the bias term, * indicates the convolution operation;
[0021] Then use the depth-separable convolution DWConv to perform convolution operations on each channel generated by the Conv module independently, and output a feature map with a channel number of C2 / 2. The formula is:
[0022] x″=D cv2 *(P cv2 *x′+b cv2 )
[0023] Where x″ represents the feature map with C2 / 2 channels after the depthwise separable convolution DWConv operation, cv1 is the depthwise separable convolution, D cv2 is the depth convolution kernel, P cv2 is the point convolution kernel, b cv2 is the bias term, * indicates the convolution operation;
[0024] And concatenate the results generated by ordinary convolution and depth-separable convolution, the formula is:
[0025] X = Concat(x′, x″)
[0026] Where x′ represents the concatenated feature map, and Concat represents concatenating two feature maps in the channel dimension;
[0027] Finally, the feature map is mixed, and the concatenated feature map X is rearranged in the channel dimension through the Shuffle operation to achieve feature mixing. The formula is:
[0028] y = Shuffle(X)
[0029] Where y is the feature map after the Shuffle operation;
[0030] Finally, the feature map processed by the GSConv module is obtained, and the formula is:
[0031] GSConv(x)=y
[0032] Output result y.
[0033] Furthermore, in step S3, the lightweight detection model is first optimized by pruning, quantization, distillation, etc. to adapt to the performance limitations of edge computing devices, and then the model is converted to ONNX format, and then TensorRT is used for reasoning acceleration, including tensor fusion of network layers (horizontally merging layers with the same parameters and vertically merging layers with the same structure but different parameters) and replacing FP32 tensors with FP16 and INT8 tensors to reduce precision and accelerate reasoning. Finally, we deploy the optimized model on Jetson TX2NX.
[0034] The present invention has the following beneficial effects:
[0035] The method of the invention achieves lightweight model while maintaining high detection accuracy and performance, and is successfully deployed on the edge computing device Jetson TX2NX. This innovation solves the problems faced by traditional fall detection methods in edge computing devices, such as slow real-time response and limited resources and environment; it also greatly expands the application potential of fall detection technology in the Internet of Things, mobile devices and various low-power scenarios, and provides strong support for the development of fall warning and health protection systems for the elderly. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 The figure is a flowchart of a lightweight fall detection method based on improved YOLOv8.
[0037] Figure 2 This is the GSConv model structure diagram.
[0038] Figure 3 This is the VoV-GSCSP model structure diagram.
[0039] Figure 4 This is the structure diagram of the GSBottleneck module.
[0040] Figure 5 This is the structural diagram of the lightweight detection head Detect_LADH model.
[0041] Figure 6 This is the network structure diagram of the YOLOv8 model.
[0042] Figure 7 This is the network structure diagram of the VGL-YOLO model.
[0043] Figure 8 This is the PR curve of the VGL-YOLO model.
[0044] Fig. 9 Deploy a flowchart for the model. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] Embodiment 1
[0048] This embodiment discloses a lightweight fall detection method and system based on improved YOLOv8, see Figure 1 , including the following steps:
[0049] S1: Build a fall detection dataset.
[0050] The dataset is a public dataset provided by Le2i, a laboratory at the University of Bourges in France, specifically for studying fall detection. It contains different indoor scenes, such as living rooms, bedrooms, kitchens, etc., simulating various activity environments of the elderly in their daily lives. We extract key frames from the video, capture key actions before, during, and after the fall, and annotate the behaviors in the dataset. The final dataset contains a total of 3766 images, of which 2636 images are used for training, 753 images are used for verification, and 377 images are used for testing.
[0051] S2: Improve the YOLOv8 model to obtain the VGL-YOLO lightweight fall detection model. The improvements to the YOLOv8 model include the following three aspects:
[0052] S21: Introduce the GSConv module in the YOLOv8 model to replace the ordinary convolution module.
[0053] GSConv is a lightweight convolutional structure. First, the Conv module performs a convolution operation to generate a feature map with a channel number of C2 / 2. Then, the DWConv module is used to perform convolution operations on each channel independently. The Concat module is used to splice the output results of the Conv module and the DWConv module. Finally, the information generated by the Conv module is infiltrated into the feature map of the DWConv through the shuffle layer operation. It adopts a uniform mixing strategy to exchange local feature information on different channels, so that the GSConv module can keep the computational efficiency of the DWConv module while being as close to the feature expression ability of the Conv as possible. The GSConv module structure diagram is shown in the figure. Figure 2 shown.
[0054] The GSConv module first downsamples the input feature map through a normal convolution Conv to generate a feature map with a channel number of C2 / 2, the formula is:
[0055] x′=W cv1 *x+b cv1
[0056] Where x′ is the feature map with C2 / 2 channels after downsampling of the ordinary convolution, x is the input feature map, cv1 is the ordinary convolution, and W cv1 is the convolution kernel, b cv1 is the bias term, * indicates the convolution operation;
[0057] Then use the depth-separable convolution DWConv to perform convolution operations on each channel generated by the Conv module independently, and output a feature map with a channel number of C2 / 2. The formula is:
[0058] x″=D cv2 *(P cv2 *x′+b cv2 )
[0059] Where x″ represents the feature map with C2 / 2 channels after the depthwise separable convolution DWConv operation, cv1 is the depthwise separable convolution, D cv2 is the depth convolution kernel, P cv2 is the point convolution kernel, b cv2 is the bias term, * indicates the convolution operation;
[0060] And concatenate the results generated by ordinary convolution and depth-separable convolution, the formula is:
[0061] X = Concat(x′, x″)
[0062] Where x′ represents the concatenated feature map, and Concat represents concatenating two feature maps in the channel dimension;
[0063] Finally, the feature map is mixed, and the concatenated feature map X is rearranged in the channel dimension through the Shuffle operation to achieve feature mixing. The formula is:
[0064] y = Shuffle(X)
[0065] Where y is the feature map after the Shuffle operation;
[0066] Finally, the feature map processed by the GSConv module is obtained, and the formula is:
[0067] GSConv(x)=y
[0068] Output result y.
[0069] S22: Introduce the cross-stage partial network module VoV-GSCSP module in the YOLOv8 model to replace the C2f module.
[0070] The VoV-GSCSP consists of two branches, one of which is composed of a Conv module, a GSBottleneck module and a Conv convolution module, and the other is a Conv module. The two branches are connected by a Concat module, wherein the GSBottleneck module is the core of the VoV-GSCSP, which is composed of two GSConv modules and one Conv module. The two GSConv modules are used to enhance the nonlinear expression of features and information reuse, while the middle Conv module is used to further process features and enhance the learning ability of the model. Each channel is responsible for processing different features and tasks to enhance the feature extraction ability and detection accuracy of the model. The GSBottleneck module structure is shown in FIG. Figure 3 shown.
[0071] S23: Introduce the lightweight detection head Detect_LADH in the YOLOv8 model to replace the original detection head. The Detect_LADH module consists of three different channels. The first channel consists of two 1×1 Conv modules, the second channel consists of three 3×3 DWConv modules and one 1×1 Conv module, and the third channel consists of two 1×1 Conv modules.
[0072] S3: Accelerate the reasoning of the lightweight detection model and embed it into the edge computing terminal Jetson TX2 NX for deployment.
[0073] Before deploying the lightweight detection model to the edge computing device, it first needs to be optimized to adapt to the performance limitations of the edge device. This includes pruning operations, which reduce the complexity of the model by removing unimportant weights or neurons in the model; quantization operations, which convert the model's weights from FP32 to INT8 to reduce the model size and accelerate reasoning; and distillation operations, which transfer the knowledge of the complex model to a smaller model through knowledge distillation technology to improve the performance of the lightweight model.
[0074] The optimized model needs to be converted to ONNX format, which is an open format that supports model conversion between different deep learning frameworks. Subsequently, we use TensorRT to further accelerate the reasoning of the ONNX model. TensorRT builds an efficient reasoning engine by performing tensor fusion of network layers, eliminating unnecessary transposition operations, and automatic kernel adjustment. In addition, TensorRT can also replace FP32 tensors with FP16 and INT8 tensors, reducing the accuracy requirements of the model, thereby accelerating the reasoning process while maintaining model performance.
[0075] Finally, we deployed the TensorRT optimized model to the Jetson TX2NX.
[0076] S4: Collect data through the camera and transmit it to the edge computing terminal Jetson TX2NX for processing;
[0077] S5: The processed data is displayed in real time through a visualization platform.
[0078] Example 2
[0079] The present invention also provides a fall detection system based on key points of the human body, the system comprising: a data acquisition module, a data processing module and an alarm module;
[0080] The data acquisition module uses visual sensor devices to collect real-time image data and transmits it to the edge computing terminal Jetson TX2 NX;
[0081] The data processing module processes the collected information through the improved algorithm deployed in Jetson TX2 NX;
[0082] The alarm module determines whether the action information is a fall, and sends an alarm message to the user when the action information is a fall.
[0083] Example 3
[0084] This embodiment performs target detection on fall images through specific experiments to verify the beneficial effects of the method of the present invention.
[0085] 1. Experimental Environment
[0086] The configuration of this experimental platform is shown in Table 1.
[0087] Table 1 Experimental platform
[0088]
[0089] The experiment uses the public dataset Le2i as the training, testing, and validation dataset, which includes different indoor scenes, such as living rooms, bedrooms, kitchens, etc., simulating various activity environments of the elderly in daily life. We extract key frames from the video to capture the key actions before, during, and after the fall, and divide the behaviors in the dataset into two categories: falls and no falls. The final dataset contains a total of 3766 images, of which 2636 images are used for training, 753 images are used for validation, and 377 images are used for testing.
[0090] 2. Evaluation Indicators
[0091] This example uses mean average precision (mAP), parameter quantity (Params) and floating point
[0092] The amount of operations (FLOPs) is used as the evaluation indicator of the model. The specific calculation formula is as follows:
[0093]
[0094]
[0095] Where, T P is the number of correctly detected falling targets, F P is the number of targets that were falsely detected as falling, F N is the number of samples where falls were not detected.
[0096]
[0097]
[0098] In the formula, N represents the total number of categories.
[0099] 3. Experimental Results
[0100] In this experiment, several advanced fall detection models are selected for comparative experiments. The experimental results are shown in Table 2.
[0101] Table 2 Comparison of different detection models and VGL-YOLO
[0102]
[0103] From the above experimental results, we can see that VGL-YOLO is 0.6% and 2.7% higher than the baseline YOLOv8 in terms of mAP@0.5 and mAP@0.5-0.95, respectively, and is superior to other detection models. Compared with other YOLO detection models and mainstream algorithms such as SSD and Faster-RCNN, the VGL-YOLO model has higher accuracy, with mAP@0.5 and mAP@0.5-0.95 values of 99.5% and 94.4% respectively, and the number of parameters and calculations are also lower, avoiding a lot of computational overhead.
[0104] In order to further verify its effectiveness, a series of ablation experiments are set up to explore the impact of each module on the model performance. The experimental results are shown in Table 3.
[0105] Table 3 Ablation experiment
[0106]
[0107] Comprehensive experimental results show that each module improves the accuracy of VGL-YOLO to varying degrees. After replacing the Conv module and C2f module in the YOLOv8 model with the GSConv module and the VoV-GSCSP module, the model performance mAP@0.5 and mAP@0.5-0.95 increased by 0.6% and 2.1% respectively. After the lightweight detection head LADH_Detect was introduced, the number of parameters decreased significantly and mAP@0.5 increased by 0.5% compared to the original model. Finally, compared with the original model, the VGL-YOLO model of the present invention increased mAP@0.5 by 0.6%, mAP@0.5-0.95 by 2.7%, reduced the number of parameters by 1.21M, and reduced the floating point number by 4.7G.
[0108] The above description is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.
Claims
1. A lightweight fall detection method and system based on improved YOLOv8, characterized in that: The steps include: S1: Build a fall detection dataset; S2: Improve the YOLOv8 model by introducing the GSConv module, VoV-GSCSP module and Detect_LADH lightweight detection head to obtain the VGL-YOLO lightweight fall detection model; S3: Accelerate the reasoning of the lightweight detection model and embed it into the edge computing terminal Jetson TX2 NX for deployment; S4: Collect data through the camera and transmit it to the computing terminal Jetson TX2 NX for processing; S5: The processed data is displayed in real time through a visualization platform.
2. The lightweight fall detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: In step S1, the public images containing human falls and the images taken by the image acquisition device are classified into categories, annotated using the Labelimg tool, and a fall detection dataset is constructed.
3. The lightweight fall detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: The method for improving the YOLOv8 model in step S2 is: Introducing lightweight convolution technology GSConv to replace ordinary Conv in the YOLOv8 model; Replace the C2f module in the YOLOv8 model with the cross-stage partial network module VoV-GSCSP; Introduce the lightweight detection head Detect_LADH in the YOLOv8 model.
4. The lightweight fall detection method based on improved YOLOv8 as claimed in claim 3, characterized in that: The GSConv module first performs convolution downsampling on the input signal, and then uses DWConv depth convolution to combine the two convolution results SC and DSC, and obtains the output result through shuffle layer processing.
5. The lightweight fall detection method based on improved YOLOv8 as claimed in claim 3, characterized in that: The VoV-GSCSP cross-stage partial network module divides the input signal into two groups of convolutions, one of which is feature processed by the GSBottleneck module and the result is integrated with the other group of unprocessed convolutions to obtain the output result.
6. The lightweight fall detection method based on improved YOLOv8 as claimed in claim 3, characterized in that: The Detect_LADH module adopts an asymmetric head design and uses 3×3 depthwise separable convolution DWConv as a substitute in the ADH network to gradually reduce the number of convolution kernels and channels.
7. The lightweight fall detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: In step S3, the lightweight detection model is converted into ONNX format, then input into the TensorRT framework for inference acceleration processing, and the model is deployed on the edge computing terminal Jetson TX2 NX.