A human fall detection method and apparatus

CN115588240BActive Publication Date: 2026-10-09AEROSPACE SCI & IND SHENZHEN GROUP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211554344.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2026-10-09
Estimated Expiration
2042-12-06

AI Technical Summary

Technical Problem

[0004]本发明所要解决的技术问题是怎样在训练样本较少的情况下,提出了一种低成本高准确率对人体跌倒检测方法及装置

Benefits of technology

本发明提供的一种对人体跌倒检测方法及装置,通过利用大规模数据集预训练模型的特征表达能力,将其应用到下游服务中,且为了避免过低的特征层没有语义信息,以及过高的特征层偏向于预训练模型的分类,只使用中间几层进行图像特征的提取,然后用较小的人体跌倒训练集使用预训练模型提取特征作为特征库,从而可以使输入视频所提取的图像特征与特征库进行比较,判断是否跌倒。使用本发明的方法需要的训练样本少,泛化到新场景的能力比较强,推理速度快。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115588240B_ABST
    Figure CN115588240B_ABST
Patent Text Reader

Abstract

The application provides a human body fall detection method and device, which utilizes the feature expression capability of a large-scale data set pre-training model, applies the pre-training model to a downstream task, and uses only middle layers to extract image features in order to avoid that a feature layer that is too low has no semantic information and a feature layer that is too high is biased to the classification of the pre-training model, and then uses a smaller human body fall training set to use the pre-training model to extract features as a feature library, so that the image features extracted from an input video can be compared with the feature library to determine whether a fall occurs. The method of the application requires fewer training samples, has strong generalization capability to new scenes, and has fast reasoning speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image recognition, and in particular relates to a method and device for detecting human falls. Background Technology

[0002] With the increasing aging of the population, the number of elderly people living alone is growing, and their health and well-being deserve serious attention. Accidental falls pose a significant threat to the lives of the elderly, and timely medical assistance after a fall can effectively reduce the risk of injury or death. According to publicly available data from relevant national surveys, the fall rate among people aged 65 and above in my country is as high as 16%, reaching 18.9% in rural areas. Some elderly people have experienced more than one fall. Furthermore, with increasing age, the risk of injury or death increases dramatically if assistance is not received promptly after a fall, potentially leading to significant disability and seriously threatening their physical and mental health. Therefore, automatically detecting and issuing alarms for accidental falls in the elderly is of practical and urgent significance.

[0003] Currently, fall detection methods at home and abroad are mainly divided into three categories: (1) Traditional feature extraction and comparison methods, which extract feature points based on designed image features and compare aspect ratios to determine whether a fall has occurred. These methods are simple to implement but have low accuracy. (2) Fall detection systems based on various sensors, usually wearable real-time fall detection systems based on microsystem triaxial accelerometers and dual-axis gyroscopes, or models human movements based on skeleton data provided by Kinect sensors, and fall recognition algorithms based on human motion feature parameters. Systems based on wearable sensors have a high false alarm rate due to the lack of overall information on human movements, while methods based on Kinect sensors have high installation costs and expensive sensors. (3) Deep learning-based target detection methods, which model the distribution of elderly fall images through a large number of training samples, train a neural network model, and detect elderly fall images under similar distributions. These methods are based on the assumption that training data and application scenario data are independently and identically distributed, require a large number of samples, and have limited generalization to new scenarios. Summary of the Invention

[0004] The technical problem to be solved by this invention is how to propose a low-cost and high-accuracy method and device for detecting human falls with a small number of training samples.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A method for detecting human falls includes the following steps: Step 1: Input video data and preprocess each frame of the image; Step 2: Extract features from each pre-processed frame using a pre-trained image classification model. The pre-trained image classification model refers to an image classification model trained on a large-scale dataset to extract image features. Step 3: Use the intermediate layers of the image classification model to extract the image features of each image in the human fall training set as the training feature library; extract the image features of each image in the human fall validation set. Step 4: Calculate the threshold for classifying human fall images based on the training feature library and the validation feature library; Step 5: Compare the image features extracted in Step 2 with the threshold for classifying human fall images. If the features are greater than the threshold, the image is judged as a fall.

[0006] Furthermore, the preprocessing method described in step 1 is: Each frame of the image is down-resolution processed; then the down-resolution image is denoised and contrast enhanced.

[0007] Furthermore, average pooling is performed on all image features in the training feature library, and the threshold for classifying human fall images is calculated using the average pooled training feature library and the validation feature library.

[0008] Furthermore, the method for calculating the threshold for classifying human fall images in step 4 is as follows: Step 4.1: For the training feature library after pooling, calculate the L2 distance between the image features of each image in the human fall verification set and all image features in the training feature library, obtain the maximum L2 distance of each image feature in the verification set and save it; Step 4.2: Calculate the classification accuracy for the maximum L2 distance of all image features in the validation set, and use the L2 distance corresponding to the highest classification accuracy as the threshold for classifying human fall images.

[0009] Furthermore, the method for comparing the image features extracted in step 2 with the threshold for classifying human fall images in step 5 is as follows: Step 5.1: Calculate the L2 distance between the image features extracted in Step 2 and each image feature in the training feature library to obtain the maximum L2 distance; Step 5.2: Compare the maximum L2 distance with the threshold for human fall image classification.

[0010] Furthermore, a sampling algorithm is used to sample the image features in the training feature library after average pooling to form a new feature library.

[0011] Furthermore, the sampling algorithm is a core set sampling algorithm, which is used to find a set of values ​​that maximize the representative features of the image features in the training feature library after average pooling processing, and to form a new feature library.

[0012] Furthermore, the method for calculating the threshold for classifying human fall images in step 4 is as follows: S4.1: Calculate the L2 distance between each image feature in the validation set and all image features in the new feature library, obtain the maximum L2 distance of each image feature in the validation set and save it; S4.2: Calculate the classification accuracy for the maximum L2 distance of all image features in the validation set, and use the L2 distance corresponding to the highest classification accuracy as the threshold for classifying human fall images.

[0013] Furthermore, the method for comparing the image features extracted in step 2 with the threshold for classifying human fall images in step 5 is as follows: S5.1: Calculate the L2 distance between the image features extracted in step 2 and each image feature in the new feature library to obtain the maximum L2 distance; S5.2: Compare the maximum L2 distance with the threshold for classifying human fall images.

[0014] The present invention also provides a human fall detection device, comprising the following modules: Input module: Used to input video data and preprocess each frame of the image; Image feature extraction module: used to extract features from each frame of the preprocessed image using a pre-trained image classification model, wherein the pre-trained image classification model refers to an image classification model trained on a large-scale dataset to extract image features; Feature library construction module: used to extract image features of each image in the human fall training set as the training feature library using the intermediate layers of the image classification model; and to extract image features of each image in the human fall validation set. Threshold calculation module: used to calculate the threshold for classifying human fall images based on the training feature library and the validation feature library; Inference module: Used to compare the image features extracted by the image feature extraction module with the threshold for classifying human fall images. If the features are greater than the threshold, the image is judged as a fall.

[0015] By adopting the above technical solution, the present invention has the following beneficial effects: This invention provides a method and apparatus for detecting human falls. It leverages the feature representation capabilities of a pre-trained model on a large dataset and applies it to downstream services. To avoid excessively low feature layers lacking semantic information and excessively high feature layers biased towards the classification of the pre-trained model, only the middle layers are used for image feature extraction. A small human fall training set is then used to extract features from the pre-trained model to create a feature library. This allows the extracted image features from the input video to be compared with the feature library to determine if a fall has occurred. The method of this invention requires fewer training samples, has strong generalization ability to new scenarios, and fast inference speed. Attached Figure Description

[0016] Figure 1 The system flowchart provided for the embodiments of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Figure 1 This invention illustrates a specific embodiment of a human fall detection method, comprising the following steps: Step 1: Input video data and preprocess each frame of the image.

[0019] In this embodiment, the preprocessing method is: Each frame of the image is down-resolution processed; then the down-resolution image is denoised and contrast enhanced.

[0020] The specific steps for resolution downscaling are as follows: Each frame is divided into four smaller images: top, bottom, left, and right. The central four-part image is then extracted to avoid the target image being completely divided, resulting in a total of five images. The longer side of each of these smaller images is uniformly scaled to 256 pixels. These images are then placed on a 256x256 pixel solid color (grayscale preferred) background image to facilitate feature extraction by the network and maintain the original image's aspect ratio without distortion. Finally, existing noise reduction and contrast enhancement algorithms are used for image preprocessing.

[0021] Step 2: Extract features from each pre-processed image frame using a pre-trained image classification model. The pre-trained image classification model refers to an image classification model trained on a large-scale dataset to extract image features.

[0022] In this embodiment, the feature representation capabilities of a pre-trained model on a large-scale dataset are utilized and applied to downstream tasks to extract image features. This eliminates the need to train a new model, saving significant training data. Consequently, the lack of a large training set of human fall detection data prevents the acquisition of a mature training model, allowing for the completion of human fall detection training with fewer training samples. The pre-trained model used in this embodiment is based on an image classification model pre-trained on the large-scale ImageNet dataset for feature extraction; however, other image classification models trained on large-scale datasets can also be used.

[0023] Step 3: Use the middle layers of the pre-trained image classification model to extract the image features of each image in the human fall training set as the training feature library; extract the image features of each image in the human fall validation set.

[0024] In this embodiment, only the middle layers of the pre-trained image classification model are used, typically two to three layers. This is because the first layer lacks semantic information, and the highest layer is biased towards ImageNet classification. Therefore, this embodiment only uses the middle two layers. Let φ represent the pre-trained network, using a network structure similar to ResNet, which has four downsampling layers and one classification head. j Let j represent the j-th layer of the network, where j = {1, 2, 3, 4}. In this embodiment, j = {2, 3} is used to extract features. This allows for the extraction of both low-level edge features and high-level semantic features such as shape and texture information.

[0025] In this embodiment, the training set for human falls contains approximately 200 normal images, while the validation set contains 50-100 normal images and a similar number of abnormal images. Image features extracted using the second and third layers of the pre-trained image classification model are stored in the training feature library M0.

[0026] In this embodiment, average pooling is performed on all image features in the training feature library. Average pooling can increase the receptive field; therefore, the image features in the training feature library are further average pooled to obtain the average pooled training feature library M1. In this embodiment, a 2x2 convolutional kernel is used for average pooling. At this point, the feature library M1 has two vectors of sizes, 32*32*128 and 16*16*256, for each training image, with a total number of features of N*(32*32*128+16*16*256), where N is the number of training samples in the human fall training set. Alternatively, a 3x3 convolutional kernel can also be used for average pooling.

[0027] Example 1: Step 4: Calculate the threshold for classifying human fall images based on the training feature library and the verification feature library.

[0028] After obtaining the training feature library M1 after average pooling, it is also necessary to determine the optimal threshold. In this embodiment, the method for calculating the threshold for classifying human fall images is as follows: Step 4.1: For the training feature library after pooling, calculate the L2 distance between the image features of each image in the human fall verification set and all image features in the training feature library, obtain the maximum L2 distance of each image feature in the verification set and save it; Step 4.2: Calculate the classification accuracy for the maximum L2 distance of all image features in the validation set, and use the L2 distance corresponding to the highest classification accuracy as the threshold for classifying human fall images.

[0029] Step 5: Compare the image features extracted in Step 2 with the threshold for classifying human fall images. If the features are greater than the threshold, the image is judged as a fall.

[0030] In this embodiment, the comparison method is as follows: Step 5.1: Calculate the L2 distance between the image features extracted in Step 2 and each image feature in the training feature library to obtain the maximum L2 distance; Step 5.2: Compare the maximum L2 distance with the threshold for human fall image classification.

[0031] Example 2: Because the number of image features in the training feature library is too large, resulting in a significant computational burden that increases linearly with the number of samples, downsampling of the features is necessary to reduce the feature library size, improve speed, and maintain accuracy. The difference between Example 2 and Example 1 lies in the use of a sampling algorithm to sample the training feature library M1 after average pooling to form a new feature library. The new feature library is used for L2 distance calculation.

[0032] In this embodiment, the sampling algorithm is the Coreset sampling algorithm, which is used to find a set of values ​​that maximize the representative features of the image features in the training feature library after average pooling, and use them to form a new feature library. This embodiment uses the Coreset sampling algorithm, setting the sampling ratio to 0.01, that is, sampling 1% of the features to form a new representative feature library M. At this time, the total number of features in the new feature library is 0.01*N*(32*32*128+16*16*256).

[0033] Another threshold algorithm in this embodiment is: S4.1: Calculate the L2 distance between each image feature in the validation set and all image features in the new feature library, obtain the maximum L2 distance of each image feature in the validation set and save it; S4.2: Calculate the classification accuracy for the maximum L2 distance of all image features in the validation set, and use the L2 distance corresponding to the highest classification accuracy as the threshold for classifying human fall images.

[0034] Step 5: Compare the image features extracted in Step 2 with the threshold for classifying human fall images. If the features exceed the threshold, it is determined to be a fall. The specific comparison method is as follows: S5.1: Calculate the L2 distance between the image features extracted in step 2 and each image feature in the new feature library to obtain the maximum L2 distance; S5.2: Compare the maximum L2 distance with the threshold for classifying human fall images.

[0035] If the image exceeds the threshold and is judged as a fall, the system issues an alarm and terminates the detection; otherwise, if the image is considered normal, the detection continues.

[0036] The present invention also provides a human fall detection device, comprising the following modules: Input module: Used to input video data and preprocess each frame of the image; Image feature extraction module: used to extract features from each frame of the preprocessed image using a pre-trained image classification model, wherein the pre-trained image classification model refers to an image classification model trained on a large-scale dataset to extract image features; Feature library construction module: used to extract image features from each image in the human fall training set as a training feature library using the intermediate layers of the image classification model; and to extract image features from each image in the human fall validation set as a validation feature library; Threshold calculation module: used to calculate the threshold for classifying human fall images based on the training feature library and the validation feature library; Inference module: Used to compare the image features extracted by the image feature extraction module with the threshold for classifying human fall images. If the features are greater than the threshold, the image is judged as a fall.

[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting human falls, characterized in that, Includes the following steps: Step 1: Input video data and preprocess each frame of the image; Step 2: Extract features from each pre-processed frame using a pre-trained image classification model. The pre-trained image classification model refers to an image classification model trained on a large-scale dataset to extract image features. Step 3: Use the middle layers of the image classification model to extract the image features of each image in the human fall training set as the training feature library, and extract the image features of each image in the human fall validation set as the validation feature library; All image features in the training feature library are subjected to average pooling. The image features in the training feature library after average pooling are sampled using a core set sampling algorithm to form a new feature library. The core set sampling algorithm is used to find a set of values ​​that maximize the representative features at the feature boundaries of the image features in the training feature library after average pooling, which are used to form the new feature library. Step 4: Calculate the threshold for classifying human fall images based on the training feature library and the validation feature library; The method for calculating the threshold for classifying human fall images is as follows: S4.1: Calculate the L2 distance between each image feature in the validation set and all image features in the new feature library, obtain the maximum L2 distance of each image feature in the validation set and save it; S4.2: Calculate the classification accuracy for the maximum L2 distance of each image feature in the validation set, and use the L2 distance corresponding to the highest classification accuracy as the threshold for classifying human fall images; Step 5: Calculate the L2 distance between the image features extracted in Step 2 and each image feature in the new feature library to obtain the maximum L2 distance; compare the maximum L2 distance with the threshold for classifying human fall images; if it is greater than the threshold, it is judged as a fall.

2. The method according to claim 1, characterized in that, The preprocessing method described in step 1 is: Each frame of the image is down-resolution processed; then the down-resolution image is denoised and contrast enhanced.

3. A human fall detection device, characterized in that, Includes the following modules: Input module: Used to input video data and preprocess each frame of the image; Image feature extraction module: used to extract features from each frame of the preprocessed image using a pre-trained image classification model, wherein the pre-trained image classification model refers to an image classification model trained on a large-scale dataset to extract image features; Feature library construction module: used to extract image features of each image in the human fall training set as the training feature library and extract image features of each image in the human fall validation set as the validation feature library using the intermediate layers of the image classification model; All image features in the training feature library are subjected to average pooling. The image features in the training feature library after average pooling are sampled using a core set sampling algorithm to form a new feature library. The core set sampling algorithm is used to find a set of values ​​that maximize the representative features at the feature boundaries of the image features in the training feature library after average pooling, which are used to form the new feature library. Threshold calculation module: used to calculate the threshold for classifying human fall images based on the new feature library and the verification feature library; the calculation method is as follows: calculate the L2 distance between each image feature in the verification set and all image features in the new feature library, obtain the maximum L2 distance of each image feature in the verification set and save it; calculate the classification accuracy for the maximum L2 distance of each image feature in the verification set, and take the L2 distance corresponding to the value with the highest classification accuracy as the threshold for classifying human fall images; Inference module: Used to calculate the L2 distance between the image features extracted by the image feature extraction module and each image feature in the new feature library, and obtain the maximum L2 distance; compare the maximum L2 distance with the threshold for classifying human fall images, and if it is greater than the threshold, it is judged as a fall.

Citation Information

Patent Citations

  • Rapid face recognition method with angle resistance and shielding interference

    CN108509862A

  • Fall detection method and device and robot

    CN112418096A

  • Road traffic accident detection method and device, computer and storage medium

    CN113378803A

  • Pavement anomaly detection method and system based on unsupervised learning

    CN114663742A