A method and device for real-time human fall detection and alarm based on video

Through real-time human fall detection and alarm methods based on video processing, the camera and local edge computing chips are used to solve the problems of equipment inconvenience and high cost in the existing fall detection technology, and the efficient and privacy-protected fall detection and alarm functions are achieved.

CN115082825BActive Publication Date: 2025-06-27CHINA-SINGAPORE INT JOINT RES INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210682605.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-06-27
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

The existing fall detection technology has the problems of the elderly not liking to wear, forgetting to wear, fixed wear position and insufficient battery life. Environmental solutions require the transformation of the living environment, and the system composition is complex and the cost is high.

Method used

采用基于视频处理的实时人体跌倒检测及报警方法,通过摄像头检测目标是否跌倒,并利用本地边沿计算芯片进行处理,避免云端检测以保护用户隐私。 This method includes establishing a fall detection data set, pruning the YOLOv2-Tiny network, building a fall detection model, and determining whether a fall occurs through the threshold of aspect ratio and center of gravity offset.

Benefits of technology

It realizes contactless fall detection, avoids user privacy leakage, reduces equipment cost and complexity, improves detection accuracy and ease of use, and is suitable for fall detection in multi-person scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082825B_ABST
    Figure CN115082825B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for real-time human fall detection and alarm based on video. The method includes the following steps: on the basis of a publicly available human detection data set, adding human image information in the fall state to establish a fall detection data set; pruning and transforming the YOLOv2-Tiny network to build a fall detection model, and training the fall detection model based on the self-established fall data set; by assigning different thresholds to the sensitivity of the aspect ratio α and the center of gravity point offset d of the human body before and after falling, obtaining new determination parameters for falling to achieve the judgment of falling; obtaining a real-time video as the video stream for fall detection and inputting it into the fall detection model. The fall detection model first performs human target detection on the frame images obtained by processing the input video, and then performs fall detection and alarm on the identified human targets. The present invention can be applied to people and places prone to falling, improving the rescue efficiency for people who fall.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of real-time object detection, and in particular to a method and device for real-time human fall detection and alarm based on video. Background Art

[0002] According to Forbes, by 2050, the global population aged 60 and above is expected to reach 2 billion. In this regard, this will represent more than one-fifth of the global population. Since the physical functions of the elderly decline significantly with age, most elderly people suffer from cardiovascular diseases, osteoporosis and other symptoms. These diseases and the side effects of drugs will increase the likelihood of the elderly falling. Falls can cause sprains, contusions, fractures, and even trigger the onset of other diseases. If the elderly cannot receive timely medical treatment and assistance after falling, their life safety will surely be seriously threatened.

[0003] With the increase in the elderly population, more advanced home monitoring is needed while still allowing people to maintain personal autonomy and privacy. According to the data of the Centers for Disease Control and Prevention in the United States, nearly one-fourth of the elderly fall each year, and falls are the main cause of trauma-related hospitalizations for the elderly. At present, the fall detections on the market are all contact-based inductive detections. One is the wearable sensor solution, and the other is the environmental solution. The wearable sensor solution is to wear a multi-axis acceleration sensor on the elderly person, and the environmental solution is to install sound and collision detection sensors on the floor of the home environment. The former judges whether a person has fallen through acceleration parameters, and the latter judges whether a person has fallen by using environmental parameters such as sound and vibration. However, both of them have some disadvantages. For example, the wearable sensor solution in a wristband touch-pressure detection fall device and its alarm system and implementation method with the publication number CN108041772A has limitations in many aspects such as the elderly not liking to wear, forgetting to wear, fixed wearing position and battery life. The environmental solution in a fall detection floor and method with the publication number CN111538264A requires the transformation of their living environment, and the system composition is complex and the cost is high. Therefore, it is very necessary to invent a device for fall detection based on video processing that can perform fall detection on multiple targets and judge whether the target has fallen. Summary of the Invention

[0004] In response to the problem of falls of people prone to falls, and from the actual needs of maintaining personal autonomy and privacy, the present invention proposes a method and device for real-time human fall detection and alarm based on video. The main starting point of the invention is that when people prone to falls fall, the device can send a fall message to the relatives of the people prone to falls or community social workers to achieve the means of alarm. Because the fall detection system uses a video image detection solution and provides non-contact sensing through accurate image data, there is no need for people prone to falls to wear sensors or install sensors in the home environment. Only the camera is needed to determine whether the target has fallen and alarm. However, the camera will involve user privacy issues, so all processing of the present invention is processed by the local edge computing chip, and the fall of the person will not be detected through the cloud server, avoiding the leakage of user privacy.

[0005] The purpose of the present invention is achieved by at least one of the following technical solutions.

[0006] A method for real-time human fall detection and alarm based on video, comprising the following steps:

[0007] S1. Based on the public human body detection dataset, add human body image information in the fall state to establish a fall detection dataset;

[0008] S2. Prune and transform the YOLOv2-Tiny network, build a fall detection model, and train the fall detection model based on the self-established fall dataset;

[0009] S3, by assigning different thresholds to the sensitivity of the image aspect ratio α and the center of gravity offset d before and after the human body falls, a new fall judgment parameter is obtained to realize the fall judgment;

[0010] S4. Real-time video is obtained as a video stream for fall detection and input into the fall detection model. The fall detection model first performs human target detection on the frame image obtained by processing the input video, and then performs fall detection and alarm on the identified human target.

[0011] Furthermore, in step S1, the public human detection data set is preliminarily screened to remove useless data, which refers to images with only partial hands and legs and no human torso features. Images with 80% of human body features appearing on the screen are useful data;

[0012] Collect and shoot multiple videos of people falling, capture frames of the videos, capture multiple frames per second, manually filter out pictures between standing and falling, obtain a data set with human falling postures, and annotate the filtered pictures, labeling the filtered pictures with person tags;

[0013] Perform data augmentation on the dataset with human fall postures. Generate multiple data-augmented images from the images in the dataset with human fall postures through image processing operations, including rotation, translation, and stretching, to obtain a data-augmented dataset with human fall postures.

[0014] Merge the data-augmented dataset with human fall postures and the screened public human detection dataset to establish a fall detection dataset.

[0015] Furthermore, in step S2, perform weight sparsification training on the original weights of the YOLOv2-Tiny network model, that is, introduce a scaling factor γ of the BN layer in each channel and multiply it with the output of the channel; then train the weights and the scaling factor, prune the channels with small scaling factors, and fine-tune the pruned network to achieve the purpose of weight sparsification.

[0016] Among them, channel pruning and fine-tuning are as follows:

[0017] After introducing the L1 regularization term of the scaling factor, the scaling factors in the obtained model will tend to 0; then sort the absolute values of the scaling factors first, take the scaling factor at the 80% position of the sorted scaling factors from small to large as the threshold, and prune the channels corresponding to the small scaling factors γ below the threshold. Pruning the channels corresponding to the small scaling factors γ is essentially directly pruning the convolution kernels corresponding to this channel; by doing so, a compact network with fewer parameters, less memory occupation during operation, and low computational complexity can be obtained: the Prune-YOLOv2-Tiny network model.

[0018] Change the category C in the detection layer of the Prune-YOLOv2-Tiny network model to a single class, that is, 1; at the same time, use the K-means algorithm to update the values of 5 anchors in the detection layer for the fall detection dataset. The values of the anchors are the width and height of the prediction box; the number of convolution kernels R in the last convolutional layer of the Prune-YOLOv2-Tiny network model is calculated using formula (1) and modified to the corresponding 30, specifically as follows:

[0019] R = Anchors * (5 + C) (1)

[0020] Among them, Anchors is the number of prediction boxes in the Prune-YOLOv2-Tiny network model, and the number is 5. The obtained R is the output channel size, which is used to obtain the tensor of the prediction output and generate the target box; thus, a fall detection model is obtained.

[0021] Furthermore, in step S2, train through the dataset on the fall detection model to obtain the weights in the fall detection model, specifically as follows:

[0022] The dataset is divided into a test set and a training set according to a ratio of 1:9. For each training image in the training set, the IOU loss, classification loss, and coordinate loss between the model prediction result and the true label are calculated on the fall detection model loaded with pre-trained weights, where the pre-trained weights are existing weights;

[0023] When the fall detection model reaches fitting, save the weights of the fall detection model, adjust the learning rate, and start the next round of training.

[0024] Furthermore, under the same test set, by comparing the weights of the fall detection models obtained from different rounds of training, comparing the accuracy and recall scores of different weights on the test set, and selecting the weight with the highest score as the human detection weight in the final fall detection model.

[0025] Furthermore, in step S3, in the fall detection model, the aspect ratio α is calculated using formula (2), specifically as follows:

[0026]

[0027] where α t is the aspect ratio of the human body detected in the image at the t-th frame, h t is the length of the human body detected in the image at the t-th frame, and w t is the width of the human body detected in the image at the t-th frame;

[0028] The center of gravity point offset d is calculated using formula (3), specifically as follows:

[0029] d t+1 = P t - P t+1 (3)

[0030] where d t+1 is the center of gravity offset of the human body detected in the image at the (t + 1)-th frame from the human body detected in the previous frame image, and P t and P t+1 are the center of gravity point positions of the human body detected in the image at the t-th frame and the (t + 1)-th frame respectively, including the position information of the vertical and horizontal coordinates;

[0031] By combining the aspect ratio and the center of gravity point offset and assigning different thresholds, new fall determination parameters are obtained, so as to judge whether the detected human body has fallen, specifically as follows:

[0032] When α t ≥ 1.1 and d t+1 ≥ 0.08 * w t it is judged as a fall;

[0033] Another situation is when the aspect ratios of the human body detected in two consecutive frames of images are both greater than a threshold, i.e., α t ≥1.5 and α t+1 ≥1.5, it is also judged as a fall; the above two thresholds are obtained through a large number of comparison tests of fall data.

[0034] Furthermore, in step S4, a real-time video is obtained as the video stream for fall detection and input into the fall detection model. The fall detection model first performs human target detection on the frame images obtained by processing the input video, and then performs fall detection on the recognized human targets. If it is determined that a human body has fallen in the real-time video, an alarm is triggered.

[0035] A device for real-time human fall detection and alarm based on video, including a camera, a video decoding and encoding device, and an edge computing chip;

[0036] Among them, the camera is used for image acquisition, the video decoding and encoding device is used for decoding and encoding the images collected by the camera, and the edge computing chip is used for training the fall detection model and real-time judging whether there is a human fall in the video.

[0037] Furthermore, the video decoding and encoding device is a professional Smart IP Camera SoC: Hi3516DV300, which processes the images collected by the camera into a video stream of 1920x1080@30fps and inputs it into the edge computing chip.

[0038] Furthermore, in the edge computing chip, a fall detection model is trained according to the publicly available human detection data set;

[0039] The edge computing chip processes the video stream input by the video decoding and encoding device into frame images and inputs them into the fall detection model. The fall detection model accelerates the calculation of whether there are human targets in the frame images and calculates whether the human targets have fallen. If it is determined that the target human body has fallen, the edge computing chip sends a signal to trigger an alarm.

[0040] Compared with the prior art, the advantages of the present invention are as follows:

[0041] The present invention specifically establishes a human detection data set for the fall human body model and prunes the model to reduce the model size, improve the detection accuracy and speed, and realize the deep learning fall detection of real-time video under local processing without the need for cloud computing, solving the problem of fall detection in multi-person scenarios without disclosing user privacy. At the same time, it overcomes the disadvantages of wearable fall detection devices such as resistance to wearing and battery life, and solves the problems of complex composition and high cost of environmental fall detection systems, improving usability and reducing detection costs. It can be applied to people and places prone to falls, improving the rescue efficiency for fall victims. Description of the Drawings

[0042] Figure 1 It is a schematic flowchart of a method for real-time human fall detection and alarm based on video in an embodiment of the present invention;

[0043] Figure 2 It is an operation flowchart of a device for real-time human fall detection and alarm based on video in an embodiment of the present invention;

[0044] Figure 3 It is a decision flowchart of a fall algorithm in an embodiment of the present invention;

[0045] Figure 4 It is a training flowchart of a fall detection model in an embodiment of the present invention;

[0046] Figure 5 It is a production flowchart of a fall data set in an embodiment of the present invention;

[0047] Figure 6 It is a structural diagram of a fall detection model in an embodiment of the present invention;

[0048] Figure 7 It is a composition structural diagram of a device for real-time human fall detection and alarm based on video in Embodiment 3 of the present invention;

[0049] Figure 8 It is a production flowchart of an elderly body posture fall data set in Embodiment 2 of the present invention. Detailed Embodiments

[0050] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following will combine the accompanying drawings and give examples to detail the specific implementation of the present invention.

[0051] Embodiment 1:

[0052] A method for real-time human fall detection and alarm based on video, as Figure 1 shown, includes the following steps:

[0053] S1. On the basis of a publicly available human detection data set, add human image information in the fall state to establish a fall detection data set, as Figure 5 shown;

[0054] Preliminarily screen the publicly available human detection data set, and eliminate useless data. Useless data refers to images with only partial hands and legs in the image and no human torso features. Images with 80% of the human body part features appearing in the picture are useful data;

[0055] Collect and capture multiple human fall videos, perform frame extraction on the human fall videos, extract 8 frames per second, manually select the pictures between the standing state and the fallen state on the ground, obtain a dataset with human fall postures, and perform data annotation on the selected pictures, and label all the selected pictures with the "person" label;

[0056] Perform data augmentation on the dataset with human fall postures. Generate 10 data-augmented pictures from the pictures in the dataset with human fall postures through image processing operations, including rotation, translation, and stretching, to obtain a dataset with human fall postures after data augmentation;

[0057] Merge the dataset with human fall postures after data augmentation with the publicly available human detection dataset after screening to establish a fall detection dataset.

[0058] S2. Prune and transform the YOLOv2-Tiny network, build a fall detection model, and train the fall detection model based on the self-established fall dataset;

[0059] Perform weight sparsification training on the original weights of the YOLOv2-Tiny network model, that is, introduce a scaling factor γ of the BN layer in each channel and multiply it by the output of the channel; then train the weights and the scaling factor, cut off the channels with small scaling factors, and fine-tune the pruned network to achieve the purpose of weight sparsification;

[0060] Among them, the channel pruning and fine-tuning are as follows:

[0061] After introducing the L1 regularization term of the scaling factor, the scaling factors in the obtained model will tend to 0; then sort the absolute values of the scaling factors first, and take the scaling factor at the 80% position of the scaling factors sorted from small to large as the threshold, and cut off the channels corresponding to the small scaling factor γ below the threshold. Cutting off the channels corresponding to the small scaling factor γ is essentially directly cutting off the convolution kernels corresponding to this channel; by doing so, a compact network with fewer parameters, less memory occupation during operation, and low computational complexity can be obtained: the Prune-YOLOv2-Tiny network model;

[0062] Change the category C in the detection layer of the Prune-YOLOv2-Tiny network model to a single category, that is, 1; at the same time, use the K-means algorithm to update the values of 5 anchors in the detection layer for the fall detection dataset. The values of the anchors are the width and height of the prediction box; the number of convolution kernels R in the last convolutional layer of the Prune-YOLOv2-Tiny network model is calculated using formula (1) and modified to the corresponding 30, as follows:

[0063] R = Anchors * (5 + C) (1)

[0064] Among them, Anchors is the number of prediction boxes in the Prune - YOLOv2 - Tiny network model, and the number is 5. The obtained R is the output channel size, which is used to obtain the tensor of the prediction output and generate the target box; thus, a fall detection model is obtained.

[0065] The fall detection model is essentially the Prune - YOLOv2 - Tiny network model, which consists of 16 layers and involves 3 types of layers: convolutional layers (9 layers), max - pooling layers (6 layers), and the final detection layer (the last 1 layer). Among them, the convolutional layers play a role in feature extraction, and the pooling layers are used for sampling and reducing the scale of the feature map. For RGB images of any resolution, each pixel is divided by 255 to be transformed into the interval [0, 1], and then scaled to 416×416 according to the original image's aspect ratio, and 0.5 is filled in the insufficient parts. The obtained 416×416×3 - sized array is input into the fall detection model, and after detection, a 13×13×55 - sized array is output. The values therein are the number of channels from 416×416 to 13×13.

[0066] As Figure 4 shown, by training on the fall detection model with the dataset, the weights in the fall detection model are obtained, specifically as follows:

[0067] The dataset is divided into a test set and a training set according to a 1:9 ratio. For each training image in the training set, the IOU loss, classification loss, and coordinate loss between the model prediction result and the true label are calculated on the fall detection model loaded with the pre - trained weights. Among them, the pre - trained weights are the pre - trained weights yolov2 - tiny.weights provided on the official website of the YOLO network model;

[0068] When the fall detection model reaches fitting, the weights of the fall detection model are saved, the learning rate is adjusted, and the next round of training is started.

[0069] Furthermore, under the same test set, by comparing the weights of the fall detection models obtained from different rounds of training, comparing the accuracy and recall scores of different weights on the test set, the weight with the highest score is selected as the human detection weight in the final fall detection model.

[0070] S3. By assigning different thresholds to the sensitivity of the aspect ratio α and the center - of - gravity point offset d of the images before and after human fall, new determination parameters for fall are obtained to achieve the judgment of fall;

[0071] In the fall detection model, the aspect ratio α is calculated using formula (2), specifically as follows:

[0072]

[0073] Among them, α t is the aspect ratio of the human body detected in the image at the t-th frame, h t is the length of the human body detected in the image at the t-th frame, and w t is the width of the human body detected in the image at the t-th frame;

[0074] The center of gravity point offset d is calculated using formula (3), specifically as follows:

[0075] d t+1 = P t - P t+1 (3)

[0076] Among them, d t+1 is the center of gravity offset of the human body detected in the image at the (t + 1)-th frame and the human body detected in the previous frame image, and P t and P t+1 are respectively the center of gravity point positions of the human body detected in the image at the t-th frame and the (t + 1)-th frame, including the position information of the vertical and horizontal coordinates;

[0077] By combining the aspect ratio and the center of gravity point offset and assigning different thresholds, new fall determination parameters are obtained, so as to judge whether the detected human body has fallen, specifically as follows:

[0078] When α t ≥ 1.1 and d t+1 ≥ 0.08 * w t , it is judged as a fall;

[0079] Another situation is when the aspect ratios of the human bodies detected in two consecutive frame images are both greater than a threshold, that is, α t ≥ 1.5 and α t+1 ≥ 1.5, it is also judged as a fall; The above two thresholds are obtained through comparison and testing of a large amount of fall data.

[0080] S4. As Figure 2 shown, obtain the real-time video as the video stream for fall detection and input it into the fall detection model. The fall detection model first performs human target detection on the frame images obtained by processing the input video, and then performs fall detection and alarm on the recognized human targets;

[0081] Obtain the real-time video as the video stream for fall detection and input it into the fall detection model. The fall detection model first performs human target detection on the frame images obtained by processing the input video, and then performs fall detection on the recognized human targets. If it is determined that someone has fallen in the real-time video, an alarm will be issued.

[0082] Example 2: Compared with Example 1, the VOC and COCO datasets on the Internet are used, such as Figure 8As shown, the falling shape is not added in Example 2. Although the data set is large and the human body features are extensive, the falling human target can be identified without adding the falling shape, but there are human targets in the public data set that only have hands and feet that are recognized as people, which can easily lead to misjudgment of falls.

[0083] Embodiment 3:

[0084] A device based on real-time video human fall detection and alarm, including a camera, a video decoding and encoding device, an edge computing chip, a cooling fan and a shell; the following tail wires are reserved: 5V1A power output, network port, and 12V2A input. The device power supply, cooling fan and shell are provided by the FPGA board, including the whole machine reset circuit, and the 5V1A output in the tail wire. The 5V1A power output provides fan power for chip heat dissipation, and the network port is used for network communication data exchange. Figure 7 shown.

[0085] Among them, the camera is used for image acquisition, the video decoding and encoding device is used to decode and encode the images captured by the camera, and the edge computing chip is used to train the fall detection model and determine in real time whether there is a human fall in the video.

[0086] The video decoding and encoding device is a professional Smart IP Camera SoC: Hi3516DV300, which processes the images collected by the camera into a 1920x1080@30fps video stream and inputs it into the edge computing chip.

[0087] In the edge computing chip, a fall detection model is trained based on the public human body detection dataset;

[0088] The edge computing chip processes the video stream transmitted by the video decoding and encoding device into frame images and inputs them into the fall detection model. The fall detection model accelerates the calculation of whether there is a human target in the frame image, and calculates whether the human target has fallen. If it is determined that the target person has fallen, the device will transmit the data through the network port connected to the network cable, and inform relatives or nearby social workers through SMS, WeChat, email, telephone and other communication methods to achieve the purpose of timely rescue of the fallen person.

Claims

1. A method for real-time human fall detection and alarm based on video, characterized in that, The following steps are involved: S1. Based on the public human body detection dataset, add human body image information in the fall state to establish a fall detection dataset; S2. Prune the YOLOv2-Tiny network, build a fall detection model, and train the fall detection model based on the self-established fall data set; perform weight sparse training on the original weights of the YOLOv2-Tiny network model, that is, introduce a scaling factor γ introduced into the BN layer in each channel and multiply it with the output of the channel; then train the weights and scaling factors, prune the channels with small scaling factors, and fine-tune the pruned network to achieve the purpose of weight sparseness; Among them, channel pruning and fine-tuning are as follows: After introducing the scaling factor L1 regularization term, the scaling factors in the obtained model will tend to 0; then sort the absolute values ​​of the scaling factors first, take the scaling factors at 80% of the positions in the scaling factors sorted from small to large as the threshold, cut off the channels corresponding to the small scaling factors γ below the threshold, and obtain the Prune-YOLOv2-Tiny network model; Change the category C in the detection layer of the Prune-YOLOv2-Tiny network model to a single category, that is, 1; at the same time, use the K-means algorithm to update the values ​​of the 5 anchors in the detection layer for the fall detection dataset. The values ​​of the anchors are the width and height of the prediction box; the number of convolution kernels R in the last convolution layer of the Prune-YOLOv2-Tiny network model is calculated using formula (1) and modified to the corresponding 30, as follows: R=Anchors*(5+C) (1) Among them, Anchors is the number of prediction boxes in the Prune-YOLOv2-Tiny network model, which is 5. The obtained R is the channel size of the output, which is used to obtain the tensor of the predicted output and generate the target box; thus, the fall detection model is obtained; S3, by assigning different thresholds to the sensitivity of the image aspect ratio α and the center of gravity offset d before and after the human body falls, a new fall judgment parameter is obtained to realize the fall judgment; S4. Real-time video is obtained as a video stream for fall detection and input into the fall detection model. The fall detection model first performs human target detection on the frame image obtained by processing the input video, and then performs fall detection and alarm on the identified human target.

2. The method for real-time human fall detection and alarm based on video according to claim 1, wherein In step S1, the public human detection data set is preliminarily screened to remove useless data. Useless data refers to images with only partial hands and legs and no human torso features. Images with 80% of human body features appearing on the screen are useful data. Collect and shoot multiple videos of people falling, capture frames of the videos, capture multiple frames per second, manually filter out pictures between standing and falling, obtain a data set with human falling postures, and annotate the filtered pictures, labeling the filtered pictures with person tags; Perform data augmentation on the dataset with human fall postures. Generate multiple data-augmented images from the images in the dataset with human fall postures through image processing operations, including rotation, translation, and stretching, to obtain a data-augmented dataset with human fall postures. Merge the data-augmented dataset with human fall postures and the screened publicly available human detection dataset to establish a fall detection dataset.

3. A method for real-time human fall detection and alarm based on video, as claimed in claim 2, wherein, In step S2, train using the dataset on the fall detection model to obtain the weights in the fall detection model, specifically as follows: Divide the dataset into a test set and a training set in a ratio of 1:

9. For each training image in the training set, calculate the IOU loss, classification loss, and coordinate loss between the model prediction result and the true label on the fall detection model loaded with pre-trained weights, where the pre-trained weights are existing weights. When the fall detection model reaches convergence, save the weights of the fall detection model, adjust the learning rate, and start the next round of training.

4. A method for real-time human fall detection and alarm based on video, according to claim 3, characterized in that Under the same test set, compare the weights of the fall detection models obtained from different rounds of training, compare the accuracy and recall scores of different weights on the test set, and select the weight with the highest score as the human detection weight in the final fall detection model.

5. A method for real-time human fall detection and alarm based on video, characterized in that, In step S3, in the fall detection model, the aspect ratio α is calculated using formula (2), specifically as follows: Among them, α t is the aspect ratio of the human body detected in the image at the t-th frame, h t is the length of the human body detected in the image at the t-th frame, and w t is the width of the human body detected in the image at the t-th frame; The center of gravity point offset d is calculated using formula (3), specifically as follows: d t+1 = P t -P t+1 (3) Among them, d t+1 is the centroid offset of the human body detected in the (t + 1)-th frame from that detected in the previous frame, and P t and P t+1 are the centroid positions of the human body detected in the t-th and (t + 1)-th frames respectively, including the position information of the ordinate and abscissa; Combine the aspect ratio and the center of gravity point offset and assign different thresholds to obtain new determination parameters for falls, thereby realizing the judgment of whether the detected human has fallen, specifically as follows: When α t ≥ 1.1 and d t+1 ≥ 0.08 * w t it is judged as a fall; Another situation is when the aspect ratios of the human body detected in two consecutive frames are both greater than a threshold, i.e., α t ≥ 1.5 and α t+1 ≥ 1.5, it is also judged as a fall; the above two thresholds are obtained through comparison and testing of a large amount of fall data.

6. A method for real-time human fall detection and alarm based on video, according to any one of claims 1 to 5, characterized in that In step S4, obtain a real-time video as the video stream for fall detection and input it into the fall detection model. The fall detection model first performs human target detection on the frame images obtained by processing the input video, and then performs fall detection on the recognized human targets. If it is determined that there is a human fall in the real-time video, an alarm is issued.

7. A device for real-time human fall detection and alarm based on video, characterized in that, It includes a camera, a video decoding and encoding device, and an edge computing chip; Among them, the camera is used for image acquisition, the video decoding and encoding device is used for decoding and encoding the images collected by the camera, and the edge computing chip is used to train the edge computing chip to train the fall detection model as described in claim 1 and to determine in real time whether there is a human fall in the video.

8. A device for real-time human fall detection and alarm based on video, according to claim 7, characterized in that, The video decoding and encoding device is a professional Smart IP Camera SoC: Hi3516DV300, and processes the images collected by the camera into a video stream of 1920x1080@30fps and inputs it into the edge computing chip.

9. An apparatus for real-time human fall detection and alarm based on video, as claimed in claim 7, wherein, In the edge computing chip, a fall detection model is trained based on the publicly available human detection dataset; The edge computing chip processes the video stream input by the video decoding and encoding device into frame images and inputs them into the fall detection model. The fall detection model accelerates the calculation of whether there are human targets in the frame images and calculates whether the human targets have fallen. If it is determined that the target human has fallen, the edge computing chip sends a signal to issue an alarm.

Citation Information

Patent Citations

  • Old man falling-down detection smart bracelet and old man falling-down detection method

    CN108041772A

  • Tumble detection floor and method

    CN111538264A

  • Indoor fall detection method and device for old people based on FPGA and deep learning

    CN113408485A

  • Fall detection method based on intelligent mobile terminal

    WO2021212883A1