A Deep 3D Convolutional Neural Network Detection Method for Identifying Fall Events in Motion Behaviors
By adopting a deep three-dimensional convolutional neural network twin network in motion scenarios, efficient and accurate detection of fall events during motion is achieved, and the problem of insufficient accuracy of fall detection in motion scenarios in the prior art is solved.
Patent Information
- Application Number
- CN202111291424.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The existing fall detection methods are insufficient in motion scenarios, especially the two-dimensional convolutional neural network-based methods are difficult to effectively identify fall events during motion.
A deep three-dimensional convolutional neural network twin network is used to monitor 6-second video clips every 3 seconds, extract 18 frames of images and perform feature fusion, generate fall probability judgment through the three-dimensional convolutional sub-neural network, and output fall probability using the sigmoid function.
It realizes efficient and accurate fall event detection in sports scenarios, and improves the recognition rate of fall events.
Smart Images

Figure CN113920588B_ABST
Abstract
Description
1. Technical Field
[0001] Object motion classification, fall detection, computer vision, artificial intelligence 2. Background Art
[0002] 2.1 Introduction to General Technical Methods
[0003] The convolutional neural network is a method of extracting features by sliding a convolutional kernel on an image or features, and it is a widely used technology.
[0004] The ResNet convolutional neural network [1] is a popular method for feature extraction of convolutional neural networks.
[0005] 2.2 Introduction to Similar Methods
[0006] Existing human fall detection is generally divided into two categories.
[0007] The first category is to judge through picture information.
[0008] The second category is to judge through video information.
[0009] Each category can be specifically divided according to the methods adopted. For example:
[0010] Solution 1. Detect and obtain the key points of the human body skeleton, judge the posture of the human body, and then judge whether it has fallen.
[0011] Solution 2. Extract features through a convolutional neural network and use the features to determine whether a person has fallen.
[0012] Therefore, all methods can be summarized as: the fall determination method based on the key points of the human body skeleton in pictures, the fall determination method based on the neural network feature extraction in pictures, the fall determination method based on the key points of the human body skeleton in videos, and the fall determination method based on the neural network feature extraction in videos.
[0013] The method of this application belongs to the fall determination method based on the neural network feature extraction in videos. However, compared with the existing methods, the method of this application uses a three-dimensional convolutional neural network for feature fusion and fall determination, while the existing methods only use two-dimensional models.
[0014] At the same time, because the training data of this application is collected from motion data, the proposed method is more applicable to fall determination during exercise; while the existing methods are more applicable to daily life scenarios. 3. Summary of the Invention
[0015] Based on the integration of innovative technologies and existing methods, this patent realizes an end-to-end method for detecting the fall events of people in motion. The method of this application monitors the next 6-second video clip every 3 seconds, and extracts 18 video images at equal intervals in the 6-second video clip; inputs the 18 images into a specific deep three-dimensional convolutional neural network Siamese network, and obtains the determination probability of whether someone has fallen through this network; determines that a fall event has occurred within this time segment for video clips with a probability greater than 0.5.
[0016] The content of this invention application mainly includes the specific steps of dealing with falls and the innovative three-dimensional convolutional feature fusion and extraction algorithms used in the steps. IV. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the structural diagram of the neural network of the method of this application.
[0018] This method achieves the purpose of fall event monitoring by judging whether a fall event occurs in each video clip.
[0019] This method extracts 18 images at equal intervals in the 6-second video clip. According to the network structure of ResNet-50 in reference [1], intermediate features are extracted for each picture, each with a dimension of 2048x7x7. The 18 images are sequentially spliced into a feature of 18x2048x7x7 dimensions and deformed into 18x2048x49.
[0020] These features will be input into the three-dimensional convolutional sub-neural network. The first convolutional kernel is 18x3x3 in size. After convolution, the feature size is 2048x49. After deforming into 2048x7x7, a convolutional kernel of 3x3 size is used to generate 2048 feature maps of 7x7 size, and the global average pooling method is used to generate one feature for each feature map, forming a 2048-dimensional feature.
[0021] Finally, this 2048-dimensional feature uses the sigmoid function to generate a probability in the range of [0, 1] for evaluating whether someone has fallen in this video clip. V. SPECIFIC IMPLEMENTATION MANNER
[0022] This application realizes the automatic determination of the fall of a moving human body through 3 steps.
[0023] Step 1: Frame extraction
[0024] Monitor the next 6-second video clip every 3 seconds, and extract 18 video images at equal intervals in the 6-second video clip.
[0025] Step 2: Intermediate feature extraction
[0026] The 18-frame video image is input into 18 sub-branches to form intermediate features.
[0027] Step 3: Use 3D convolution for feature fusion and extraction
[0028] The intermediate features are input into the 3D convolutional sub-neural network module proposed by the method of this application to generate the final determination features, and the sigmoid function is used to obtain the fall probability.
[0029] References:
[0030] [1]Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun: Deep ResidualLearning for Image Recognition.CVPR 2016: 770-778
Claims
1. A deep three-dimensional convolutional network detection method for human fall events during movement, characterized in that The following steps: (1) Monitor the next 6-second video clip every 3 seconds, and extract 18 video images at equal intervals within the 6-second video clip; (2) Input the 18 images into a specific deep three-dimensional convolutional neural network Siamese network, and obtain the determination probability of whether someone has fallen through this network, and its value range is from 0 to 1; Determine that a fall event has occurred within the time segment for the video clip with a probability greater than 0.
5. For the specific deep three-dimensional convolutional neural network Siamese network, there are 18 two-dimensional convolutional branches at the input of the neural network. Each branch is a two-dimensional feature extraction sub-network with a ResNet-50 structure, and each branch outputs a feature vector with a dimension of 2048x7x7; The features output by all two-dimensional convolutional branch sub-networks are concatenated to form a feature with a size of 18x2048x7x7, and after feature stitching, it is input into a three-dimensional convolutional sub-network with a feature dimension of 18x2048x49. The first convolutional kernel of the three-dimensional convolutional sub-network is an 18x3x3 3D convolutional kernel. After convolution, the feature size is 2048x49. After deformation into 2048x7x7, a 2D convolutional kernel with a size of 3x3 is used to generate 2048 feature maps with a size of 7x7, and the global average pooling method is used. Each feature map generates a feature to form a 2048-dimensional feature. Finally, the sigmoid function is used to give the probability determination of whether someone has fallen.
Citation Information
Patent Citations
Method for identifying fall behavior of old people based on deep learning
CN108549841A
Tumble detection method and system based on space-time hybrid convolutional network
CN110942009A