A method, system, device and storage medium for fall detection of the elderly at home based on a high-resolution neural network model
By adopting a high-resolution neural network model in fall detection, combining spatial and temporal convolution, and the Coordinate Attention mechanism, the problem of inability to effectively handle timing features in the existing technology is solved, and the accuracy of fall detection is significantly improved.
Patent Information
- Application Number
- CN202310128991.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-02-17
AI Technical Summary
The prior art cannot effectively handle the timing characteristics in continuous actions in fall detection, resulting in poor detection effect.
Using a high-resolution neural network model, the spatial and temporal features of the video are captured through spatial convolution and temporal convolution, and combined with the Coordinate Attention mechanism, the position, direction and time perception information are captured to generate a high-resolution timing feature map.
On the basis of maintaining high resolution, it can effectively deal with timing problems and improve the accuracy and effectiveness of fall detection.
Smart Images

Figure CN116052058B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, system, device and storage medium for detecting falls of the elderly at home based on a high-resolution neural network model, and belongs to the technical field of computer vision abnormal behavior detection. Background Art
[0002] In daily life, there are mainly two reasons for human falls. One is being tripped or slipping due to inconvenient legs and feet, and the other is falling caused by diseases. If timely assistance is not obtained during a fall, the injury often gets worse and even results in the loss of life. Therefore, it is particularly important to detect human falls.
[0003] The fall detection task based on computer vision is position-sensitive. In order to make the obtained position information more accurate, first of all, the most straightforward method is to maintain a high-resolution feature map. Most existing methods encode the input from a high-resolution representation to a low-resolution representation and then restore it to a high-resolution representation, such as SegNet (a deep convolutional encoder-decoder architecture for image segmentation), DeconvNet (Deconvolution network), U-Net (a fully convolutional neural network), etc. Although this approach obtains a high resolution, a large amount of effective information is lost during the continuous upsampling and downsampling processes, thus affecting the final experimental results. HRNet (High-Resolution Net) parallelly connects the convolutional streams from high resolution to low resolution and always maintains a high resolution, so the representation is more accurate spatially. At the same time, HRNet repeats multi-resolution fusion to improve the high-resolution representation with the help of the low-resolution representation, thereby obtaining strong semantic information. Secondly, the method of using the Attention mechanism can be adopted. The essence of the attention mechanism is to locate the information of interest and suppress the useless information, enabling the neural network model to notice what it needs to notice when doing a specific thing. Compared with the SE (Sequeeze-and-Excitation) and CBAM (Convolutional Block AttentionModule) attention methods, coordinate attention has more advantages. It can not only capture cross-channel information but also capture direction-aware and position-aware information, which helps the model more accurately locate and identify the target of interest.
[0004] However, in practice, falling is a dynamic and continuous process. The falling detection task not only needs to analyze the features of each frame of the video, but also needs to mine clues from the temporal information between video frames, expanding the feature extraction problem from two-dimensional space to three-dimensional space-time. HRNet and coordinate attention are a high-resolution network and an attention mechanism for processing image spatial features, both of which process each image independently and cannot capture long-term temporal evolution. Therefore, they cannot process the temporal features in continuous actions, resulting in poor performance in dealing with the falling detection task. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, system, device and storage medium for detecting the falls of the elderly at home based on a high-resolution neural network model, so as to solve the problem of poor falling detection effect caused by the inability to process the temporal features in continuous actions in the prior art.
[0006] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0007] In the first aspect, the present invention provides a method for detecting the falls of the elderly at home based on a high-resolution neural network model, including:
[0008] Obtain the video to be detected;
[0009] Perform size preprocessing on the video;
[0010] Input the video after size preprocessing into the trained high-resolution neural network model to output the falling detection result.
[0011] Combined with the first aspect, further, the size preprocessing includes:
[0012] Set the size of the video to a preset size.
[0013] Combined with the first aspect, further, the high-resolution neural network model includes a preprocessing module, and the preprocessing module includes a spatial convolution and a temporal convolution;
[0014] In the preprocessing module:
[0015] First, capture the spatial features of the video through spatial convolution and reduce the size of the feature map finally generated by the preprocessing module;
[0016] Then input the features output by the spatial convolution into the temporal convolution to capture the temporal features of the video, thereby obtaining the preprocessing feature map.
[0017] Combined with the first aspect, further, the high-resolution neural network model includes an attention module, and the attention module is Coordinate Attention;
[0018] In the attention module:
[0019] Encode the channels of the input pre-processed feature map along the horizontal, vertical, and temporal directions respectively through a pooling kernel to obtain intermediate feature maps of horizontal direction perception, vertical direction perception, and temporal direction perception;
[0020] Perform a concatenate operation on the three intermediate feature maps, and then use a convolution transformation function for transformation operations to obtain three intermediate feature maps;
[0021] Decompose the three intermediate feature maps into three separate tensors after passing through the BN and ReLU activation functions, restore the three tensors to the same number of channels as the input pre-processed feature map through convolution respectively, and finally weight the three tensors with the number of channels restored through the Sigmoid activation function to obtain the output features.
[0022] Combined with the first aspect, further, the high-resolution neural network model includes a backbone network, the backbone network includes four stages, the first stage includes a sub-network of the first level, the second stage includes sub-networks of the first level and the second level in parallel, the third stage includes sub-networks of the first level and the second level in parallel, and the fourth stage includes sub-networks of the first level, the second level, and the third level; among them, the resolutions of the sub-networks of the first level, the second level, and the third level decrease in turn; an exchange unit across parallel sub-networks is also provided between each stage to enable each sub-network to repeatedly receive information from other parallel sub-networks, and a high-resolution feature map is output in the fourth stage.
[0023] Combined with the first aspect, further, the high-resolution neural network model includes a Softmax classifier. In the Softmax classifier:
[0024] Calculate the score values of the high-resolution feature map output by the input backbone network belonging to the two categories of fall and non-fall respectively, and then use the Sigmoid function to map all the score values into a probability value, and select the category with the largest probability value as the fall detection result.
[0025] Combined with the first aspect, further, the high-resolution neural network model uses a binary cross-entropy loss function with an L2 regularization term as the target loss function and is trained through an Adam optimizer; the target loss function is:
[0026]
[0027] where L is the target loss function, m is the number of samples, a is the sample subscript, is the sample label, where the positive class is 1 and the negative class is 0, y a is the probability predicted as positive, λ‖θ‖ 2 is the L2 regularization term.
[0028] In a second aspect, the present invention also provides a home elderly fall detection system based on a high-resolution neural network model, including:
[0029] Video acquisition module: used to acquire the video to be detected;
[0030] Video preprocessing module: used to perform size preprocessing on the video;
[0031] Fall detection module: used to input the video after size preprocessing into the trained high-resolution neural network model and output the fall detection result.
[0032] In a third aspect, the present invention also provides a home elderly fall detection device based on a high-resolution neural network model, including a processor and a storage medium;
[0033] The storage medium is used to store instructions;
[0034] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first aspects.
[0035] In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to any one of the first aspects.
[0036] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0037] A home elderly fall detection method, system, device and storage medium provided by the present invention designs a high-resolution neural network model, enabling it to handle temporal problems while maintaining high resolution, and can well meet the high-resolution requirements of the fall detection task, improving the accuracy of fall detection;
[0038] Expand Coordinate Attention, which can not only capture direction-aware and position-aware information, but also capture time-aware information, so that Coordinate Attention can focus on the position temporal information in the dynamic continuous process and improve the effect of the model in handling the fall detection task;
[0039] In the backbone network, let multiple parallel network branches with different resolutions exchange information, so that the output high-resolution feature map contains more information and improves the accuracy of fall detection. Description of the Drawings
[0040] Figure 1 It is a flowchart of a method for detecting falls of the elderly at home based on a high-resolution neural network model provided by an embodiment of the present invention;
[0041] Figure 2 It is a framework diagram of a high-resolution neural network model provided by an embodiment of the present invention;
[0042] Figure 3 It is a schematic structural diagram of Coordinate Attention provided by an embodiment of the present invention;
[0043] Figure 4 It is a schematic structural diagram of the backbone network provided by an embodiment of the present invention. Specific embodiments
[0044] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0045] Embodiment 1
[0046] As Figure 1 shown, a method for detecting falls of the elderly at home based on a high-resolution neural network model provided by an embodiment of the present invention includes the following steps:
[0047] S1. Obtain the video to be detected.
[0048] The source of the video to be detected is generally a camera device installed at home. After the camera device captures the video, it is uploaded to the storage unit, and the video within the time period to be detected can be obtained from the storage unit.
[0049] S2. Perform size preprocessing on the video.
[0050] Set the size of the video to a preset size through image resampling to facilitate processing by the neural network.
[0051] S3. Input the video after size preprocessing into the trained high-resolution neural network model and output the fall detection result.
[0052] In the trained high-resolution neural network model (the structure is as Figure 2 shown), the following processes are performed:
[0053] First, the pre - processed video is input into the pre - processing module of 3D HRNet. The pre - processing module consists of spatial convolution and temporal convolution, generating a high - resolution pre - processed feature map. Specifically, the input video first passes through the spatial convolution composed of Conv 2d (two - dimensional convolution operation), BN (Batch Normalization), and ReLU (Rectified Linear Unit) to capture the spatial features of the input video (captured by the spatial convolution), while reducing the size of the feature map finally generated by the pre - processing module (implemented by Conv 2d). Then, the features output by the spatial convolution are input into the temporal convolution composed of BN, ReLU, Conv 2d, and Dropout (random inactivation) to capture the temporal features of the input video, obtaining a high - resolution pre - processed feature map.
[0054] The above - mentioned pre - processed feature map is input into Figure 3 the Coordinate Attention (a new attention mechanism designed for lightweight networks proposed by Qibin Hou et al. from the National University of Singapore, which embeds position information into channel attention). Coordinate Attention encodes channel relationships and long - term dependencies through precise position information. The specific operation is divided into two steps: Coordinate information embedding and Coordinate Attention generation. Specifically, given the input, first, the pooling kernel is used to encode the channels along the horizontal, vertical, and temporal directions respectively, obtaining intermediate feature maps of horizontal perception, vertical perception, and temporal perception. Then, the above - mentioned intermediate feature maps are concatenated, and then a convolution transformation function is used to perform a transformation operation on them, obtaining an intermediate feature map that encodes position information in the horizontal, vertical, and temporal directions. This intermediate feature map can be expressed as:
[0055] f = δ(F([z h ,z w ,z t ))
[0056] where f represents the intermediate feature map, δ represents the activation function, F represents the concatenate operation, zh represents the feature of the channel along the h direction, zw represents the feature of the channel along the w direction, and zt represents the feature of the channel along the t direction.
[0057] The intermediate feature map is decomposed into three separate tensors after passing through the BN (Batch Normalization) and ReLU (Rectified Linear Unit) activation functions. The three tensors are respectively restored to the same number of channels as the input preprocessed feature map through convolution. Finally, the three tensors with restored channels are weighted by the Sigmoid activation function as attention weights. The output feature of Coordinate Attention can be written as:
[0058] y = x × g h × g w × g t
[0059] where y represents the output feature, x represents the input feature, and g h represents the weight of the channel along the h direction, and g w represents the weight of the channel along the w direction, and g t represents the weight of the channel along the t direction.
[0060] Then, the above output feature is input into the backbone network of 3D HRNet, as Figure 4 shown. The backbone network contains four stages. stage1 to stage4 respectively represent the first stage to the fourth stage. The first stage includes the sub-network of the first level, the second stage includes the sub-networks of the first level and the second level in parallel, the third stage includes the sub-networks of the first level and the second level in parallel, and the fourth stage includes the sub-networks of the first level, the second level, and the third level, which is the output stage. Among them, the resolutions of the sub-networks of the first level, the second level, and the third level decrease in turn. Figure 4 The downward diagonal arrow in it represents the resolution decrease, and the upward diagonal arrow represents the resolution increase. Figure 4 The sub-networks in the first row in it are all sub-networks of the first level, the sub-networks in the second row are all sub-networks of the second level, and the sub-networks in the third row are all sub-networks of the third level. At the same time, in the backbone network, the outputs of the parallel sub-networks are fused with each other and then input into the sub-networks of the next stage, so that the sub-networks of the next stage receive information from the parallel sub-networks of the previous stage, so that the output high-resolution feature map contains more information.
[0061] The high-resolution feature map output by the backbone network is input into the Softmax classifier to calculate the score values of the high-resolution feature map belonging to the two categories of fall and non-fall respectively. Then, the Sigmoid function is used to map all the score values into a probability value, and the category with the largest probability value is selected as the fall detection result.
[0062] In this embodiment, the high-resolution neural network model uses the binary cross-entropy loss function with L2 regularization term (to further avoid model overfitting) as the target loss function, and uses the Adam optimizer to train the optimal model, which is suitable for solving problems with large-scale data or parameters, and has high computational efficiency and low memory requirements; the target loss function containing L2 regularization is as follows:
[0063]
[0064] where L is the target loss function, m is the number of samples, a is the sample subscript, is the sample label, the positive class is 1, the negative class is 0, and y a is the probability predicted as positive, and λ‖θ‖ 2 is the L2 regularization term.
[0065] Embodiment 2
[0066] The embodiment of the present invention provides a home elderly fall detection system based on a high-resolution neural network model, including:
[0067] Video acquisition module: used to acquire the video to be detected;
[0068] Video preprocessing module: used to perform size preprocessing on the video;
[0069] Fall detection module: used to input the video after size preprocessing into the trained high-resolution neural network model and output the fall detection result.
[0070] Embodiment 3
[0071] The embodiment of the present invention provides a home elderly fall detection device based on a high-resolution neural network model, including a processor and a storage medium;
[0072] The storage medium is used to store instructions;
[0073] The processor is used to operate according to the instructions to execute the steps of the following method:
[0074] Acquire the video to be detected;
[0075] Perform size preprocessing on the video;
[0076] Input the video after size preprocessing into the trained high-resolution neural network model and output the fall detection result.
[0077] Embodiment 4
[0078] The embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the following method are implemented:
[0079] Obtain the video to be detected;
[0080] Perform size preprocessing on the video;
[0081] Input the video after size preprocessing into the trained high-resolution neural network model to output the fall detection result.
[0082] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0083] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be realized by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0084] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for realizing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0086] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for detecting falls of the elderly at home based on a high-resolution neural network model, characterized in that Including: Obtain the video to be detected; Perform size preprocessing on the video; Input the video after size preprocessing into the trained high-resolution neural network model to output the fall detection result; The high-resolution neural network model includes an attention module, and the attention module is CoordinateAttention; In the attention module: Encode the channels separately along the horizontal, vertical, and temporal directions on the input preprocessed feature map through a pooling kernel to obtain intermediate feature maps with horizontal direction perception, vertical direction perception, and temporal direction perception; Perform a concatenate operation on the three intermediate feature maps, and then use a convolutional transformation function for transformation operations to obtain three intermediate feature maps; Decompose the three intermediate feature maps into three separate tensors after passing through the BN and ReLU activation functions, restore the channels of the three tensors to the same number of channels as the input preprocessed feature map through convolution respectively, and finally weight the three tensors with restored channel numbers through the Sigmoid activation function to obtain the output features; The high-resolution neural network model includes a backbone network, and the backbone network includes four stages. The first stage includes a sub-network of the first level, the second stage includes sub-networks of the first level and the second level in parallel, the third stage includes sub-networks of the first level and the second level in parallel, and the fourth stage includes sub-networks of the first level, the second level, and the third level; among them, the resolutions of the sub-networks of the first level, the second level, and the third level decrease in turn; at the same time, in the backbone network, the outputs of the parallel sub-networks are fused with each other and then input into the sub-networks of the next stage, so that the sub-networks of the next stage receive information from the parallel sub-networks of the previous stage.
2. The method for detecting falls of the elderly at home based on a high-resolution neural network model according to claim 1, characterized in that, The size preprocessing includes: Set the size of the video to a preset size.
3. A method for detecting falls of the elderly at home based on a high-resolution neural network model according to claim 1, characterized in that, The high-resolution neural network model includes a preprocessing module, and the preprocessing module includes spatial convolution and temporal convolution; In the preprocessing module: First, capture the spatial features of the video through spatial convolution and reduce the size of the feature map finally generated by the preprocessing module; Then input the features output by the spatial convolution into the temporal convolution to capture the temporal features of the video, thereby obtaining the preprocessed feature map.
4. A method for detecting falls of the elderly at home based on a high-resolution neural network model according to claim 1, characterized in that, The high-resolution neural network model includes a Softmax classifier. In the Softmax classifier: Calculate the score values of the high-resolution feature map output by the input backbone network belonging to the two categories of fall and non-fall respectively, then use the Sigmoid function to map all the score values into a probability value, and select the category with the largest probability value as the fall detection result.
5. A method for detecting falls of the elderly at home based on a high-resolution neural network model according to claim 1, characterized in that, The high-resolution neural network model uses a binary cross-entropy loss function with an L2 regularization term as the target loss function and is trained through an Adam optimizer; the target loss function is: ; Among them, is the target loss function, is the number of samples, is the sample subscript, is the sample label, where the positive class is 1 and the negative class is 0, is the probability predicted as positive, is the L2 regularization term.
6. A fall detection system for the elderly at home based on a high-resolution neural network model, characterized in that, Including: Video acquisition module: used to obtain the video to be detected; Video preprocessing module: used to perform size preprocessing on the video; Fall detection module: used to input the video after size preprocessing into the trained high-resolution neural network model to output the fall detection result; Among them, the high-resolution neural network model includes an attention module, and the attention module is CoordinateAttention; In the attention module: Encode the channels of the input preprocessed feature map along the horizontal, vertical, and temporal directions respectively through a pooling kernel to obtain intermediate feature maps with horizontal direction perception, vertical direction perception, and temporal direction perception; Perform a concatenate operation on the three intermediate feature maps, and then use a convolutional transformation function for transformation operations to obtain three intermediate feature maps; Decompose the three intermediate feature maps into three separate tensors after passing through the BN and ReLU activation functions, restore the channels of the three tensors to the same number of channels as the input preprocessed feature map respectively through convolution, and finally weight the three tensors with restored channel numbers through the Sigmoid activation function to obtain the output features; Among them, the high-resolution neural network model includes a backbone network, and the backbone network includes four stages. The first stage includes a sub-network of the first level, the second stage includes sub-networks of the first level and the second level in parallel, the third stage includes sub-networks of the first level and the second level in parallel, and the fourth stage includes sub-networks of the first level, the second level, and the third level; among them, the resolutions of the sub-networks of the first level, the second level, and the third level decrease in sequence; at the same time, in the backbone network, the outputs of the parallel sub-networks are fused with each other and input into the sub-networks of the next stage, so that the sub-networks of the next stage receive information from the parallel sub-networks of the previous stage.
7. A device for detecting falls of the elderly at home based on a high-resolution neural network model, characterized in that, It includes a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Behavior recognition method and device and storage medium
CN114764902A
Method and system for motion analysis and fall prevention
US20170352240A1