Human body tumble identification method and device
By obtaining the motion characteristics of the key points of the human body and building a deep learning model, and extracting the multi-level features of time series features, the existing technology's shortcomings in timing dependency modeling, complex environment adaptability and real-time performance are solved, and higher robustness and real-time performance are achieved in fall recognition.
Patent Information
- Application Number
- CN202510305209.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
The existing visual fall recognition technology has shortcomings in timing dependency modeling, complex environment adaptability and real-time performance, making it difficult to design a more robust and real-time fall recognition method.
Time series features are obtained by obtaining the motion characteristics of key points of the human body and processing these characteristics based on the sliding window. Then, a deep learning model is constructed based on time encoding and dynamic discarding techniques, multi-level features of time series features are extracted, and these features are finally input into a fully connected classifier to identify the human body's fall results.
It improves the recognition accuracy of the model in fall detection, enhances robustness and real-timeness, and can more effectively identify human falls in complex environments.
Smart Images

Figure CN120236328A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fall recognition, and in particular to a method and device for human fall recognition. Background Art
[0002] As the aging process accelerates, falls have become one of the main causes of accidental injuries among the elderly. Traditional fall recognition technologies are roughly divided into wearable devices, visual sensors, and environmental sensors. However, the wearable device method is limited by the comfort and dependence of the device, and the environmental sensor method is sensitive to external conditions.
[0003] Methods based on visual sensors have attracted widespread attention due to their non-invasiveness and convenience. However, existing visual methods have shortcomings in temporal dependency modeling, adaptability to complex environments, and real-time performance.
[0004] Therefore, how to design a more robust and real-time fall recognition method is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of this, it is necessary to provide a human fall recognition method and device to improve the robustness and real-time performance of the fall recognition method.
[0006] In order to solve the above problems, the present invention provides a method for identifying a human fall, comprising:
[0007] Acquire motion features of key points of a human body, and process the motion features based on a sliding window to obtain time series features;
[0008] Building a deep learning model based on time encoding and dynamic discarding, and extracting multi-level features of the time series features based on the deep learning model;
[0009] The multi-level features are input into a fully connected classifier to identify human fall results.
[0010] In a possible implementation, the step of obtaining motion features of key points of a human body includes:
[0011] The coordinates of the key points of the human body are obtained, and the motion features of the key points of the human body are calculated based on the coordinates of the key points of the human body.
[0012] In a possible implementation, the processing of the motion features based on a sliding window to obtain time series features includes:
[0013] The motion features are input into a sliding window for mean processing to obtain the time series features, wherein the length and step length of the sliding window can be adaptively adjusted based on the time length of human behavior.
[0014] In a possible implementation, the multi-level features of the time series features include time embedding sequence features, short-term dependence features, position time series features, and global dependence features. Extracting the multi-level features of the time series features based on the deep learning model includes:
[0015] Performing non-linear embedding on the time series features based on the Time2Vec layer to obtain time embedding sequence features;
[0016] Extracting short-term dependence features of the time series features based on the LSTM layer;
[0017] Performing position embedding on the time series features based on the position encoding layer, and adding the features after position embedding to the short-term dependence features to obtain position time series features;
[0018] Extracting global dependence features of the time series features based on the Transformer encoder, and dynamically adjusting the dropout rate of the time series features in the Transformer encoder based on the dynamic Dropout layer.
[0019] In a possible implementation, the formula of the Time2Vec layer is:
[0020] T(t) = sin(W·X + b)
[0021] where T(t) is the time embedding sequence feature, W is the first weight matrix corresponding to the Time2Vec layer, b is the first bias vector corresponding to the Time2Vec layer, and X is the feature matrix corresponding to the time series features.
[0022] In a possible implementation, the Transformer encoder includes a multi-head attention mechanism, a feed-forward neural network, a residual connection, and normalization; extracting the global dependence features of the time series features based on the Transformer encoder includes:
[0023] Inputting the time series features into the multi-head attention mechanism, outputting multi-head attention features, and performing a residual connection on the multi-head attention features and the position time series features to obtain residual features;
[0024] Inputting the residual features into the feed-forward neural network, outputting feed-forward features, and performing a residual connection on the feed-forward features and the residual features to obtain the global dependence features of the time series features.
[0025] In a possible implementation, the formula for calculating the dropout rate is:
[0026]
[0027] Among them, α is the initial discard rate, β is the final discard rate, t is the current step number, and T is the total number of training steps.
[0028] In a possible implementation, the fully connected classifier includes a fully connected layer and a Softmax activation function. The process of inputting the multi-level features into the fully connected classifier to identify the human fall result includes:
[0029] Linearly map the multi-level features to the category space based on the fully connected layer, and obtain the probability distribution of human falls based on the activation function.
[0030] In a second aspect, the present invention also provides a human fall recognition device, including:
[0031] A feature acquisition module, configured to acquire the motion features of human key points, and process the motion features based on a sliding window to obtain time series features;
[0032] A feature extraction module, configured to build a deep learning model, and extract multi-level features of the time series features based on the deep learning model;
[0033] A feature recognition module, configured to input the multi-level features into a fully connected classifier to identify the human fall result.
[0034] The beneficial effects of the present invention are:
[0035] The present invention obtains the motion features of human key points, processes the motion features based on a sliding window to obtain time series features, then builds a deep learning model based on time encoding and dynamic discard technology, extracts multi-level features of the time series features based on the deep learning model, and finally inputs the multi-level features into a fully connected classifier to identify the human fall result, improving the recognition accuracy of the model in fall detection. Description of the Drawings
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 It is the method flow chart of an embodiment of the human fall recognition method provided by the present invention;
[0038] Figure 2 For Figure 1 It is the method flow chart of an embodiment of step S102 in
[0039] Figure 3 Schematic structural diagram of an embodiment of the human fall recognition device provided by the present invention. Specific embodiments
[0040] The following will specifically describe the preferred embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0041] In the embodiments of the present invention, the descriptions such as "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0042] Referring to "embodiment" herein means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0043] Figure 1 Schematic flowchart of an embodiment of the human fall recognition method provided by the present invention. As Figure 1 shown, the human fall recognition method includes:
[0044] S101. Obtain the motion features of human key points, and process the motion features based on a sliding window to obtain time series features;
[0045] Among them, the human key points include 13 joint points such as shoulders, hips, knees, and heads. Obtaining the motion features of human key points includes: obtaining the coordinates of human key points and calculating the motion features of human key points based on the coordinates of human key points.
[0046] It can be understood that the two-dimensional coordinates of human key points can be extracted from each frame of RGB video data based on the human key point detection program provided by the Mediapipe library. The data format is as follows:
[0047] P t = {(x1, y1), (x2, y2), …, (x 13 , y 13 )}
[0048] P t represents the set of human key points at time step t, (x i , y i) are the two-dimensional coordinates of the i-th key point. Then, based on the key point coordinates, the frame difference method is used to calculate the motion characteristics of the human key points.
[0049] By inputting the motion characteristics into a sliding window for mean processing, time series characteristics are obtained. Among them, the length and step size of the sliding window can be adaptively adjusted based on the time length of human behavior.
[0050] It should be noted that the length of the sliding window is w, the step size is s, w and s are adaptively adjusted according to the time length of human behavior, and the average value, standard deviation, etc. of the data within each sliding window length are calculated as the result of the current window to generate time series characteristics.
[0051] S102. Construct a deep learning model based on time encoding and dynamic dropout, and extract multi-level features of the time series features based on the deep learning model;
[0052] By combining Time2Vec, that is, time encoding and dynamic Dropout technology to construct a deep learning model, the recall rate and accuracy of the model in fall detection can be significantly improved.
[0053] S103. Input the multi-level features into a fully connected classifier to identify the human fall result.
[0054] In the present invention, by obtaining the motion characteristics of human key points, processing the motion characteristics based on a sliding window to obtain time series characteristics, then constructing a deep learning model based on time encoding and dynamic dropout technology, extracting multi-level features of the time series features based on the deep learning model, and finally inputting the multi-level features into a fully connected classifier to identify the human fall result, the recognition accuracy of the model in fall detection is improved.
[0055] In an embodiment of the present invention, the motion characteristics of human key points include speed, acceleration, centroid speed, and torso angle. Calculating the motion characteristics of human key points based on the coordinates of human key points includes:
[0056] Human key point acceleration formula:
[0057]
[0058] Among them, v(t) is the human key point speed, a(t) is the human key point acceleration, and D(t) is the pixel motion change amount between adjacent frames of the key point;
[0059] Human centroid speed formula:
[0060]
[0061] Among them, v CoM is the human centroid speed, N is the number of key points participating in the calculation, xi and y i are the velocity components of each key point;
[0062] Trunk angle calculation formula:
[0063]
[0064] where θ is the trunk angle, is the vector from the shoulder to the hip, is the vertical vector.
[0065] In some embodiments of the present invention, as Figure 2 shown, step S102 includes:
[0066] S201. Non-linearly embed the time series features based on the Time2Vec layer to obtain time-embedded sequence features;
[0067] It can be understood that when using the Time2Vec technology to embed time series data, temporal information is added to each time step to generate time-embedded features. Specifically, the formula of the Time2Vec layer is:
[0068] T(t) = sin(W·X + b)
[0069] where T(t) is the time-embedded sequence feature, W is the first weight matrix corresponding to the Time2Vec layer, b is the first bias vector corresponding to the Time2Vec layer, and X is the feature matrix corresponding to the time series feature.
[0070] Furthermore, the Time2Vec layer enhances the periodic and linear features of the time series. First, the time stamp t is mapped to a continuous space through a linear transformation. For example, after the time stamp t undergoes a linear transformation with a weight w0 and a bias b0, a linear embedding can be obtained.
[0071] Linear(t) = w0t + b0
[0072] Then, the time stamp is subjected to a periodic transformation using sine and cosine functions, and these two transformations capture the cyclic patterns in the time series.
[0073] Sine(t) = sin(w1b + t1)
[0074] Cosine(t) = cos(w2t + b2)
[0075] The parameters of these transformations, such as w1, b1, w2, b2, are learnable.
[0076] Finally, merge the linear and periodic features, and concatenate the results of the linear part and the periodic part into a feature vector, so that the model can utilize both the linear trend and the periodic pattern of time simultaneously.
[0077] Time2Vec(t) = [w0t + b0, sin(w1t + b1), cos(w2t + b2), …].
[0078] S202. Extract short-term dependence features of time series features based on the LSTM layer;
[0079] LSTM (Long Short-Term Memory) is a special recurrent neural network that can effectively capture short-term features in time series data while alleviating the long-term dependence problem. In LSTM, the key formulas include the input gate, forget gate, output gate, and a state update mechanism:
[0080] Among them, the forget gate controls the discarding of information, and the forget gate formula is:
[0081] f t = σ(W f · [h t-1 , X t + b f )
[0082] The input gate controls the addition of new information, and the input gate formula is:
[0083] i t = σ(W i · [h t-1 , X t + b i )
[0084] The output gate controls the current output, and the output gate formula is:
[0085] o t = σ(W o · [h t-1 , X t + b o )
[0086] The state update mechanism is:
[0087] C t = f t ⊙ C t-1 + i t ⊙ tanh(W c · [h t-1 , X t + b c )
[0088] h t = ot ⊙tanh(C t )
[0089] Among them, X t is the input at the current time step, h t is the hidden state at the current time step, C t is the cell state at the current time step, f t is the activation value of the forget gate, i t is the activation value of the input gate, o t is the activation value of the output gate, W f is the second weight matrix corresponding to the LSTM layer, W i is the third weight matrix corresponding to the LSTM layer, W o is the fourth weight matrix corresponding to the LSTM layer, W c is the fifth weight matrix corresponding to the LSTM layer, b f is the second bias vector corresponding to the LSTM layer, b i is the third bias vector corresponding to the LSTM layer, b o is the fourth bias vector corresponding to the LSTM layer, b c is the fifth bias vector corresponding to the LSTM layer.
[0090] S203. Perform positional embedding on the time series features based on the positional encoding layer, and add the features after positional embedding to the short-term dependency features to obtain positional time series features;
[0091] It should be noted that in Transformer, the original architecture does not directly capture the sequence order. Therefore, positional encoding (Positional Encoding) explicitly adds time position information by performing positional embedding on the time steps, enhancing the ability to perceive the order of time series. Specifically, perform positional embedding on the time series features through the positional encoding layer, and add the features after positional embedding to the short-term dependency features to obtain positional time series features;
[0092] Among them, the common formula for positional encoding is constructed based on sine and cosine functions:
[0093]
[0094] PE is the positional encoding, pos is the position of the time step, i is the index in the embedding dimension, and d is the total embedding dimension.
[0095] S204. Extract the global dependency features of the time series features based on the Transformer encoder, and dynamically adjust the dropout rate of the time series features in the Transformer encoder based on the dynamic Dropout layer.
[0096] The Transformer encoder includes a multi-head attention mechanism, a feed-forward neural network, residual connections, and normalization. Based on the Transformer encoder, the global dependence features of time series features are extracted, and the dropout rate of time series features in the Transformer encoder is dynamically adjusted based on a dynamic Dropout layer, specifically including:
[0097] Input the time series features into the multi-head attention mechanism, output the multi-head attention features, and perform a residual connection on the multi-head attention features and the position time series features to obtain residual features;
[0098] Input the residual features into the feed-forward neural network, output the feed-forward features, and perform a residual connection on the feed-forward features and the residual features to obtain the global dependence features of the time series features.
[0099] It can be understood that the Transformer encoder is used to capture long-term dependence relationships, strengthen feature interactions by combining the multi-head attention mechanism, and adaptively adjust the feature dropout probability by adding a dynamic Dropout mechanism in the Transformer. The dropout rate calculation formula is:
[0100]
[0101] where α is the initial dropout rate, β is the final dropout rate, t is the current step, and T is the total number of training steps.
[0102] It should be noted that the multi-head attention formula is:
[0103]
[0104] The feed-forward neural network formula is:
[0105] FFN(x) = ReLU(xW1 + b1)W2 + b2
[0106] The residual connection and normalization formula is:
[0107] X′ = LayerNorm(X + Attention(Q, K, V))
[0108] where Q is the query matrix, K is the key matrix, V is the value matrix, d k is the feature dimension, W1 is the sixth weight matrix corresponding to the feed-forward neural network, and W2 is the seventh weight matrix corresponding to the feed-forward neural network.
[0109] In an embodiment of the present invention, the fully connected classifier includes a fully connected layer and a Softmax activation function. Input multi-level features into the fully connected classifier to identify the human fall result, including:
[0110] Based on the fully connected layer, the multi-level features are linearly mapped to the category space, and the probability distribution of human falls is obtained based on the activation function.
[0111] It can be understood that the fully connected classifier is used to receive the feature representation output by the model finally and generate the classification result, that is, the probability distribution of fall and non-fall.
[0112] Among them, the fully connected classifier includes a fully connected layer and a Softmax activation function;
[0113] Formula of the fully connected layer:
[0114] z = XW + b
[0115] Where z is the linear output, X is the input feature, W is the weight matrix, and b is the bias term.
[0116] Formula of the Softmax activation function:
[0117]
[0118] The Softmax activation function can convert the linear output z into a probability distribution.
[0119] Furthermore, the specific content of the post-processing of the sliding window in the present invention is as follows:
[0120] The sliding window queue Q stores the prediction results of the most recent T frames. Each prediction result is a binary value y i , where y i = 1 indicates that a fall is detected in the i-th frame, and y i = 0 indicates that no fall is detected in the i-th frame. Therefore, the queue Q stores the fall prediction results of the most recent T frames: Q = [y1, y2,..., y T , and T is the length of the sliding window. Then, the number of fall frames is counted within the window, and the number of frames marked as "fall" y i = 1 within the current sliding window is counted. This can be done by summing all the prediction results within the window: T fall is the number of frames marked as "fall" within the window.
[0121] To determine whether a fall has occurred within the current time period, the number of fall frames T fall is divided by the window length T to calculate the proportion of fall frames: where p fall represents the proportion of "fall" frames within the sliding window.
[0122] By setting a threshold α, usually a constant between 0 and 1, such as 0.5, we determine according to p fallto make a final fall judgment. If the proportion of fall frames is greater than α, it is considered that a fall has occurred; otherwise, it is considered that no fall has occurred:
[0123]
[0124] If p fall >α, that is, most of the frames in the current window are fall frames, then it is judged as a fall and output 1; otherwise, it is judged as no fall and output 0. In addition, every time a frame of data is processed, the sliding window queue needs to be updated. This means that the new prediction result y new is added to the queue, and the old prediction result y old is removed. Therefore, the formula for queue update is: Q = [y2, y3, …, y c , y new . In this way, the window always only contains the prediction results of the most recent T frames.
[0125] To better implement the human fall recognition method in the embodiments of the present invention, correspondingly, based on the human fall recognition method, as Figure 3 shown, the embodiments of the present invention further provide a human fall recognition device 300, and the human fall recognition device 300 includes:
[0126] A feature acquisition module 301, configured to acquire the motion features of human key points, and process the motion features based on a sliding window to obtain time series features;
[0127] A feature extraction module 302, configured to construct a deep learning model based on time encoding and dynamic dropout, and extract multi-level features of the time series features based on the deep learning model;
[0128] A feature recognition module 303, configured to input the multi-level features into a fully connected classifier to recognize the human fall result.
[0129] The human fall recognition device 300 provided in the above embodiments can implement the technical solutions described in the embodiments of the above human fall recognition method. The specific implementation principles of the above modules or units can be referred to the corresponding content in the embodiments of the above human fall recognition method, and will not be elaborated here.
[0130] Those skilled in the art can understand that all or part of the processes for implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory or a random access memory, etc.
[0131] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for identifying a human fall, characterized in that: include: Acquire motion features of key points of a human body, and process the motion features based on a sliding window to obtain time series features; Building a deep learning model based on time encoding and dynamic discarding, and extracting multi-level features of the time series features based on the deep learning model; The multi-level features are input into a fully connected classifier to identify human fall results.
2. The human fall recognition method according to claim 1, characterized in that: The step of obtaining motion features of key points of a human body includes: The coordinates of the key points of the human body are obtained, and the motion characteristics of the key points of the human body are calculated based on the coordinates of the key points of the human body, wherein the motion characteristics of the key points of the human body include speed, acceleration, center of mass speed and trunk angle.
3. The human fall recognition method according to claim 2, characterized in that: The step of processing the motion features based on a sliding window to obtain time series features includes: The motion feature is input into a sliding window for mean processing to obtain the time series feature, wherein the length and step length of the sliding window are adaptively adjusted based on the time length of human behavior.
4. The human fall recognition method according to claim 1, characterized in that: The deep learning model includes a Time2Vec layer, an LSTM layer, a position encoding layer, a Transformer encoder and a dynamic Dropout layer, and the multi-level features of the time series features include time embedding sequence features, short-term dependency features, position time series features and global dependency features.
5. The human fall recognition method according to claim 4, characterized in that: The multi-level features of the time series features extracted based on the deep learning model include: Based on the Time2Vec layer, the time series features are nonlinearly embedded to obtain time embedded series features; Extracting short-term dependency features of the time series features based on the LSTM layer; Based on the position encoding layer, the time series feature is positionally embedded, and the positionally embedded feature is added to the short-term dependency feature to obtain a position time series feature; The global dependency features of the time series features are extracted based on the Transformer encoder, and the dropout rate of the time series features in the Transformer encoder is dynamically adjusted based on the dynamic Dropout layer.
6. The human fall recognition method according to claim 5, characterized in that: The Time2Vec layer formula is: T(t)=sin(W·X+b) Among them, T(t) is the time embedding sequence feature, W is the first weight matrix corresponding to the Time2Vec layer, b is the first bias vector corresponding to the Time2Vec layer, and X is the feature matrix corresponding to the time series feature.
7. The human fall recognition method according to claim 5, characterized in that: The Transformer encoder includes a multi-head attention mechanism, a feedforward neural network, a residual connection and normalization; the global dependency feature of the time series feature extracted based on the Transformer encoder includes: Inputting the time series feature into the multi-head attention mechanism, outputting the multi-head attention feature, and performing a residual connection between the multi-head attention feature and the position time series feature to obtain a residual feature; The residual feature is input into the feedforward neural network, the feedforward feature is output, and the feedforward feature and the residual feature are residually connected to obtain the global dependency feature of the time series feature.
8. The human fall recognition method according to claim 5, characterized in that: The discard rate calculation formula is: Among them, α is the initial drop rate, β is the final drop rate, t is the current step number, and T is the total training step number.
9. The human fall recognition method according to claim 5, characterized in that: The fully connected classifier includes a fully connected layer and a Softmax activation function, and the multi-level features are input into the fully connected classifier to identify the result of a human fall, including: The multi-level features are linearly mapped to the category space based on the fully connected layer, and the probability distribution of a human body falling is obtained based on the activation function.
10. A human fall recognition device, characterized in that: include: A feature acquisition module is used to acquire motion features of key points of a human body and process the motion features based on a sliding window to obtain time series features; A feature extraction module, used for constructing a deep learning model based on time encoding and dynamic discarding, and extracting multi-level features of the time series features based on the deep learning model; The feature recognition module is used to input the multi-level features into a fully connected classifier to identify the result of a human fall.