Method of indoor old man falling detection system based on mixed mode
By integrating video and audio information in the indoor elderly fall detection system, and using advanced algorithms and models for identity identification, fall detection and emotion recognition, the problems of high detection accuracy and cost in the existing technology are solved, and an efficient, reliable, economical and practical fall detection effect is achieved.
Patent Information
- Application Number
- CN202510303715.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing indoor elderly fall detection technology has the problems of a single sensor being susceptible to environmental interference, low accuracy and high cost of complex systems, and it is difficult to widely use.
The indoor elderly fall detection system based on a hybrid method is adopted. By integrating video and audio information, advanced algorithms and models are used to achieve accurate identification of the elderly’s identity, efficient detection of fall behavior, and timely detection of abnormal emotions, and can quickly issue alarm signals.
It significantly improves the accuracy and reliability of fall detection, reduces misjudgments and misjudgments caused by single data, ensures the safety of the elderly, and is systematic, economical and practical.
Smart Images

Figure CN120164302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of indoor elderly safety monitoring, and specifically to a method for an indoor elderly fall detection system based on a hybrid approach. Background Art
[0002] With the deepening of the aging degree, the safety issues of the elderly living alone have become increasingly prominent.
[0003] When the elderly are moving indoors, fall accidents occur from time to time. If not discovered and rescued in time, serious consequences may be caused. Existing fall detection technologies have many defects. For example, the detection method based on a single sensor is easily affected by the environment and has a low accuracy rate; some complex monitoring systems are costly and difficult to be widely applied. Therefore, it is crucial to develop an accurate, reliable, economical and practical indoor elderly fall detection technology. Thus, a method for an indoor elderly fall detection system based on a hybrid approach is proposed. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a method for an indoor elderly fall detection system based on a hybrid approach. By integrating video and audio information and applying advanced algorithms and models, it can accurately identify the identity of indoor elderly people, efficiently detect fall behaviors, promptly detect abnormal emotions, and quickly send out alarm signals to ensure the life safety and health of the elderly. It solves the problem that when the elderly are moving indoors, fall accidents occur from time to time. If not discovered and rescued in time, serious consequences may be caused. Existing fall detection technologies have many defects. For example, the detection method based on a single sensor is easily affected by the environment and has a low accuracy rate; some complex monitoring systems are costly and difficult to be widely applied. Therefore, it is crucial to develop an accurate, reliable, economical and practical indoor elderly fall detection technology.
[0006] (2) Technical Solutions
[0007] To achieve the above object of accurately identifying the identity of the elderly indoors, efficiently detecting fall behaviors, and promptly detecting abnormal emotions by integrating video and audio information and applying advanced algorithms and models, and being able to quickly send out alarm signals to ensure the life safety and health of the elderly, the present invention provides the following technical solutions: An indoor elderly fall detection system based on a hybrid method, including a video acquisition module, an audio acquisition module, an identity recognition module, a fall detection module, an emotion recognition module, and an alarm module; the video acquisition module is used to acquire indoor video data; the audio acquisition module is used to acquire indoor audio data; the identity recognition module is connected to the video acquisition module and is used to identify the identity of the person in the video; the fall detection module is respectively connected to the video acquisition module and the identity recognition module and is used to perform fall detection after determining the target elderly person living alone; the emotion recognition module is connected to the audio acquisition module and is used to perform emotion recognition on the acquired audio; the alarm module is connected to the fall detection module and the emotion recognition module and is used to send out an alarm signal when an abnormality is detected.
[0008] Preferably, the identity recognition module uses the ResNet34 model for face recognition, specifically including:
[0009] Input face data, capture and detect the face in the image using the Dlib library, and extract the 128D features of the input face.
[0010] Obtain the current frame of the video stream in real time, perform face detection, calculate the 128D features of the face in the current frame, and compare them with the input face data using the Euclidean distance. If the Euclidean distance is less than the set threshold of 0.4, it is determined as the target person; if it exceeds the threshold, it is determined as a non-target person.
[0011] Preferably, in the fall detection module, the key point coordinates of the human body skeleton are obtained through AlphaPose and used as the input of ST-GCN. A spatio-temporal graph is constructed with joints as graph nodes and the natural connection of the human skeleton and the temporal relationship of the same joints as graph edges.
[0012] Preferably, the ST-GCN uses spatial graph convolution and temporal graph convolution to process data. The spatial graph convolution constructs an intra-frame spatial graph convolution according to the natural connectivity of human joints.
[0013] G s =(V S ,E S )
[0014] where V S ={v ti |i = 1,…,N S} represents all the joint points in the skeleton, and E S =v ti v tj|(i,j)∈H represents the connection between joint points, and each node is described by the feature vector F(V tt ) to represent the spatial features; the temporal graph convolution connects the same joint points in consecutive multi-frame images on the spatial graph to form the spatio-temporal graph of the skeleton sequence:
[0015] G T =(V T , E T )
[0016] where V T =v ti |t = 1, ……, N represents the joint point sequence of the same part; E T =v ti v (t+1)i represents the connection between them.
[0017] Preferably, the overall ST-GCN model is composed of 9 layers of STGCNs, which input the joint coordinate vector of the graph nodes, extract the 256-dimensional node features of each node, where the dimension of the key frame is 38, perform global pooling on the data, use backpropagation to train the model end-to-end, and analyze the behavior category with the highest probability through the softMax classifier. The loss function is calculated using cross-entropy.
[0018] Preferably, each ST-GCN layer adopts the ResNet structure. At the same time, to solve the problem of gradient explosion, a dropout strategy is added to the ST-GCN layer; the loss function of the ST-GCN is cross-entropy, and the calculation formula is:
[0019]
[0020] In the formula, c represents the category, M represents the number of categories, y c represents the variable (0 or 1), which is 1 when the category is the same as the sample category, and 0 otherwise; represents the predicted probability of category c.
[0021] The indoor elderly fall detection method based on the hybrid method includes the indoor elderly fall detection system based on the hybrid method, and also includes the following steps:
[0022] S1. Use the video acquisition module and the audio acquisition module to collect the video data and audio data in the room;
[0023] S2. The identity recognition module recognizes the identity of the person in the video;
[0024] S3. If it is the target elderly living alone, the fall detection module performs fall detection based on the video data;
[0025] S4. The emotion recognition module performs emotion recognition on the audio data;
[0026] S5. When a fall or abnormal emotion is detected, the alarm module emits an alarm signal.
[0027] Preferably, the emotion recognition module first preprocesses the collected voice signal, including setting a low-pass filter to remove noise, using a windowing operation to convert the voice signal into a digital signal, and selecting a Hamming window as the window function:
[0028]
[0029] After that, voice features are extracted, including short-time energy, short-time average amplitude, zero-crossing rate, pitch period and frequency, formants, and Mel-frequency cepstral coefficients (MFCC).
[0030] Preferably, the emotion recognition module automatically extracts the features of the voice signal using OPENSMILE, calculates the weights using an attention mechanism, inputs the voice signal through LSTM, and calculates the weight values for each frame; the weight coefficients corresponding to each keyword are obtained by calculating the correlation between each query value and each keyword, and the calculation methods include the vector dot product method, the cosine function method, and an additional neural network evaluation method; the Softmax function is used to normalize the weights, and the formula is:
[0031]
[0032] where Lχ represents the length of the corresponding data source; the weighted sum of the weight coefficients and the corresponding key values is used to obtain the attention mechanism, and the result of multiplying the weight values obtained by the attention mechanism by the input matrix H is input into the fully connected layer for classification output to complete the voice emotion recognition.
[0033] (III) Beneficial effects
[0034] Compared with the prior art, the present invention provides a method for an indoor elderly fall detection system based on a hybrid method, having the following beneficial effects:
[0035] 1. The method for the indoor elderly fall detection system based on the hybrid method, by innovatively integrating video and audio data, realizes multi-dimensional monitoring of the elderly's state. The video acquisition module captures the movement postures of the elderly, and the audio acquisition module obtains voice and environmental sound information. The two complement each other. When the movement changes of the person in the video are not obvious, abnormal sounds in the audio, such as the collision sound of a fall, a shout, etc., can assist in the judgment. This multi-dimensional data acquisition and analysis method greatly reduces the misjudgment and missed judgment situations caused by single data, significantly improves the accuracy and reliability of detection, and comprehensively guarantees the safety of the elderly.
[0036] 2. The hybrid-based indoor elderly fall detection system adopts advanced deep learning models such as ResNet34, ST-GCN and LSTM combined with attention mechanism. ResNet34 effectively avoids the gradient vanishing problem through its unique residual structure, accurately extracts facial features in face recognition tasks, and accurately identifies the identity of the elderly. ST-GCN can perform in-depth analysis of the space-time graph composed of human skeletal point coordinates and accurately detect fall actions. LSTM combined with attention mechanism can focus on key voice information and improve emotion recognition accuracy. These models work together to greatly improve the system's ability to process and identify complex data, building a solid technical barrier for elderly safety monitoring.
[0037] 3. The method of the indoor elderly fall detection system based on a hybrid approach has a scientific and reasonable system architecture design, which includes video acquisition, audio acquisition, identity recognition, fall detection, emotion recognition and alarm modules. The modules have clear division of labor and close cooperation. The video and audio acquisition modules are responsible for data acquisition, the identity recognition module determines whether it is the target elderly person, the fall and emotion recognition modules perform abnormality detection, and the alarm module responds in a timely manner. When the identity recognition determines that the target is an elderly person living alone, the fall and emotion recognition modules work immediately. Once an abnormality is detected, the alarm module quickly issues an alarm. This architecture enables the system to operate efficiently and collaboratively, and respond to the elderly's falls and abnormal emotions in a timely manner. It has extremely high practicality and application value, and provides solid protection for the safety of the elderly. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a structural schematic diagram of the indoor elderly fall detection system of the present invention;
[0039] Figure 2 Schematic diagram of the residual network model of the present invention;
[0040] Figure 3 It is a schematic diagram of the attention calculation process of the present invention;
[0041] Figure 4 It is a schematic flow chart of the indoor elderly fall detection method of the present invention. DETAILED DESCRIPTION
[0042] The following will be combined with the embodiments of the present invention and the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] See also Figures 1-4, An indoor elderly fall detection system based on a hybrid method, comprising a video acquisition module, an audio acquisition module, an identity recognition module, a fall detection module, an emotion recognition module, and an alarm module; the video acquisition module is used to acquire indoor video data; the audio acquisition module is used to acquire indoor audio data; the identity recognition module is connected to the video acquisition module and is used to identify the identity of the person in the video; the fall detection module is respectively connected to the video acquisition module and the identity recognition module and is used to perform fall detection after determining the target elderly living alone; the emotion recognition module is connected to the audio acquisition module and is used to perform emotion recognition on the acquired audio; the alarm module is connected to the fall detection module and the emotion recognition module and is used to emit an alarm signal when an abnormality is detected.
[0044] An indoor elderly fall detection method based on a hybrid method, comprising the indoor elderly fall detection system based on a hybrid method, and further comprising the following steps:
[0045] S1. Use the video acquisition module and the audio acquisition module to acquire indoor video data and audio data;
[0046] S2. The identity recognition module identifies the identity of the person in the video;
[0047] S3. If it is the target elderly living alone, the fall detection module performs fall detection based on the video data;
[0048] S4. The emotion recognition module performs emotion recognition on the audio data;
[0049] S5. When a fall or abnormal emotion is detected, the alarm module emits an alarm signal.
[0050] Embodiment 1:
[0051] The present invention adopts a Deep Residual Networks (ResNet), which is a deep learning architecture. Compared with shallower models, deeper networks can capture higher-level abstractions and complex features, thus more comprehensively understanding the semantic information in images. In this structure, each layer will transform and refine the input features, gradually sublimating simple information such as edges and textures into high-level descriptions such as shapes, local structures, and even categories. As the number of layers increases, the network can learn more complex features, thus having an obvious advantage in expressing image semantics.
[0052] However, in a simply stacked network, due to fewer parameters and low computational complexity per layer, the gradient tends to gradually weaken or even vanish during backpropagation. This phenomenon leads to poor update effects of the underlying weights, and further causes the entire network to converge slowly or even stagnate. Although the use of non-linear activation functions can improve the model's expressive ability, their saturation effect in deep networks may make the activation values extremely small, exacerbating the problem of gradient decay. Meanwhile, as the number of network layers and parameters increases, it becomes more difficult to update the underlying parameters, thus affecting the overall performance of the model.
[0053] The ResNet34 model of the present invention approximates the ideal mapping H(x) by stacking non-linear layers. In the limit case, the residual part F(x) = H(x) - x approaches zero, enabling the network to capture the target mapping more accurately. Figure 2 The structural schematic of the residual network is shown. F(x) + x introduces skip connections in the feedforward network, enabling the network to directly learn the residual mapping F(x) = H(x) - x. The so-called skip connection means directly merging the output of the previous layer with the input of the next layer. This design constructs a direct transmission channel between different layers, thus avoiding the gradual attenuation of information. This not only simplifies the training process but also effectively alleviates the problem of gradient vanishing.
[0054] Example 2:
[0055] The spatio-temporal graph convolutional neural network (ST-GCN) proposed by the present invention is a deep learning model specifically designed for spatio-temporal data. Its basic idea is similar to traditional graph convolution methods. By performing convolution operations on nodes and edges in spatio-temporal data, it efficiently extracts the relevant information of time and space. Combining motion analysis, the model classifies the spatial graph into three categories: static, centripetal, and centrifugal motion.
[0056] ST-GCN processes data using spatial graph convolution and temporal graph convolution. The spatial graph convolution constructs an intra-frame spatial graph convolution based on the natural connectivity of human joints.
[0057] G s =(V S , E S )
[0058] where V S ={v ti | i = 1, …, N S} represents all joint points in the skeleton, and E S = v ti v tj | (i, j) ∈ H represents the connections between joint points. Each node is represented by a feature vector F(V tt)Describe spatial features; Temporal graph convolution connects the same joint points in consecutive multi-frame images on the spatial graph to form a spatio-temporal graph of the skeleton sequence:
[0059] G T =(V T ,E T )
[0060] where V T =v ti |t = 1, ……, N represents the joint point sequence of the same part; E T =v ti v (t+1)i represents the connection between them.
[0061] Combined with motion analysis, the model classifies the spatial graph into three categories: static, centripetal, and centrifugal motions. Specifically, the skeleton joint point itself as the root node represents the static feature; the nodes adjacent to the root node and closer to the skeleton center correspond to centripetal motion; while the nodes adjacent to the root node and far from the skeleton center represent centrifugal motion. Performing convolution operations on these three subsets respectively helps capture action features at different scales.
[0062] In terms of input, the joint coordinate vectors of each node are directly used as the original data. The overall network consists of 9 layers of ST-GCN, gradually mining deeper features, and finally generating a 256-dimensional feature representation for each node (the key frame dimension is 38). Then, the feature tensor is integrated through global pooling, and end-to-end training is achieved using backpropagation. Finally, the softMax classifier outputs the behavior category with the highest probability.
[0063] In addition, to optimize the gradient propagation effect, each ST-GCN layer adopts the ResNet structure; at the same time, to prevent gradient explosion, a dropout strategy is introduced in these layers, thus ensuring the stability of model training.
[0064] Example 3:
[0065] 1. Speech signal preprocessing.
[0066] In the present invention, the speech signals collected by the audio acquisition module are often interfered by transmission noise. Therefore, preprocessing operations such as filtering, framing, and windowing need to be performed before subsequent processing to reduce the risk of aliasing and distortion and ensure more accurate and clear speech data.
[0067] First, to remove noise, a low-pass filter is used to preprocess the signal. Although the filter can effectively reduce noise interference, simply setting a fixed threshold is not the best method. If the threshold is set too high, some useful speech signals may be erroneously filtered out; if the threshold is too low, the noise cannot be fully eliminated. In addition, directly setting the signals below the threshold to zero may also cause distortion and affect subsequent speech analysis. Therefore, for different noise situations, a more optimal filtering scheme should be selected and corresponding optimization should be carried out.
[0068] The second step is windowing and framing. To convert continuous speech into digital signals that are easy to process, the signal amplitude needs to be extracted at fixed time intervals. Framing is a key step in preprocessing, which divides the complex time-varying signal into multiple relatively stable short-time frames for subsequent feature extraction. Framing mainly includes two processes:
[0069] 1. Window function processing: Use window functions (such as Hamming window, rectangular window, Hanning window, etc.) to separate each frame of the signal and ensure a certain overlap between adjacent frames, so as to realize the conversion of continuous signals into discrete signals.
[0070] 2. Frame shift: Introduce partial overlap (usually 2 to 3 times the frame length) at the edges of adjacent frames to ensure the correlation and continuity between frames.
[0071] After the above steps, the speech signal is divided into multiple stable short-time frames, and corresponding feature parameters can be extracted for each frame. In the present invention, the Hamming window is selected as the window function w(n), and the windowed speech signal s w (n) = s(n) * w(n), and the formula of the Hamming window function is as follows:
[0072]
[0073] After obtaining the framed features, training can be carried out for classification.
[0074] II. Speech Feature Analysis
[0075] Traditional emotion recognition methods use artificially designed low-level descriptors and high-level statistical functions to extract features, which are prone to discarding a large amount of information. In the present invention, OPENSMILE is used to automatically extract the features of speech signals, which has applications in the fields of speech recognition, sentiment analysis, audio classification, and speech conversion.
[0076] In the process of emotion recognition, the attention mechanism is used to calculate weights. The attention mechanism enables the network to learn more effective information and improve the performance of the model. The self-attention mechanism calculates the importance of elements based on the interaction between the elements of the input sequence, dynamically weights the input sequence. For a given input sequence, it is first converted into a vector representation, then the similarity scores between elements are calculated and normalized, and the context representation vector is obtained by dynamically weighted summation according to the scores, such asFigure 3 For the specific calculation process of the attention mechanism.
[0077] In speech signal processing, after input through LSTM, the weight values of each frame are calculated. There are three methods for calculating the weight coefficients: the vector dot product method, the cosine function method, and an additional neural network evaluation method. The Softmax function is used to normalize the weights, and the formula is:
[0078]
[0079] Among them, Lχ represents the length of the corresponding data source; the weighted sum of the weight coefficients and the corresponding key values is used to obtain the attention mechanism. The result of multiplying the weight value obtained from the attention mechanism by the input matrix H is input into the fully connected layer for classification output to complete speech emotion recognition.
[0080] In summary, the method of the indoor elderly fall detection system based on the hybrid method innovatively fuses video and audio data to achieve multi-dimensional monitoring of the elderly's state. The video acquisition module captures the action postures of the elderly, and the audio acquisition module obtains voice and environmental sound information. The two complement each other. When the changes in the actions of the people in the video are not obvious, abnormal sounds in the audio, such as the collision sound and shouting sound of a fall, can assist in the judgment. This multi-dimensional data acquisition and analysis method greatly reduces the misjudgment and missed judgment situations caused by single data, significantly improves the accuracy and reliability of detection, and comprehensively ensures the safety of the elderly.
[0081] Moreover, the method of the indoor elderly fall detection system based on the hybrid method adopts advanced deep learning models such as ResNet34, ST-GCN, and LSTM combined with the attention mechanism. ResNet34 effectively avoids the problem of gradient disappearance through its unique residual structure, accurately extracts face features in face recognition tasks, and accurately identifies the identity of the elderly. ST-GCN can deeply analyze the spatio-temporal graph composed of the coordinates of human body bone points and accurately detect fall actions. LSTM combined with the attention mechanism can focus on the key information of speech and improve the emotion recognition accuracy. These models work together to greatly improve the system's processing and recognition capabilities for complex data, and build a technical barrier for the safety monitoring of the elderly.
[0082] Moreover, the method of the indoor elderly fall detection system based on the hybrid mode has a scientific and reasonable system architecture design, including video acquisition, audio acquisition, identity recognition, fall detection, emotion recognition, and alarm modules. Each module has a clear division of labor and close cooperation. The video and audio acquisition modules are responsible for data acquisition, the identity recognition module determines whether it is the target elderly person, the fall and emotion recognition modules conduct anomaly detection, and the alarm module responds in a timely manner. When the identity recognition determines that it is the target elderly person living alone, the fall and emotion recognition modules start working immediately. Once an anomaly is detected, the alarm module quickly issues an alarm. This architecture enables the system to operate efficiently and collaboratively, respond in a timely manner to the elderly's fall and abnormal emotion situations, has extremely high practicality and application value, provides a solid guarantee for the safety of the elderly, and solves the problem that when the elderly are moving indoors, fall accidents occur frequently. If they cannot be discovered and rescued in time, serious consequences may be caused. Existing fall detection technologies have many defects. For example, the detection method based on a single sensor is easily affected by the environment and has a low accuracy rate; some complex monitoring systems are costly and difficult to be widely applied. Therefore, it is crucial to develop a precise, reliable, and cost-effective indoor elderly fall detection technology.
[0083] All relevant modules involved in this system are either hardware system modules or functional modules that combine computer software programs or protocols in the prior art with hardware. The computer software programs or protocols themselves involved in this functional module are all well-known technologies to those skilled in the art, and they are not the improvements of this system; the improvement of this system is the interaction relationship or connection relationship between each module, that is, the overall structure of the system is improved to solve the corresponding technical problems to be solved by this system.
[0084] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An indoor elderly fall detection system based on a hybrid approach, comprising a video acquisition module, an audio acquisition module, an identity recognition module, a fall detection module, an emotion recognition module and an alarm module; characterized in that: The video acquisition module is used to collect indoor video data; the audio acquisition module is used to collect indoor audio data; the identity recognition module is connected to the video acquisition module and is used to identify the identity of the person in the video; the fall detection module is connected to the video acquisition module and the identity recognition module respectively, and is used to perform fall detection after determining that the target is an elderly person living alone; the emotion recognition module is connected to the audio acquisition module and is used to perform emotion recognition on the collected audio; the alarm module is connected to the fall detection module and the emotion recognition module, and is used to send an alarm signal when an abnormality is detected.
2. The indoor elderly fall detection system based on a hybrid method according to claim 1 is characterized in that: The identity recognition module uses the ResNet34 model for face recognition, which specifically includes: Enter face data, use the Dlib library to capture and detect faces in images, and extract 128D features of the entered faces; The current frame of the video stream is obtained in real time, face detection is performed, the 128D features of the face in the current frame are calculated, and the Euclidean distance is compared with the recorded face data. If the Euclidean distance is less than the set threshold of 0.4, it is determined to be the target person. If it exceeds the threshold, it is determined to be a non-target person.
3. The indoor elderly fall detection system based on a hybrid method according to claim 1 is characterized in that: In the fall detection module, the coordinates of the key points of the human skeleton are obtained through AlphaPose and used as the input of ST-GCN to construct a spatiotemporal graph with joints as graph nodes and the natural connections of the human skeleton and the temporal relationships of the same joints as graph edges.
4. The indoor elderly fall detection system based on a hybrid method according to claim 3 is characterized in that: The ST-GCN uses spatial graph convolution and temporal graph convolution to process data. The spatial graph convolution constructs intra-frame spatial graph convolution according to the natural connectivity of human joints. G s =(V S ,E S ) Where V S = {v ti |i=1,…,N S } represents all the nodes in the skeleton, E S =v ti vt j| (i,j)∈H represents the connection between joints, and each node is represented by a feature vector F(V tt ) describes the spatial features; the temporal graph convolution connects the same joint points in multiple consecutive frames of images on the spatial graph to form a spatiotemporal graph of the skeleton sequence: G T =(V T ,E T ) Among them, V T =v ti |t=1,……,N represents the sequence of joint points in the same part; E T =v ti v (t+1)i Indicates the connection between them.
5. The indoor elderly fall detection system based on a hybrid method according to claim 4 is characterized in that: The ST-GCN model as a whole consists of 9 layers of STGCN. The joint coordinate vectors of the graph nodes are input, and the 256-dimensional node features of each node are extracted, where the key frame dimension is 38. The data is globally pooled, and the model is trained end-to-end using back propagation. The behavior category with the highest probability is obtained through softMax classifier analysis, and the loss function is calculated using cross entropy.
6. The indoor elderly fall detection system based on a hybrid method according to claim 5 is characterized in that: Each ST-GCN layer adopts the ResNet structure, and in order to solve the gradient explosion problem, a dropout strategy is added to the ST-GCN layer; the loss function of the ST-GCN is the cross entropy, and the calculation formula is: In the formula, c represents the category, M represents the number of categories, and y c Represents a variable (0 or 1), which is 1 when the category is consistent with the sample category, otherwise it is 0; it represents the predicted probability of category c.
7. A hybrid-based indoor elderly fall detection method, characterized in that: The hybrid indoor elderly fall detection system according to claims 1 to 6 further comprises the following steps: S1, using the video acquisition module and the audio acquisition module to collect indoor video data and audio data; S2, the identity recognition module recognizes the identity of the person in the video; S3. If the target is an elderly person living alone, the fall detection module performs fall detection based on video data; S4, the emotion recognition module performs emotion recognition on the audio data; S5. When a fall or abnormal emotion is detected, the alarm module sends an alarm signal.
8. The indoor elderly fall detection method based on a hybrid approach according to claim 7 is characterized in that: The emotion recognition module pre-processes the collected speech signal, including setting a low-pass filter to remove noise, converting the speech signal into a digital signal by using a windowing operation, and selecting a Hamming window as the window function: Then the speech features are extracted, including short-time energy, short-time average amplitude, zero-crossing rate, pitch period and frequency, formant and Mel-frequency cepstral coefficient (MFCC).
9. The indoor elderly fall detection method based on a hybrid approach according to claim 7 is characterized in that: The emotion recognition module uses OPENSMILE to automatically extract voice signal features, uses the attention mechanism to calculate weights, and calculates the weight value of each frame after the voice signal is input through LSTM; the weight coefficient of the value corresponding to each keyword is obtained by calculating the correlation between each query value and each keyword, and the calculation methods include vector dot product method, cosine function method and additional neural network evaluation method; Use the Softmax function to normalize the weights. The formula is: Where L χ Indicates the length of the corresponding data source; The weight coefficient and the corresponding key value are weighted and summed to obtain the attention mechanism. The weight value obtained by the attention mechanism is multiplied by the input matrix H and input into the fully connected layer for classification and output to complete the speech emotion recognition.
Citation Information
Cited By
Home abnormal event detection method based on multi-modal data and edge intelligent device
CN121170965A
A home abnormal event detection method based on multi-modal data and edge intelligent device
CN121170965B
Non-contact anti-falling monitoring system for old people based on cloud intelligence
CN121564879A
Old person falling-down detection system and method based on double-flow convolutional neural network
CN121725388A