Real-time status monitoring method for elderly people at home based on temporal deformable attention mechanism
By adopting a real-time status monitoring method based on the time-deformable attention mechanism in the home elderly monitoring system, the problems of high difficulty in deploying equipment and high cost of use are solved, efficient and convenient monitoring and emotional analysis are achieved, and monitoring accuracy and humanistic care are improved.
Patent Information
- Application Number
- CN202311388239.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-10-25
AI Technical Summary
The existing home elderly monitoring system has problems such as high difficulty in deploying equipment, inconvenient life for monitoring objects, and high cost of use.
The real-time state monitoring method based on the time-deformable attention mechanism is adopted to monitor the elderly at home in real time through the camera, and the improved yolov7 extracts the 2D posture map of the video human body, combines the time-deformable attention mechanism and 3D convolution to build an action recognition model, and uses a multi-head attention network to build an expression recognition model.
It reduces the difficulty of equipment deployment, improves monitoring accuracy and efficiency, reduces computing resource consumption, enhances the ability to recognize actions and expressions, realizes a real-time emotional scoring system, and improves the convenience of monitoring and humanistic care.
Smart Images

Figure CN117218709B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a real-time status monitoring method for elderly people at home based on a time-deformable attention mechanism. Background Art
[0002] In recent years, action recognition and expression recognition have attracted extensive research attention in academia and industry in the field of artificial intelligence. They have had a positive impact in many fields, such as human-computer interaction, intelligent monitoring systems, virtual reality, etc.
[0003] Action recognition is generally divided into dynamic and static action recognition. Action features include human posture, motion trajectory, action speed, etc., and the same action may present various forms due to different environmental conditions, observation angles, and individual differences of the actor. Therefore, the use of dynamic action recognition improves the detection accuracy. Compared with static images, video action sequences contain more spatiotemporal information, so they are more challenging for computer vision systems.
[0004] Expression recognition mainly relies on features such as facial muscle movement, eye and mouth morphology changes to determine the real-time expression of the detected object. However, due to the diversity and individual differences of facial expressions, as well as the influence of factors such as lighting, angle, occlusion, and posture, expression recognition needs to be continuously optimized and updated.
[0005] The use of deep neural networks and computer vision technology can greatly improve the efficiency of action recognition and expression recognition. For action recognition tasks, by taking the time dimension into consideration, a network structure that adapts to the spatiotemporal relationship can be designed to extract temporal features in action sequences. For expression recognition tasks, network structures such as convolutional neural networks can be used to extract features from facial images. Due to the slight differences between actions and expressions, it is necessary to learn feature representations with distinguishing capabilities in order to accurately classify different action or expression categories. Therefore, metric learning, attention mechanisms, etc. have emerged to enhance the model's perception of key features. Summary of the invention
[0006] In view of this, the purpose of the present invention is to provide a real-time status monitoring method for elderly people at home based on a time-deformable attention mechanism, which monitors the elderly at home in real time through a camera, aiming to solve the problems of high difficulty in equipment deployment, inconvenience in life of the monitored objects and high cost of use under the current monitoring mode.
[0007] To achieve the above purpose, the present invention adopts the following technical solution: a real-time status monitoring method for elderly people at home based on a time-deformable attention mechanism, comprising the following steps:
[0008] Step S1: Extract the 2D posture graph of the human body in the video based on the improved yolov7, stack it into a 3D heat map along the time dimension, and use multiple methods such as subject center cropping and uniform sampling to preprocess the data;
[0009] Step S2: Using the time-deformable attention mechanism module and the feedforward neural network, using network hidden frame weighting, combined with 3D convolution, to build an action recognition model;
[0010] Step S3: extract the face position based on the Harr cascade classifier, combine the island loss function with the feature clustering network, multi-head attention network and attention fusion network to build the expression recognition model;
[0011] Step S4: Perform iterative training according to the specified training parameters, update the parameters of the action recognition model and the expression recognition model by optimizing the combined loss, and continuously save the optimal model according to the verification accuracy; use the action recognition model to build a multi-level action discrimination warning system, and combine the expression recognition model to build a real-time emotion scoring system.
[0012] In a preferred embodiment, step S1 specifically includes the following steps:
[0013] Step S11: Use yolov7-pose to perform target detection, merge low-level features with high-level features, and thus improve the feature representation capability of the yolov7 model; then perform 2D human pose estimation and extract up to 17 key points;
[0014] Step S12: After extracting the 2D human posture key points, a 3D heat map volume stacked along the time dimension is formulated; we represent the 2D posture as a heat map of size K×H×W, where K is the number of joints, H and W are the height and width of the video frame; in the case of the corresponding bounding box given by the yolov7 object detector, the heat map is padded with zeros to match the size of the original frame; the human joint coordinates (x k ,y k ) and the confidence score c k , combine K Gaussian maps centered on each joint to obtain the joint heat map J:
[0015]
[0016] σ1 is the variance of the Gaussian graph, (x o ,y o ) represents the coordinates of the points around the joint coordinates, and e is a natural constant; and, using the extracted human key points, a human limb heat map L is constructed:
[0017]
[0018] Function D calculates point (x o ,yo ) to line segment seg(a k , b k ), a k , b k Indicates the two ends of a limb. Represents the confidence of the joint points at both ends; finally, all heat maps (J or L) are superimposed along the time dimension to obtain a three-dimensional heat map body, whose size is K×Ti×H×W, where Ti is the time length;
[0019] Step S13: First, the center cropping technology is used to crop all frames according to the size of the minimum target bounding box of all 2D pose estimates, and adjust them to the size of the detected target, which can not only retain all action information but also reduce the size of the 3D heat map volume space; since processing each frame of the video will cause a lot of computational overhead, the uniform sampling method is then used to evenly segment the video into n′ segments with the same number of frames, and one frame is extracted from each segment and spliced into a shorter video to reduce the length in the time dimension; the data is processed using flipping, deformation, and scaling methods.
[0020] In a preferred embodiment: step S2 specifically includes the following steps:
[0021] Step S21: Use the temporal deformable attention mechanism; take a set of video features as query input; then it will output a set of action predictions; each action prediction is represented as a tuple of time segment, confidence score and label; use the temporal deformable attention module TDA to adaptively focus on the features of the temporal positions around the reference position in the input feature sequence; first assume the input video represents the real number space; T S Refers to the length of the time dimension, and C represents the dimension of a certain frame; therefore, the features in the feature sequence are feature vectors extracted from each frame of the video. Next, the features of each frame will be enhanced so that each frame has temporal context features;
[0022] Let query vector t q ∈[0,1] is the normalized coordinate of the corresponding reference point, where the reference point is a frame of the video; the input is The output of the mth TDA module header is It is calculated as a weighted sum of a set of key elements sampled from X:
[0023]
[0024] k n represents the number of sampling points, a mqk∈[0, 1] is the normalized attention weight of each sampling point, reflecting the degree of attention to different sampling points; Δt mqk ∈[0,1] is relative to t q The sampling offset of X((t q +Δt mqk )T S ) means that in (t q +Δt mqk )T S The linear interpolation feature at ; then the query feature z is obtained by linear projection q Predict the attention weight a mqk and sampling offset Δt mqk ; Use the softmax function to normalize the attention weights, is the weight value of each frame, which is a learnable parameter; the output of TDA is calculated by the linear combination of the outputs of different TDA heads:
[0025] TDA(z q , t q , X) = W O concat(h1,h2,...,h m )
[0026] It is also a set of learnable weights, and concat represents a linear combination;
[0027] When calculating the t'th frame in the output sequence, the query point and the reference point are both the t'th frame in the input sequence, and the query feature is the sum of the input feature of the frame and other position features embedded at that position; position embedding is used to distinguish different positions in the input sequence, and the sinusoidal position embedding method is used to determine the embedding position:
[0028]
[0029] γ=1, 2, 3…, set according to actual situation;
[0030] Step S22: Weighting each frame Assign, weight the features of all frames of all identified segments; by calculating the video coding vector c′ and the hidden layer representation k i The similarity is calculated to obtain the weight coefficient corresponding to each frame feature; the calculation formula is as follows:
[0031]
[0032] T represents matrix transpose, T0 is the number of input video frames, ξ i is the weight of the i-th frame, and V0 is a learnable parameter;
[0033] Step S23: The decoding layer uses the self-attention mechanism combined with the temporal deformable attention (TDA) to transform the former TDA (z q , t q , X) as input, and through connecting the pooling layer and the feedforward neural network, the prediction result of the decoding layer can be obtained;
[0034] Step S24: The previous step is to use the attention mechanism to improve the network's ability to recognize videos. Here, we introduce a skeleton-based 3D convolutional network as the backbone network for action recognition. Among various 3D convolutions, the slowonly network is selected as the main network component, and the previously proposed attention mechanism is embedded in the network layer; in the slowonly network, the parameters used for 3D convolution are different. Here, the dimension of the convolution kernel is expressed as Respectively represent the time step, space step, and channel size. We use different types of convolution to extract video features. The usage of each layer of convolution is as follows:
[0035] The first convolutional layer is: 1×7 2 , 64
[0036] The second convolution residual connection layer uses:
[0037]
[0038] The third convolution residual connection layer uses:
[0039]
[0040] The fourth convolutional residual connection layer uses:
[0041]
[0042] In a preferred embodiment: step S3 specifically includes the following steps:
[0043] Step S31: extracting the face position based on Haar cascade classifier; Haar cascade classifier is a cascade structure composed of a large number of weak classifiers, each weak classifier is used to detect a specific feature of the image; the cascade structure allows non-face areas to be quickly filtered out, and only the areas that may contain faces are detected in more detail; after the corresponding face is detected, it is cropped according to the minimum target frame of face detection, and only the face part is retained; random noise, blurring, and color change processing methods are added to part of the data;
[0044] Step S32: To build a multi-head attention network, the first part of the network uses a feature clustering network; the entire network is based on a residual network. We use two loss functions, one is called affinity loss and the other is island loss. The purpose of using two loss functions is to maximize the boundaries between different classes and the distance between the centers of different classes while making the distances within the same category as close as possible; we assume that the input of the network is x i , the input label is y i , the output of this part of the network is x′ i :
[0045] x′ i =F(w r .x i )
[0046] F represents the part of the network, w r Represents the network parameters; then uses affinity loss:
[0047]
[0048] is the class center matrix, each column corresponds to the center of a specific class, is the column vector in c, representing the actual label, N′ is the number of images trained in this batch, σ c represents the standard deviation of each class center, Y represents the label space, and D0 is the class center dimension; the island loss function is used at the same time:
[0049]
[0050] τ is a custom threshold;
[0051] Step S33: The second part is a multi-head attention network. Our method constructs 1×1, 1×3, 3×1 and 3×3 convolution kernels to capture multi-scale local features; the channel attention unit consists of a global average pooling layer, two linear layers and an activation function, and two linear layers are used to encode channel information;
[0052] K a Pay attention to the space head, K a spatial attention map; since the output of the first part is x′ i , the output of the j-th spatial attention unit is:
[0053] s j = x′ i ⊙H j (w s , x′ i ), j∈{1,...,Ka}
[0054] w s represents the network parameters, and assumes is the channel attention head, is the final attention feature vector output by the channel attention head, then the j-th output is:
[0055] a j =s j ⊙H′ j (w s ,s j ), j∈{1,...,K a}
[0056] Step S34: The third part uses the attention fusion network; the attention fusion network scales the attention feature vector by applying the log-softmax function; because in the second part of the multi-head attention network, the output attention vector feature The feature scaling result is:
[0057]
[0058] Take 512 here in L0, and then use the partition loss method;
[0059]
[0060] for The variance of is calculated to guide the attention heads to focus on different key areas and avoid attention overlap. Finally, the normalized attention feature vectors are merged into one, and then a linear layer is used to calculate the class confidence.
[0061] In a preferred embodiment: step S4 specifically includes the following steps:
[0062] Step S41: For the action recognition model, we directly use the cross entropy loss and gradient descent method to optimize the model; for the expression recognition model, we use four loss functions to form a new loss function:
[0063]
[0064] in, For affinity loss, For the island's loss, is the partition loss, is the cross entropy loss of the prediction result; λ1, λ2, and λ3 represent the coefficients of the corresponding loss function, which are adjusted as needed; then the model is continuously iterated, the model parameters are updated using the gradient descent method, the model accuracy is continuously verified, and the optimal model parameters are retained;
[0065] Step S42: After the model training of action recognition and expression recognition is performed, the trained model is deployed in the real-time monitoring system for elderly people at home; the actions are graded into three levels, which represent three situations: no danger, possible danger, and danger:
[0066] (1) We consider actions such as waving, sitting, walking, standing, and lying down to be normal and non-dangerous actions;
[0067] (2) We regard actions such as headaches, back discomfort, knee discomfort, coughing, sneezing, etc. as potentially dangerous actions, maintain strict monitoring of the people in the picture, and remind family members of potential dangers or possible diseases;
[0068] (3) We consider actions such as falling or calling for help as dangerous actions and use alarms to alert family members;
[0069] Step S43: Real-time emotion scoring system based on expression recognition. We use the expression recognition model combined with the camera to capture the elderly’s facial expression once per second and predict the elderly’s mood. At the same time, we score the elderly’s mood every second according to the confidence of the prediction, and calculate the average score of the elderly’s real-time mood on the day in real time. We assume that the current mood score is score i
[0070] When the mood is disgust and contempt:
[0071] score1=100-60*pro
[0072] When you are happy or excited:
[0073] score2=90+10*pro
[0074] When the mood is neutral or surprised:
[0075] score3=60+25*pro
[0076] When you are angry or sad:
[0077] score4=100-80*pro
[0078] pro is the confidence of the prediction, and the confidence interval is [0, 1], which can be used to calculate the average real-time sentiment score of the day.
[0079] Compared with the prior art, the present invention has the following beneficial effects:
[0080] 1. The present invention relates to a real-time status monitoring method for elderly people at home based on a time-deformable attention mechanism. The system monitors the elderly at home in real time through a camera, aiming to solve the problems of high difficulty in equipment deployment, inconvenience in life of the monitored objects and high cost of use under the current monitoring mode.
[0081] 2. For action recognition based on human posture estimation, the present invention uses yolov7-pose to extract human posture, and uses the time-deformable attention mechanism and 3D convolution to perform action recognition on the posture in the video. Compared with the recognition in RGB mode, it reduces the consumption of computing resources and improves the recognition speed. And compared with static action recognition, it pays more attention to the changing characteristics of the action and improves the accuracy of action recognition.
[0082] 3. For expression recognition, the present invention uses a multi-head attention network to strengthen the focus on the features of different parts of the face, reduce the interference of similar features between expressions, enhance the ability to recognize the potential differences between multiple similar facial expressions, and more accurately identify the emotions of the monitored objects.
[0083] 4. The combination of motion recognition and expression recognition adds an emotion scoring mode to real-time monitoring, adding an element of humanistic care and being closer to actual monitoring needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 It is a schematic diagram of the principle of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0085] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0086] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0087] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.
[0088] like Figure 1 As shown, this embodiment provides a real-time status monitoring method for elderly people at home based on a time-deformable attention mechanism, which specifically includes the following steps:
[0089] Step S1: Extract the 2D posture graph of the human body in the video based on the improved yolov7, stack it into a 3D heat map along the time dimension, and use multiple methods such as subject center cropping and uniform sampling to preprocess the data;
[0090] Step S2: Using the time-deformable attention mechanism module and the feedforward neural network, using network hidden frame weighting, combined with 3D convolution, to build an action recognition model;
[0091] Step S3: extract the face position based on the Harr cascade classifier, combine the island loss function with the feature clustering network, multi-head attention network and attention fusion network to build the expression recognition model;
[0092] Step S4: Perform iterative training according to the specified training parameters, update the parameters of the action recognition model and the expression recognition model by optimizing the combined loss, and continuously save the optimal model according to the verification accuracy. Use the action recognition model to build a multi-level action discrimination warning system, and combine the expression recognition model to build a real-time emotion scoring system.
[0093] In this embodiment, step S1 specifically includes the following steps:
[0094] Step S11: Use yolov7-pose for target detection. In order to make the model sensitive to the changes in the original data features, add jump connections to merge low-level features with high-level features, thereby improving the feature representation ability of the yolov7 model. Then perform 2D human pose estimation and extract up to 17 key points.
[0095] Step S12: After extracting the 2D human pose key points, a 3D heatmap volume is constructed by stacking along the time dimension. We represent the 2D pose as a heatmap of size K×H×W, where K is the number of joints, and H and W are the height and width of the video frame. Given the corresponding bounding box given by the yolov7 object detector, the heatmap is padded with zeros to match the size of the original frame. We use the human joint coordinates (x k ,y k ) and the confidence score c k , combine K Gaussian maps centered on each joint to obtain the joint heat map J:
[0096]
[0097] σ1 is the variance of the Gaussian graph, (x o ,y o ) represents the coordinates of the points around the joint coordinates, and e is a natural constant. In addition, the human limb heat map L can be constructed using the extracted human key points:
[0098]
[0099] Function D calculates point (x o ,y o ) to line segment seg(a k , b k ), a k , b k Indicates the two ends of a limb. Represents the confidence of the joint points at both ends. Finally, all heat maps (J or L) are superimposed along the time dimension to obtain a three-dimensional heat map body with a size of K×Ti×H×W, where Ti is the time length.
[0100] Step S13: First, the center cropping technique is used to crop all frames according to the size of the minimum target bounding box of all 2D pose estimates, and adjust them to the size of the detected target, which can not only retain all action information, but also reduce the size of the 3D heat map volume space. Since processing each frame of the video will cause a lot of computational overhead, the uniform sampling method is then used to evenly segment the video, divide the video into n′ segments with the same number of frames, and extract one frame from each segment to splice into a shorter video to reduce the length in the time dimension. In order to ensure the generalization ability of the model, we need to enhance the data set to ensure that the model has high recognition accuracy at different angles and different distances. Therefore, the data is processed using flipping, deformation, scaling and other processing methods.
[0101] In this embodiment, step S2 specifically includes the following steps:
[0102] Step S21: Use the temporal deformable attention mechanism. Take a set of video features as query input. Then it will output a set of action predictions. Each action prediction is represented as a tuple of time segment, confidence score and label. Use the temporal deformable attention module TDA to adaptively focus on the features of the temporal positions around the reference position in the input feature sequence. First, assume that the input video represents the real number space. S Refers to the length of the time dimension, and C represents the dimension of a frame. Therefore, the features in the feature sequence are feature vectors extracted from each frame of the video. Next, each frame will be feature enhanced so that each frame has temporal context features.
[0103] Let query vector t q ∈[0, 1] is the normalized coordinate of the corresponding reference point, where the reference point is a frame of the video. The input is The output of the mth TDA module header is It is calculated as a weighted sum of a set of key elements sampled from X:
[0104]
[0105] k n represents the number of sampling points, a mqk ∈[0, 1] is the normalized attention weight of each sampling point, reflecting the degree of attention to different sampling points. mqk ∈[0,1] is relative to t q The sampling offset of X((t q +Δt mqk )T S ) means that in (t q +Δt mqk )T S Then, we use linear projection to get the linear interpolation feature from the query feature z q Predict the attention weight a mqk and sampling offset Δt mqk . Use the softmax function to normalize the attention weights. is the weight value of each frame, which is a learnable parameter. The output of TDA is calculated by linear combination of the outputs of different TDA heads:
[0106] TDA(z q , t q , X) = W O concat(h1,h2,...,h m )
[0107] It is also a set of learnable weights, and concat represents a linear combination.
[0108] When calculating the t'th frame in the output sequence, the query point and the reference point are both the t'th frame in the input sequence, and the query feature is the sum of the input feature of the frame and other position features embedded at that position. Position embedding is used to distinguish different positions in the input sequence, and the embedded position is determined using the sinusoidal position embedding method:
[0109]
[0110] γ=1, 2, 3…, set according to actual situation.
[0111] Step S22: Weighting each frame Assign weights to the features of all frames of all identified segments. By calculating the video coding vector c′ and the hidden layer representation k i The similarity of each frame is calculated to obtain the weight coefficient corresponding to each frame feature. The calculation formula is as follows:
[0112]
[0113] T represents matrix transpose, T0 is the number of input video frames, ξi is the weight of the i-th frame, and V0 is a learnable parameter.
[0114] Step S23: The decoding layer uses the self-attention mechanism combined with the temporal deformable attention (TDA) to transform the former TDA (z q , t q , X) as input, and through connecting the pooling layer and the feedforward neural network, the decoding layer prediction result can be obtained.
[0115] Step S24: The previous step is to use the attention mechanism to improve the network's ability to recognize videos. Here, we introduce a skeleton-based 3D convolutional network as the backbone network for action recognition. Among various 3D convolutions, we select the slowonly network as the main network component and embed the previously proposed attention mechanism in the network layer. In the slowonly network, the parameters used for the 3D convolution are different. Here, the dimension of the convolution kernel is expressed as Respectively represent the time step, space step, and channel size. We use different types of convolution to extract video features. The usage of each layer of convolution is as follows:
[0116] The first convolutional layer is: 1×7 2 , 64
[0117] The second convolution residual connection layer uses:
[0118]
[0119] The third convolution residual connection layer uses:
[0120]
[0121] The fourth convolutional residual connection layer uses:
[0122]
[0123] In this embodiment, step S3 specifically includes the following steps:
[0124] Step S31: Extract the face position based on Haar cascade classifier. Haar cascade classifier is a cascade structure composed of a large number of weak classifiers, each of which is used to detect a specific feature of the image. The cascade structure allows non-face areas to be quickly filtered out, and only the areas that may contain faces are detected in more detail. After detecting the corresponding face, in order to reduce the computational overhead, we crop it according to the minimum target box of face detection and only retain the face part. In order to improve the generalization ability of the model and ensure that the accuracy of expression detection is maintained at a high level under different circumstances, random noise, blurring, color change and other processing methods are added to some data.
[0125] Step S32: To build a multi-head attention network, the first part of the network uses a feature clustering network. The entire network is based on a residual network. We use two loss functions, one is called affinity loss and the other is island loss. The purpose of using two loss functions is to maximize the boundaries between different classes and the distance between the centers of different classes while making the distances within the same category as close as possible. We assume that the input of the network is x i , the input label is y i , the output of this part of the network is x′ i :
[0126] x′ i =F(w r .x i )
[0127] F represents the part of the network, w r Represents the network parameters. Then use affinity loss:
[0128]
[0129] is the class center matrix, each column corresponds to the center of a specific class, is the column vector in c, representing the actual label, N′ is the number of images trained in this batch, σ c represents the standard deviation of each class center, Y represents the label space, and D0 is the class center dimension. At the same time, the island loss function is used:
[0130]
[0131] A custom threshold.
[0132] Step S33: The second part is a multi-head attention network. Our method constructs 1×1, 1×3, 3×1 and 3×3 convolution kernels to capture multi-scale local features. The channel attention unit consists of a global average pooling layer, two linear layers and an activation function, and two linear layers are used to encode channel information.
[0133] K a Pay attention to the space head, K a A spatial attention map. Since the output of the first part is x′ i , the output of the j-th spatial attention unit is:
[0134] s j = x′ i ⊙H j (w s , x′ i), jv{1,...,K a}
[0135] w s represents the network parameters, and assumes is the channel attention head, is the final attention feature vector output by the channel attention head, then the j-th output is:
[0136] a j =s j ⊙H′ j (w s ,s j ), j∈{1,...,K a}
[0137] Step S34: The third part uses the attention fusion network. The attention fusion network scales the attention feature vector by applying the log-softmax function. Because in the second part of the multi-head attention network, the attention vector feature is output The feature scaling result is:
[0138]
[0139] Take 512 here in L0, and then use the partition loss method;
[0140]
[0141] for The variance of is calculated to guide the attention heads to focus on different key areas and avoid attention overlap. Finally, the normalized attention feature vectors are merged into one, and then a linear layer is used to calculate the class confidence.
[0142] In this embodiment, step S4 specifically includes the following steps:
[0143] Step S41: For the action recognition model, we directly use the cross entropy loss and gradient descent method to optimize the model. For the expression recognition model, we use 4 loss functions to form a new loss function:
[0144]
[0145] in, For affinity loss, For the island's loss, is the partition loss, is the cross entropy loss of the prediction result. λ1, λ2, and λ3 represent the coefficients of the corresponding loss function, which can be adjusted as needed. Then, the model is continuously iterated, and the model parameters are updated using the gradient descent method to continuously verify the model accuracy and retain the optimal model parameters.
[0146] Step S42: After training the models for action recognition and expression recognition, the trained models are deployed in the real-time monitoring system for elderly people at home. We classify actions into three levels, which represent three situations: no danger, possible danger, and danger.
[0147] (1) We consider actions such as waving, sitting, walking, standing, and lying down to be normal and non-hazardous.
[0148] (2) We regard actions such as headaches, back discomfort, knee discomfort, coughing, sneezing, etc. as potentially dangerous actions, maintain strict monitoring of the people in the picture, and remind family members of potential dangers or possible diseases.
[0149] (3) We consider actions such as falling or calling for help as dangerous and use alarms to alert family members.
[0150] Step S43: Real-time emotion scoring system based on expression recognition. We use the expression recognition model combined with the camera to capture the elderly’s facial expression once per second and predict the elderly’s mood. At the same time, we score the elderly’s mood every second according to the confidence of the prediction, and calculate the average score of the elderly’s real-time mood on the day in real time. We assume that the current mood score is score i
[0151] When the mood is disgust and contempt:
[0152] score1=100-60*pro
[0153] When you are happy or excited:
[0154] score2=90+10*pro
[0155] When the mood is neutral or surprised:
[0156] score3=60+25*pro
[0157] When you are angry or sad:
[0158] score4=100-80*pro
[0159] pro is the confidence of the prediction, and the confidence interval is [0, 1], which can be used to calculate the average real-time sentiment score of the day.
[0160] In particular, most existing home-based elderly monitoring systems are real-time monitoring based on the combination of multiple intelligent devices, which have the problems of high difficulty in equipment deployment, inconvenience in the life of the monitored object, and high cost of use. The present invention hopes to rely on relevant technologies in the field of computer vision to achieve a more efficient and convenient monitoring mode. This example uses motion recognition to identify the human posture of the monitored object, and uses expression recognition to perform real-time analysis of the mood of the monitored object, adding humanistic care elements on the basis of monitoring. For action recognition based on human posture estimation, the present invention uses yolov7-pose to extract human posture, and uses time-deformable attention mechanism and 3D convolution to perform action recognition on the posture in the video. Compared with recognition in RGB mode, it reduces the consumption of computing resources and improves the recognition speed. And compared with static action recognition, it pays more attention to the changing characteristics of the action and improves the accuracy of action recognition. For expression recognition, the present invention uses a multi-head attention network to strengthen the attention to the features of different parts of the face, reduce the interference of similar features between expressions, enhance the recognition ability of the potential differences of multiple similar facial expressions, and more accurately identify the emotions of the monitored object.
[0161] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.
Claims
1. A real-time status monitoring method for elderly people at home based on a time-deformable attention mechanism, characterized by: The following steps are involved: Step S1: Extract the 2D posture graph of the human body in the video based on the improved yolov7, stack it into a 3D heat map along the time dimension, and use the subject center cropping and uniform sampling to preprocess the data in multiple ways; Step S2: Using the time-deformable attention mechanism module and the feedforward neural network, using network hidden frame weighting, combined with 3D convolution, to build an action recognition model; Step S3: extract the face position based on the Harr cascade classifier, combine the island loss function with the feature clustering network, multi-head attention network and attention fusion network to build the expression recognition model; Step S4: Perform iterative training according to the specified training parameters, update the parameters of the action recognition model and the expression recognition model by optimizing the combined loss, and continuously save the optimal model according to the verification accuracy; use the action recognition model to build a multi-level action discrimination warning system, and combine the expression recognition model to build a real-time emotion scoring system; Step S1 specifically includes the following steps: Step S11: Use yolov7-pose to perform target detection, merge low-level features with high-level features, and thus improve the feature representation capability of the yolov7 model; then perform 2D human pose estimation and extract up to 17 key points; Step S12: After extracting the 2D human posture key points, a 3D heat map volume stacked along the time dimension is formulated; we represent the 2D posture as a heat map of size K×H×W, where K is the number of joints, H and W are the height and width of the video frame; in the case of the corresponding bounding box given by the yolov7 object detector, the heat map is padded with zeros to match the size of the original frame; the human joint coordinates (x k ,y k ) and the confidence score c k , combine K Gaussian maps centered on each joint to obtain the joint heat map J: σ1 is the variance of the Gaussian graph, (x o ,y o ) represents the coordinates of the points around the joint coordinates, and e is a natural constant; and, using the extracted human key points, a human limb heat map L is constructed: Function D calculates point (x o ,y o ) to line segment seg(a k , b k ), a k , b k Indicates the two ends of a limb. Represents the confidence of the joint points at both ends; finally, all heat maps are superimposed along the time dimension to obtain a three-dimensional heat map body, whose size is K×Ti×H×W, where Ti is the time length; Step S13: First, the center cropping technique is used to crop all frames according to the size of the minimum target bounding box of all 2D pose estimates, and adjust them to the size of the detected target, which can not only retain all action information but also reduce the size of the 3D heat map volume space; since processing each frame of the video will cause a lot of computational overhead, the uniform sampling method is then used to evenly segment the video, and the video is divided into n′ segments with the same number of frames, and one frame is extracted from each segment and spliced into a shorter video to reduce the length in the time dimension; the data is processed using flipping, deformation, and scaling methods; Step S2 specifically includes the following steps: Step S21: Use the temporal deformable attention mechanism; take a set of video features as query input; then it will output a set of action predictions; each action prediction is represented as a tuple of time segment, confidence score and label; use the temporal deformable attention module TDA to adaptively focus on the features of the temporal positions around the reference position in the input feature sequence; first assume the input video represents the real number space; T S Refers to the length of the time dimension, and C represents the dimension of a certain frame; therefore, the features in the feature sequence are feature vectors extracted from each frame of the video. Next, the features of each frame will be enhanced so that each frame has temporal context features; Let query vector t q ∈[0, 1] is the normalized coordinate of the corresponding reference point, where the reference point is a frame of the video; the input is The output of the mth TDA module header is It is calculated as a weighted sum of a set of key elements sampled from X: k n represents the number of sampling points, a mqk ∈[0, 1] is the normalized attention weight of each sampling point, reflecting the degree of attention to different sampling points; Δt mqk ∈[0,1] is relative to t q The sampling offset of X((t q +Δt mqk )T s ) means that in (t q +Δt mpk )T S The linear interpolation feature at ; then the query feature z is obtained by linear projection q Predict the attention weight a mqk and sampling offset Δt mqk ; Use the softmax function to normalize the attention weights, is the weight value of each frame, which is a learnable parameter; the output of TDA is calculated by the linear combination of the outputs of different TDA heads: TDA(z q ,t q ,X)=W O concat(h1,h2,...,h m ) It is also a set of learnable weights, and concat represents a linear combination; When calculating the t'th frame in the output sequence, the query point and the reference point are both the t'th frame in the input sequence, and the query feature is the sum of the input feature of the frame and other position features embedded at that position; position embedding is used to distinguish different positions in the input sequence, and the sinusoidal position embedding method is used to determine the embedding position: γ=1, 2, 3…, set according to actual situation; Step S22: Weighting each frame Assign, weight the features of all frames of all identified segments; by calculating the video coding vector c′ and the hidden layer representation k i The similarity is calculated to obtain the weight coefficient corresponding to each frame feature; the calculation formula is as follows: T represents matrix transpose, T0 is the number of input video frames, ξ i is the weight of the i-th frame, and V0 is a learnable parameter; Step S23: The decoding layer uses the self-attention mechanism combined with the temporal deformable attention (TDA) to transform the former TDA (z q , t q , X) as input, and through connecting the pooling layer and the feedforward neural network, the prediction result of the decoding layer can be obtained; Step S24: The previous step is to use the attention mechanism to improve the network's ability to recognize videos. Here, we introduce a skeleton-based 3D convolutional network as the backbone network for action recognition. Among various 3D convolutions, the slowonly network is selected as the main network component, and the previously proposed attention mechanism is embedded in the network layer; in the slowonly network, the parameters used for 3D convolution are different. Here, the dimension of the convolution kernel is expressed as Respectively represent the time step, space step, and channel size. We use different types of convolution to extract video features. The usage of each layer of convolution is as follows: The first convolutional layer is: 1×7 2 , 64 The second convolution residual connection layer uses: The third convolution residual connection layer uses: The fourth convolutional residual connection layer uses: Step S3 specifically includes the following steps: Step S31: extracting the face position based on Haar cascade classifier; Haar cascade classifier is a cascade structure composed of a large number of weak classifiers, each weak classifier is used to detect a specific feature of the image; the cascade structure allows non-face areas to be quickly filtered out, and only the areas that may contain faces are detected in more detail; after the corresponding face is detected, it is cropped according to the minimum target frame of face detection, and only the face part is retained; random noise, blurring, and color change processing methods are added to part of the data; Step S32: To build a multi-head attention network, the first part of the network uses a feature clustering network; the entire network is based on a residual network. We use two loss functions, one is called affinity loss and the other is island loss. The purpose of using two loss functions is to maximize the boundaries between different classes and the distance between the centers of different classes while making the distances within the same category as close as possible; we assume that the input of the network is x i , the input label is y i , the output of this part of the network is x′ i : x′ i =F(w r ,x i ) F represents the part of the network, w r Represents the network parameters; then uses affinity loss: is the class center matrix, each column corresponds to the center of a specific class, is the column vector in c, representing the actual label, N′ is the number of images for batch training, σ c represents the standard deviation of each class center, Y represents the label space, and D o is the class center dimension; and the island loss function is used: τ is a custom threshold; Step S33: The second part is a multi-head attention network. Our method constructs 1×1, 1×3, 3×1 and 3×3 convolution kernels to capture multi-scale local features; the channel attention unit consists of a global average pooling layer, two linear layers and an activation function, and two linear layers are used to encode channel information; K a Pay attention to the space head, K a spatial attention map; since the output of the first part is x′ i , the output of the jth spatial attention unit is: s j =x′ i ⊙H j (w s ,x′ i ),j∈{1,...,K a } w s represents the network parameters, and assumes is the channel attention head, is the final attention feature vector output by the channel attention head, then the j-th output is: a j =s j ⊙H′ j (w s ,s j ),j∈{1,...,K a } Step S34: The third part uses the attention fusion network; the attention fusion network scales the attention feature vector by applying the log-softmax function; because in the second part of the multi-head attention network, the output attention feature vector The feature scaling result is: Take 512 here in L0, and then use the partition loss method: for The variance of is used to guide the attention heads to focus on different key areas and avoid attention overlap. Finally, the normalized attention feature vectors are merged into one, and then the class confidence is calculated using a linear layer. Step S4 specifically includes the following steps: Step S41: For the action recognition model, we directly use the cross entropy loss and gradient descent method to optimize the model; for the expression recognition model, we use four loss functions to form a new loss function: in, For affinity loss, For the island's loss, is the partition loss, is the cross entropy loss of the prediction result; λ1, λ2, and λ3 represent the coefficients of the corresponding loss function, which are adjusted as needed; then the model is continuously iterated, the model parameters are updated using the gradient descent method, the model accuracy is continuously verified, and the optimal model parameters are retained; Step S42: After the model training of action recognition and expression recognition is performed, the trained model is deployed in the real-time monitoring system for elderly people at home; the actions are graded into three levels, which represent three situations: no danger, possible danger, and danger: (1) We consider waving, sitting, walking, standing, and lying down to be normal movements. Considered as a non-dangerous action; (2) We regard headaches, back discomfort, knee discomfort, coughing, and sneezing as potentially dangerous actions, maintain strict monitoring of the people on the screen, and alert family members of potential dangers or possible diseases; (3) We consider falling and calling for help as dangerous actions and use alarms to alert family members; Step S43: Real-time emotion scoring system based on expression recognition. We use the expression recognition model combined with the camera to capture the elderly’s facial expression once per second and predict the elderly’s mood. At the same time, we score the elderly’s mood every second according to the confidence of the prediction, and calculate the average score of the elderly’s real-time mood on the day in real time. We assume that the current mood score is score i When the mood is disgust or contempt: score1=100-60*pro When you are happy or excited: score2=90+10*pro When the mood is neutral or surprised: score3=60+25*pro When you are angry or sad: score4=100-80*pro pro is the confidence of the prediction, and the confidence interval is [0, 1], which can be used to calculate the average real-time sentiment score of the day.
Citation Information
Patent Citations
Human motion prediction method based on adversarial training attention mechanism
CN114386582A
Expression recognition method based on attention-modulated contextual spatial information
WO2023185243A1