Multi-mode getting-up intention recognition method and device, terminal and medium
Through the integration of multimodal sensing technology of millimeter wave radar and vision cameras, combined with the hierarchical timing attention model HTAM, the accuracy and robustness of the elderly’s intention to stand up are solved, and the accurate identification of the elderly’s movements are achieved, and the safety and efficiency of elderly care are improved.
Patent Information
- Application Number
- CN202510262484.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-04
AI Technical Summary
The existing elderly intention recognition system cannot accurately judge the elderly’s intention to stand up in the absence of light, small posture changes or complex scenarios, and a single sensor scheme is prone to false alarms or missed reports, and cannot capture the multi-stage nature of the elderly’s stand up movement and dependence on support objects.
The millimeter wave radar and vision camera multimodal sensing technology are used to combine the hand bedside handrail to grasp features, and the elderly’s intention to stand up through the hierarchical timing attention model HTAM, including data preprocessing, feature extraction, alignment and fusion, the time series is processed using the LSTM module, and multi-stage identification is performed in combination with the attention mechanism.
It improves the accuracy and robustness of the elderly's intention to stand up, adapts to the specific action characteristics of the elderly, reduces misjudgment, and improves the safety and efficiency of elderly care.
Smart Images

Figure CN120260112A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical technologies, and particularly to a multi-modal getting-up intention recognition method, device, terminal and medium adapted to the characteristics of the elderly. Background Art
[0002] For some elderly people with certain mobility, when attempting to get up and stand on their own (such as for the need to get up at night), they may encounter a situation of being unable to exert enough force or sudden physical discomfort and getting into a posture stalemate. Existing elderly getting-up intention recognition systems mainly rely on a single sensor (such as a pressure sensor CN109833045B, a vision sensor CN114724078B) to monitor the getting-up behavior of the elderly.
[0003] This type of solution may not be able to accurately judge the getting-up intention of the elderly under specific conditions (such as insufficient light, small posture changes). The ZA202305013B patent uses an IMU to detect the posture of the elderly. The problem is that a pure IMU can only recognize the relative relationship of joints and cannot recognize the position relationship in space (such as the position relationship between the elderly's hand and the handrail), and the elderly do not like to wear sensors.
[0004] In addition, a single-sensor solution often cannot handle complex scenarios or noise interference (such as sheet stacking, device accidental touch, etc.), resulting in false alarms or missed alarms. Compared with other groups, the getting-up actions of the elderly have certain particularities, which are mainly manifested in two aspects:
[0005] Firstly, the getting-up actions of the elderly are slow and have a relatively long pause stage. For example, the elderly may sit on the edge of the bed for a while to relieve dizziness before getting up, or move their limbs first for preparatory actions. This multi-stage nature may cause traditional getting-up intention recognition solutions to misjudge it as an ordinary turning over or other actions; secondly, the elderly need to rely more on supports to maintain balance during the getting-up process, while young people usually do not. This difference makes the elderly look for the edge of the bed or the handrail as a support when getting up, and traditional recognition systems may ignore this feature and cannot accurately judge when the elderly start to get up. Summary of the Invention
[0006] This application provides a multi-modal getting-up intention recognition method adapted to the characteristics of the elderly. The advantage is that through the fusion of different modal sensing technologies such as millimeter-wave radar and vision cameras, it provides a more accurate, efficient and adaptable data input for the recognition of the elderly's getting-up intention. Combining the auxiliary information of the hand grasping feature of the bedside armrest, and by introducing a multi-stage hierarchical temporal model HTAM (Hierarchical Temporal Attention Model), the ability to recognize the elderly's getting-up intention is improved. It has a broad market application prospect. This system can not only enhance the safety of elderly care, but also reduce the burden on caregivers and provide better care services for the elderly.
[0007] The technical solution of this application is as follows: A multi-modal getting-up intention recognition method adapted to the characteristics of the elderly, including the following steps:
[0008] S1: Obtain the millimeter-wave radar data of the user and generate a point cloud, and map the distance information into point cloud coordinates;
[0009] S2: Obtain the visual data of the user, identify the human key points and object bounding boxes from the user's visual data, and judge the user's getting-up intention according to the overlapping relationship between the human key points and the object bounding boxes;
[0010] S3: Align and fuse the millimeter-wave radar feature vector and the visual feature vector to obtain a fused feature vector;
[0011] S4: Construct a hierarchical temporal attention model HTAM, combine the attention mechanism and the hierarchical structure on the basis of the temporal model, and recognize the user's getting-up intention through the temporal attention model HTAM.
[0012] Further, step S1 includes the following steps:
[0013] S11: Denoise the millimeter-wave radar data to remove high-frequency noise;
[0014] S12: Data standardization, standardize the denoised distance, speed, and acceleration data;
[0015] S13: Calculate the position of the three-dimensional point in the polar coordinate system according to the distance and angle information, and map the distance information into point cloud coordinates.
[0016] Further, step S2 includes the following steps:
[0017] S21: Detection of human key points and grasping object bounding boxes;
[0018] Among them, the grasping object is several objects that the user often grasps when getting up and is preset. Identify the bounding boxes and class labels of the above grasping objects, expressed as:
[0019] O objects = {(x min , y min , x max , y max , p) | p ∈ {bedside armrest, chair armrest}};
[0020] The described human key points include 17 human skeletal key points and 4 key points h of the hand hand ;
[0021] S22: Judgment of getting-up intention;
[0022] First, judge whether there is contact by calculating the spatial overlap relationship between the key points of the hand and the bounding box of the grasped object:
[0023]
[0024] If d min < ε, where ε is the design threshold, it is determined that the hand is in contact with the object. If the key points of the hand are continuously detected within the object border for multiple frames, it is marked as "grasping possible";
[0025] The probability P grasp ∈ [0, 1] represents whether a grasping action is detected; if P grasp > τ, indicating that the grasping probability is higher than the set threshold τ, the judgment of "getting-up intention" is enhanced; if P grasp < τ indicates that the grasping probability is relatively low, and the judgment of the hierarchical temporal attention model HTAM is relied on.
[0026] Furthermore, in step S22, the probability P grasp is obtained in the following way:
[0027] According to the judgment result of d min < ε, a regional judgment is made on 4 hand feature points. If the key points of the hand are within the armrest border, the contact point count is incremented by 1;
[0028] Calculate the bending angle of the finger:
[0029] vector1 = fingertip coordinate - finger root coordinate;
[0030] vector2 = finger root coordinate - wrist coordinate;
[0031] Bending angle = angle_between(vector1, vector);
[0032] Define the feature weight, with the contact accounting for 70% and the bending accounting for 30%;
[0033] After normalizing each eigenvalue, perform weighted summation, and then convert it into probability P through the Sigmoid function grasp :
[0034] Contact point normalization: Contact = number of contact points / 4;
[0035] Bending normalization: Bend = 1 - bending angle / 180;
[0036] Weighted summation: Raw = 0.7 * Contact + 0.3 * Bend;
[0037] Sigmoid probability output: P grasp = 1 / (1 + e^(-0.5 * Raw - 0.5)).
[0038] Furthermore, step S3 includes the following steps:
[0039] S31: Data alignment;
[0040] Calibrate between the millimeter-wave radar and the vision camera to determine the key point coordinate mapping relationship between the two: In a unified coordinate system, calibrate the point cloud coordinates (x radar , y radar , z radar ) of the millimeter-wave radar and the key point coordinates (x vision , y vision , z vision ) of the vision camera; Use the calibration tool to obtain the corresponding point set offline and calculate the transformation matrix T;
[0041] The calibration process is as follows:
[0042] Define the target: Set 9 sampling points to sample the corresponding three-dimensional points in the millimeter-wave radar coordinate system and the vision coordinate system respectively; And record that the transformation matrix T is composed of the rotation matrix R and the translation vector t; So that Pvisual = R Pradar + t.
[0043] Among them, the coordinates of the 9 sampling points are P radar,i = [x radar,i , y radar,i , z radar,i T and [x vision,i , y vision,i , z vision,i T ;
[0044] Construct the objective function: To minimize the error between the transformed points and the actual vision points, define the error function as:
[0045]
[0046] Least squares method solution: First, expand the error as: Take the derivatives of R and t respectively and set them to zero to obtain the minimum value; among them, the rotation matrix R can be solved by SVD,
[0047] First, sample points at 9 different positions and construct two matrices:
[0048] P radar =[P radar,1 ,P radar,2 ,...,P radar,9
[0049] P vision =[P vision,1 ,P vision,2 ,...,P vision,9
[0050] Calculate the covariance matrix
[0051] H = P vision T P radar
[0052] Perform SVD decomposition on the H matrix
[0053] H = U∑V T
[0054] The final rotation matrix R = VU T , translation vector
[0055] Align the millimeter-wave radar point cloud data with the visual human key point information. For each visual key point P vision ( i ), search for the nearest point P radar ( j ) in the radar point cloud; if a visual key point corresponds to multiple point cloud clusters, take the mean; judge whether the visual recognition result at this point is invalid according to a confidence value provided by mediapipe for each key point recognition; if the confidence of this key point is lower than the set threshold τ = 0.5, then consider this key point invalid; then the historical data of the previous time needs to be used as the visual virtual coordinates of this key point for matching with the radar point cloud data;
[0056] Step S32: Feature extraction and matching:
[0057] In each period t, extract the feature vectors of the millimeter-wave point cloud and the visual camera respectively:
[0058] f radar (t)=[P radar,1 ,P radar,2 ,...,P radar,17
[0059] f vision (t)=[P vision,1 ,P vision,2 ,...,P vision,17
[0060] Feature weighted fusion:
[0061] f fusion =a*f radar +(1 - a)*f vision ; α is a parameter;
[0062] The fused feature vector f fusion Through time - series stacking, it is input into the hierarchical temporal attention model HTAM, and then successively input into the LSTM modules of each stage to complete multi - stage dynamic intention recognition..
[0063] Furthermore, in the hierarchical temporal attention model HTAM:
[0064] (1) The input data is the feature sequence Fk fused from millimeter - wave radar and vision camera, representing the temporal feature segment of this stage, with a shape of (m, 17). Where m is the time window length;
[0065] (2) The hierarchical temporal attention model HTAM includes multi - stage LSTM modules. The LSTM processes each feature vector f fusion (t) of the time series step by step through internal gating units, and maintains a dynamic hidden state h(t) and cell state C(t); The purpose of the input gate i(t) is to determine which information should be added to the cell state, the forget gate determines which information is discarded from the cell state, and the output gate o(t) determines the update of the hidden state;
[0066] i(t)=σ(W i ·[f fusion (t),h(t - 1)]+b i )
[0067] f(t)=σ(W f ·[f fusion (t),h(t - 1)]+b f )
[0068] o(t)=σ(W o ·[f fusion (t),h(t - 1)]+b o )
[0069] C(t)=f(t)·C(t - 1)+i(t)·tanh(Wc · [f fusion (t), h(t - 1)] + b C )
[0070] h(t) = o(t) · tanh(C(t))
[0071] W i , W f , W o , W c are weight matrices, b i , b f , b o , b C is the bias term. σ is the sigmoid activation function. This patent initializes W i , W f , W o , W c , using the overall distribution with a mean of 0 and a standard deviation of : where n is the number of connections between the input and the previous hidden state;
[0072] Use a fully connected layer to map the hidden state h(t) at the last moment to the intention prediction probability
[0073] p k = σ(W k · h(t) + b k )
[0074] (3) The attention mechanism calculates the final predicted intention. First, calculate the attention weights. After being processed by each LSTM module, the output of each stage is an intention probability vector p k ; Use a learnable weight vector related to the stage output to calculate the attention score:
[0075] e k = W att · h t + b att ;
[0076] where W att and b att are the parameters to be trained;
[0077] Use softmax to normalize the score to the weight αk:
[0078]
[0079] Then perform the final intention probability fusion. According to the attention weight α k weighted sum the prediction probabilities p k of each stage:
[0080]
[0081] Set the threshold τ = 0.8. When P final ≥ τ, then predict the getting-up intention as Yes, otherwise as No.
[0082] Furthermore, it further includes step S5: After recognizing the user's getting-up intention, output the recognition result to the user terminal.
[0083] This application also provides a multi-modal getting-up intention recognition device adapted to the characteristics of the elderly, including:
[0084] The first data acquisition device is used to acquire the user's millimeter-wave radar data and generate a point cloud, and map the distance information into point cloud coordinates;
[0085] The second data acquisition device is used to acquire the user's visual data, identify human key points and object bounding boxes from the user's visual data, and judge the user's getting-up intention according to the overlapping relationship between the human key points and the object bounding boxes;
[0086] The data fusion module is used to align and fuse the millimeter-wave radar feature vector and the visual feature vector to obtain the fused feature vector;
[0087] And the getting-up intention recognition module is used to construct a hierarchical temporal attention model HTAM, and identify the user's getting-up intention through the temporal attention model HTAM.
[0088] This application also provides a getting-up intention recognition terminal, including a processor and a memory. When the computer program stored in the memory is called and executed by the processor, it implements the multi-modal getting-up intention recognition method adapted to the characteristics of the elderly as described above.
[0089] This application also provides a computer-readable medium. When the computer program stored in the computer-readable medium is called and executed by the computer, it implements the multi-modal getting-up intention recognition method adapted to the characteristics of the elderly as described above.
[0090] In summary, the beneficial effects of this application are:
[0091] 1. The present invention provides a more accurate, efficient, and adaptable data input for recognizing the intention of the elderly to get up through the fusion of different modal sensing technologies of millimeter-wave radar and visual cameras. Combining the auxiliary information of the grasping features of the handrail beside the bed, the ability to recognize the intention of the elderly to get up is improved by introducing a multi-stage hierarchical temporal model HTAM (Hierarchical Temporal Attention Model). It has broad market application prospects. This system can not only enhance the safety of elderly care but also reduce the burden on caregivers, providing better care services for the elderly.
[0092] 2. In the present invention, by decomposing the intention of the elderly to get up into multiple stages, the HTAM model can more precisely capture the progress of the intention, which is suitable for recognizing the subtle changes in the elderly getting up; combining the characteristics of millimeter-wave and visual data enriches the input information, thus significantly improving the accuracy and robustness of recognition; using the LSTM temporal model can capture the temporal dependencies, ensuring more continuity and consistency in intention recognition; and realizing personalized judgment according to the hand grasping action characteristics of the elderly getting up. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 It is a schematic diagram of the overall principle of a specific embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0094] The following will describe in detail the specific embodiments of the present application with reference to the accompanying drawings.
[0095] A specific embodiment of the present application provides a multi-modal getting-up intention recognition method suitable for the characteristics of the elderly, as Figure 1 , including the following steps:
[0096] S1: Obtain the millimeter-wave radar data of the user and generate a point cloud, and map the distance information to the point cloud coordinates; the original output data of the millimeter-wave radar includes distance, speed, and acceleration, and the point cloud information is obtained after a fast Fourier transform.
[0097] Step S1 includes the following steps:
[0098] S11: Denoise the millimeter-wave radar data to remove high-frequency noise.
[0099] Due to high-frequency noises such as reflection and multipath interference, the high-frequency noise is removed through a low-pass filter, and the motion-related signals are retained. A Butterworth low-pass filter is selected, which has a smoothing characteristic. Denote the sampling rate fs of the millimeter-wave radar and the cut-off frequency fc as fs / 10. For the sampled data x[n], calculate and update y[n] point by point using the above formula in real time.
[0100] y[n] = b0x[n] + b1x[n - 1] + b2x[n - 2] - a1y[n - 1] - a2y[n - 2]
[0101] Where x[n] is the distance input signal; y[n] is the filtered signal; b0, b1, b2, a1, a2 are filter coefficients.
[0102] S12: Data standardization, standardize the denoised distance, speed, and acceleration data;
[0103] Where μ is the mean and σ is the standard deviation.
[0104] S13 Point cloud generation, map the distance information to point cloud coordinates: According to the distance d and angle information (θ, φ), calculate the position of the three-dimensional point in the polar coordinate system, and map the distance information to point cloud coordinates:
[0105] x = d·cos(φ)·cos(θ)
[0106] y = d·cos(φ)·sin(θ)
[0107] z = d·sin(φ)
[0108] The result (x, y, z) is the point cloud coordinate, where d represents the distance, the horizontal angle θ, and the vertical angle φ.
[0109] S2: Obtain the user's visual data, identify the human body key points and object bounding boxes from the user's visual data, and judge the user's intention to get up according to the overlapping relationship between the human body key points and the object bounding boxes.
[0110] The visual data is mainly used for pose detection and action recognition. Through a visual sensor (such as an RGBD camera), we can identify the pose changes of the elderly and judge whether they have the intention to get up.
[0111] Step S2 includes the following steps:
[0112] S21: Detection of human body key points and the bounding box of the grasped object; The present invention particularly defines several objects (bedside handrail, edge of the chair) that the elderly often grasp when getting up.
[0113] Wherein, the grasped object is several objects that are preset for the user to often grasp when getting up, and the bounding box and class label of the above-mentioned grasped object are identified, expressed as:
[0114] O objects ={(x min ,y min ,x max ,y max ,p)|p ∈ {bedside handrail, chair handrail}};
[0115] The described human key points include 17 human skeletal key points and 4 key points on the hand h hand ;
[0116] {0: Nose, 1: Left eye, 2: Right eye, 3: Left ear, 4: Right ear, 5: Left shoulder, 6: Right shoulder, 7: Left elbow, 8: Right elbow, 9: Left wrist, 10: Right wrist, 11: Left hip, 12: Right hip, 13: Left knee, 14: Right knee, 15: Left ankle, 16: Right ankle, 17: Wrist joint, 18: Root of thumb, 19: Tip of thumb, 20: Tip of index finger}
[0117] Perform pose key point detection through the deep learning model MediaPipe to extract the three-dimensional coordinates (x, y, z) of the joints. To match with the millimeter-wave radar regional feature data, the center points and normal vectors of the torso, thigh region, calf region, and hand region are calculated based on the coordinates of the body key points (excluding hand key points) here. The center point (Centroid) takes the average coordinates of the key points in this region. The normal vector is usually obtained by calculating the cross product between two adjacent vectors.
[0118] S22: Judgment of the intention to get up;
[0119] First, judge whether there is contact by calculating the spatial overlap relationship between the hand key points and the bounding box of the grasped object:
[0120]
[0121] If d min < ε, where ε is the design threshold, it is determined that the hand is in contact with the object. If the hand key points are continuously detected within the object border for multiple frames, it is marked as "possibly grasping";
[0122] The probability P grasp ∈[0, 1] represents whether a grasping action is detected; if P grasp > τ, indicating that the grasping probability is higher than the set threshold τ, then enhance the judgment of the "intention to get up"; if P grasp < τ indicates that the grasping probability is relatively low, and then rely on the judgment of the hierarchical temporal attention model HTAM.
[0123] Among them, the probability P grasp is obtained in the following way:
[0124] According to the judgment result of d min < ε, perform regional judgment on 4 hand feature points. If the hand key points are within the armrest border, the contact point count +1;
[0125] Calculate the bending angle of the finger:
[0126] vector1 = fingertip coordinate - knuckle coordinate;
[0127] vector2 = knuckle coordinate - wrist coordinate;
[0128] Bending angle = angle_between(vector1, vector);
[0129] Define feature weights, with contact accounting for 70% and bending accounting for 30%;
[0130] After normalizing each eigenvalue, perform weighted summation and then convert it to probability P through the Sigmoid function grasp :
[0131] Normalize the contact points: Contact = number of contact points / 4;
[0132] Normalize the bending: Bend = 1 - bending angle / 180;
[0133] Weighted summation: Raw = 0.7 * Contact + 0.3 * Bend;
[0134] Sigmoid probability output: Pgrasp = 1 / (1 + e^(-0.5 * Raw - 0.5)).
[0135] S3: Align and fuse the millimeter-wave radar feature vector and the visual feature vector to obtain the fused feature vector.
[0136] The dimensions of the millimeter-wave radar feature vector and the visual feature vector are usually inconsistent because there are significant differences in data types, resolutions, and scales between millimeter-wave radar data and visual data. To perform weighted processing during fusion, it is usually necessary to align the dimensions or perform embedding mapping on these two feature vectors to ensure consistent feature representation and dimensions.
[0137] Step S3 includes the following steps:
[0138] S31: Data alignment;
[0139] Calibrate between the millimeter-wave radar and the visual camera to determine the key point coordinate mapping relationship between the two: In a unified coordinate system, calibrate the point cloud coordinates (x radar , y radar , z radar ) of the millimeter-wave radar and the key point coordinates (x vision , y vision , z vision ) of the visual camera; Use the calibration tool to obtain the corresponding point set offline and calculate the transformation matrix T;
[0140] For the specific calibration process, the relative positions of the camera and the millimeter-wave radar are fixed. If they change, recalibration is required. The specific calibration process is as follows:
[0141] Define the target: Set 9 sampling points, and sample the corresponding three-dimensional points in the millimeter-wave radar coordinate system and the visual coordinate system respectively; and record that the transformation matrix T consists of the rotation matrix R and the translation vector t; such that Pvisual = R Pradar + t, where Pvisual is the position of the visual sampling point in the visual coordinate system; Pradar is the position of the radar sampling point in the radar coordinate system.
[0142] Among them, the coordinates of the 9 sampling points are P radar,i = [x radar,i , y radar,i , z radar,i T and [x vision,i , y vision,i , z vision,i T ;
[0143] Construct the objective function: To minimize the error between the transformed points and the actual visual points, define the error function as:
[0144]
[0145] Solve by the least squares method: First, expand the error: Take the derivatives of R and t respectively and set them to zero to obtain the minimum value; among them, the rotation matrix R can be solved by SVD,
[0146] First, sample points at 9 different positions and construct two matrices:
[0147] P radar = [P radar,1 , P radar,2 ,..., P radar,9
[0148] P vision = [P vision,1 , P vision,2 ,..., P vision,9
[0149] Calculate the covariance matrix:
[0150] H = P vision T P radar
[0151] Perform SVD decomposition on the H matrix:
[0152] H = UΣV T ;
[0153] The SVD decomposition is implemented through the prior art, for example, calculated through the numpy library in python. The calculation example code is as follows:
[0154] import numpy as np
[0155] H = np.array([[3, 2], [2, 3]])
[0156] U, S, Vt = np.linalg.svd(H)
[0157] print("U:", U)
[0158] print("Σ:", np.diag(S))
[0159] print("V:", Vt.T) # Vt is the transpose of V.
[0160] Finally, the rotation matrix R = VU T , the translation vector
[0161] Align the millimeter-wave radar point cloud data with the visual human key point information. For each visual key point P vision ( i ), search for the nearest point P radar ( j ) in the radar point cloud; if a visual key point corresponds to multiple point cloud clusters, take the average value; judge whether the recognition result of the vision at this point is invalid according to a confidence value provided by mediapipe for each key point recognition; if the confidence of this key point is lower than the set threshold τ = 0.5, it is considered that this key point is invalid; then the historical data of the previous time needs to be used as the visual virtual coordinates of this key point for matching with the radar point cloud data.
[0162] Step S32: Feature extraction and matching:
[0163] In each period t, extract the feature vectors of the millimeter-wave point cloud and the visual camera respectively:
[0164] f radar (t) = [P radar,1 , P radar,2 ,..., P radar,17
[0165] f vision (t) = [P vision,1 , P vision,2 ,..., P vision,17
[0166] Feature weighted fusion:
[0167] f fusion = a * f radar + (1 - a) * f vision ; α is a parameter; due to the indoor light scene, α is taken as 0.6.
[0168] The fused feature vector f fusion Through time series stacking, it is input into the hierarchical temporal attention model HTAM, and then sequentially input into the LSTM modules of each stage to complete multi-stage dynamic intention recognition.
[0169] S4: Construct the hierarchical temporal attention model HTAM, combine the attention mechanism and the hierarchical structure on the basis of the temporal model, and identify the user's getting-up intention through the temporal attention model HTAM.
[0170] In the "multi-stage intention recognition algorithm", predicting the progress of the elderly's getting-up intention through a temporal model (such as LSTM or GRU) can help solve the problem of difficult recognition caused by slow movement and multiple pauses during the elderly's getting-up process. However, the standard LSTM still has certain limitations in dealing with multi-stage intention recognition. For example, it only considers the dependency relationship within a short time and cannot well capture the complex stage changes during the elderly's getting-up process. Therefore, an improved "multi-stage temporal intention recognition model" is proposed, which combines the attention mechanism and the hierarchical structure to improve the recognition ability of the elderly's getting-up intention.
[0171] In the hierarchical temporal attention model HTAM:
[0172] 1. The intention prediction process of the current stage is as follows:
[0173] (1) The input data is the feature sequence Fk fused from the millimeter-wave radar and the visual camera, representing the temporal feature segment of this stage, with a shape of (m, 17). Where m is the time window length; in this patent, m is 100 ms.
[0174] (2) The hierarchical temporal attention model HTAM includes multi-stage LSTM modules. The LSTM gradually processes each feature vector f fusion (t) of the time series through internal gating units, and maintains a dynamic hidden state h(t) and cell state C(t); the purpose of the input gate i(t) is to determine which information should be added to the cell state, the forget gate determines which information is discarded from the cell state, and the output gate o(t) determines the update of the hidden state; the cell state is the core of the LSTM, storing long-term information, and the hidden state is the output of the LSTM, used to pass information downstream.
[0175] i(t) = σW i · [f fusion(t), h(t - 1)] + b i )
[0176] f(t) = σ(W f ·[f fusion (t), h(t - 1)] + b f )
[0177] o(t) = σ(W o ·[f fusion (t), h(t - 1)] + b o )
[0178] C(t) = f(t)·C(t - 1) + i(t)·tanh(W c ·[f fusion (t), h(t - 1)] + b C )
[0179] h(t) = o(t)·tanh(C(t))
[0180] W i , W f , W o , W c are weight matrices, b i , b f , b o , b C are bias terms. σ is the sigmoid activation function. This patent initializes W from the standard normal distribution i , W f , W o , W c , using the overall distribution with a mean of 0 and a standard deviation of : where n is the number of connections between the input and the previous hidden state; typically n = d + h. In this patent, d is the input dimension, which is 17, and h is the hidden state dimension, taking 64. n = 81. b i , b f , b o , b C The bias terms affect the convergence of the model and are usually initialized to a small constant. This patent considers initializing them to 0.1, which helps to avoid the saturation state of the model in the initial stage, especially for the Sigmoid activation function.
[0181] Use a fully connected layer to map the hidden state h(t) at the last moment to the intention prediction probability. In the multi-stage model, the intention prediction at each stage obtains a probability distribution p(t) through the output (hidden state) of the LSTM, which represents the possibility of the getting-up intention at that stage. Since the dimension of the hidden state is 64 and the dimension of the output intention category is 2 ("getting-up intention" and "no getting-up intention"), so W k is a 2x64 weight matrix, and b k is a 2-dimensional bias vector. Here, it is considered that the number of input nodes is 51 and the number of output nodes is 2.
[0182] p k = σ(W k ·h(t) + b k );
[0183] 2. The attention mechanism calculates the final predicted intention. The key role of the attention mechanism is to dynamically weight according to the relative importance of the outputs of the LSTM at each stage. In this way, the model can pay more attention to those stages that have a greater impact on the final judgment.
[0184] The attention mechanism is a technique that can dynamically weight input information. For the multi-stage intention recognition task, the attention mechanism can assign different weights according to the prediction importance of each stage, enabling the model to pay more attention to the stage predictions that contribute more to the final intention decision. It calculates the attention scores of each stage and combines the intention probabilities of each stage weighted according to these scores to finally make a more accurate prediction. In multi-stage intention prediction, the advantages of using the attention mechanism are mainly reflected in the following aspects:
[0185] ① Dynamic weighted attention: The predictions at different stages contribute differently to the final intention decision at different time points. Some stages may be more important in certain situations, while other stages may be relatively less important. Through the attention mechanism, the model can dynamically adjust the weight of each stage according to the current task requirements, thereby more accurately judging the getting-up intention of the elderly.
[0186] ② Reduce the interference of redundant information: The prediction of each stage may contain redundant information. The attention mechanism can effectively reduce the interference of redundant information to the decision-making by weighted selection of useful stage outputs, especially when some stages have no obvious contribution to the decision-making, the attention mechanism can reduce the influence of these stages.
[0187] ③ Enhance the attention to key stages: The recognition of the getting-up intention is usually a gradual process. Some stages (such as when the elderly's body starts to turn or when holding an object) may be more critical for the final judgment. Through the attention mechanism, the model can adaptively strengthen the attention to these key stages and improve the prediction accuracy.
[0188] ④Improve the interpretability of the model: The attention mechanism provides a way to visualize the impact of each stage on the final prediction. By examining the attention weights at each stage, we can intuitively understand which stages play important roles in the judgment results under different scenarios. This helps to enhance the interpretability of the model, especially in practical applications, particularly in fields such as elderly health monitoring, where understanding how the model makes decisions is very important.
[0189] ⑤Resolve multi-stage decision conflicts: In a multi-stage model, different stages may have different prediction results, especially when there is noise or uncertainty in the data. The attention mechanism can alleviate the conflicts between the prediction results of different stages by weighted aggregation of information from multiple stages, and finally arrive at a more reliable and stable decision.
[0190] (1) First, calculate the attention weights. After being processed by each LSTM module, the output of each stage is an intention probability vector p k ; Use a learnable weight vector related to the stage output to calculate the attention scores:
[0191] e k = W att ·h t + b att ;
[0192] where W att and b att are parameters to be trained, and the initialization is the same as above;
[0193] Use softmax to normalize the scores into weights αk:
[0194]
[0195] (2) Then, perform the final intention probability fusion. According to the attention weights α k weighted sum the prediction probabilities p k of each stage:
[0196]
[0197] Set the threshold τ = 0.8. When P final ≥ τ, then predict the getting-up intention as Yes, otherwise as No.
[0198] By introducing a hierarchical structure, the model decomposes the elderly's getting-up action into different sub-stages; at the same time, adding the attention mechanism enables the model to focus on the key information during the getting-up process to improve the accurate judgment of whether the elderly are really about to get up.
[0199] Step S5: After recognizing the user's intention to get up, output the recognition result to the user terminal. If it is recognized that the elderly person has the intention to get up, notify the nurse station to intervene in the care in time to ensure the safety of the elderly person. In a possible embodiment, access the edge side through wifi and notify the nurse station of the recognition result information.
[0200] Another specific embodiment of the present application provides a multi-modal getting-up intention recognition device adapted to the characteristics of the elderly, including:
[0201] The first data acquisition device is used to acquire the millimeter-wave radar data of the user and generate a point cloud, and map the distance information into point cloud coordinates;
[0202] The second data acquisition device is used to acquire the visual data of the user, identify the human body key points and object bounding boxes from the user's visual data, and judge the user's getting-up intention according to the overlapping relationship between the human body key points and the object bounding boxes;
[0203] The data fusion module is used to align and fuse the millimeter-wave radar feature vector and the visual feature vector to obtain the fused feature vector;
[0204] And the getting-up intention recognition module is used to construct a hierarchical temporal attention model HTAM, and recognize the user's getting-up intention through the temporal attention model HTAM.
[0205] Another specific embodiment of the present application provides a getting-up intention recognition terminal, including a processor and a memory. When the computer program stored in the memory is called and executed by the processor, it implements the multi-modal getting-up intention recognition method adapted to the characteristics of the elderly as described above.
[0206] Another specific embodiment of the present application provides a computer-readable medium, and the computer-readable medium stores a computer program. When the computer program is called and executed by a computer, it implements the multi-modal getting-up intention recognition method adapted to the characteristics of the elderly as described above.
[0207] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A multi-modal getting-up intention recognition method adapted to the characteristics of the elderly, characterized in that It includes the following steps: S1: Obtain the millimeter-wave radar data of the user and generate a point cloud, and map the distance information into point cloud coordinates; S2: Obtain the visual data of the user, identify the human body key points and object bounding boxes from the user's visual data, and judge the user's getting-up intention according to the overlapping relationship between the human body key points and the object bounding boxes; S3: Align and fuse the millimeter-wave radar feature vector and the visual feature vector to obtain the fused feature vector; S4: Construct a hierarchical temporal attention model HTAM, combine the attention mechanism and the hierarchical structure on the basis of the temporal model, and identify the user's getting-up intention through the temporal attention model HTAM.
2. The multimodal getting-up intention recognition method adapted to the characteristics of the elderly according to claim 1, characterized in that, Step S1 includes the following steps: S11: Denoise the millimeter-wave radar data to remove high-frequency noise; S12: Data standardization, standardize the denoised distance, speed, and acceleration data; S13: Calculate the position of the three-dimensional point in the polar coordinate system according to the distance and angle information, and map the distance information into point cloud coordinates.
3. The multimodal getting-up intention recognition method adapted to the characteristics of the elderly according to claim 1, wherein, Step S2 includes the following steps: S21: Detect human body key points and the bounding boxes of the grabbed objects; Among them, the grabbed objects are several objects that the user often grabs when getting up as preset, and the bounding boxes and class labels of the above grabbed objects are identified, expressed as: O objects = {(x min , y min , x max , y max , p) | p ∈ {bedside armrest, chair armrest}}; The described human body key points include 17 human body skeletal key points and 4 key points h of the hand hand ; S22: Judge the getting-up intention; First, judge whether there is contact by calculating the spatial overlapping relationship between the hand key points and the bounding boxes of the grabbed objects: If d min < ε, where ε is the design threshold, it is determined that the hand is in contact with the object. If the hand key points are continuously detected within the object border for multiple frames, it is marked as "grasp possible"; Through probability P grasp ∈ [0, 1] indicates whether a grasping action is detected; if P grasp > τ, indicating that the grasping probability is higher than the set threshold τ, then enhance the judgment of the "standing up intention"; if P grasp < τ indicates that the grasping probability is low, then make a judgment according to the hierarchical temporal attention model HTAM.
4. The multimodal getting-up intention recognition method adapted to the characteristics of the elderly according to claim 3, characterized in that, In step S22, the probability P grasp is obtained in the following manner: According to d min Based on the judgment result of ε, perform area judgment on 4 hand feature points. If the hand key point is within the armrest border, the contact point count is incremented by 1; Calculate the bending angle of the finger: vector1 = fingertip coordinate - finger root coordinate; vector2 = finger root coordinate - wrist coordinate; Bending angle = angle_between(vector1, vector); Define the feature weights, with the contact accounting for 70% and the bending accounting for 30%; After normalizing each eigenvalue, perform weighted summation and then convert it into probability P through the Sigmoid function grasp : Contact normalization: Contact = number of contact points / 4; Bending normalization: Bend = 1 - bending angle / 180; Weighted summation: Raw = 0.7 * Contact + 0.3 * Bend; Sigmoid probability output: P grasp = 1 / (1 + e^(-0.5 * Raw - 0.5)).
5. The multimodal getting-up intention recognition method adapted to the characteristics of the elderly according to claim 1, characterized in that, Step S3 includes the following steps: S31: Data alignment; Calibrate between the millimeter-wave radar and the vision camera to determine the key-point coordinate mapping relationship between the two: In a unified coordinate system, calibrate the point-cloud coordinates (x radar , y radar , z radar ) of the millimeter-wave radar and the key-point coordinates (x vision , y vision , z vision ) of the vision camera; Use the calibration tool to obtain the corresponding point set offline and calculate the transformation matrix T; The calibration process is as follows: Define the target: Set 9 sampling points, and sample the corresponding three-dimensional points in the millimeter-wave radar coordinate system and the visual coordinate system respectively; and record that the transformation matrix T is composed of the rotation matrix R and the translation vector t; so that Pvisual = R Pradar + t. Among them, the coordinates of 9 sampling points are P radar,i = [x radar,i , y radar,i , z radar,i T and [x vision,i , y vision,i , z vision,i T ; Construct the objective function: In order to minimize the error between the transformed points and the actual visual points, the error function is defined as: Solve by the least squares method: First expand the error: Take the derivatives of R and t respectively and set them to zero to obtain the minimum value; among them, the rotation matrix R can be solved by SVD, First, sample points at 9 different positions and construct two matrices: P radar = [P radar,1 , P radar,2 ,..., P radar,9 P vision = [P vision,1 , P vision,2 ,..., P vision,9 Calculate the covariance matrix H = P vision T P radar Perform SVD decomposition on the H matrix H = U∑V T The final rotation matrix R = VU T , the translation vector Align the millimeter-wave radar point cloud data with the visual human key point information, and for each visual key point P through the KNN (K-Nearest Neighbor) search algorithm vision ( i ), search for the nearest point P radar ( j ) in the radar point cloud; if a visual key point corresponds to multiple point cloud clusters, take the average value; judge whether the recognition result of vision at this point is invalid according to a confidence value provided by mediapipe for each key point recognition; if the confidence of this key point is lower than the set threshold τ = 0.5, it is considered that this key point is invalid; then the historical data of the previous time needs to be used as the visual virtual coordinate of this key point for matching with the radar point cloud data; Step S32: Feature extraction and matching: In each period t, extract the feature vectors of the millimeter-wave point cloud and the visual camera respectively: f radar (t) = [P radar,1 , P radar,2 ,..., P radar,17 f vision (t) = [P vision,1 , P vision,2 ,..., P vision,17 Feature weighted fusion: f fusion = a * f radar + (1 - a) * f vision ; where α is a parameter; The fused feature vector f fusion Through time series stacking, it is input into the hierarchical temporal attention model HTAM, and then successively input into the LSTM modules of each stage to complete multi-stage dynamic intention recognition.
6. The multimodal getting-up intention recognition method adapted to the characteristics of the elderly according to claim 1, wherein, In the hierarchical temporal attention model HTAM: (1) The input data is the feature sequence Fk fused from the millimeter-wave radar and the visual camera, representing the temporal feature segment of this stage, with a shape of (m, 17); where m is the length of the time window; (2) The hierarchical temporal attention model HTAM includes multi-stage LSTM modules. The LSTM processes each feature vector f of the time series step by step through internal gating units, and maintains a dynamic hidden state h(t) and cell state C(t). The purpose of the input gate i(t) is to determine which information should be added to the cell state, the forget gate determines which information is discarded from the cell state, and the output gate o(t) determines the update of the hidden state. fusion (t), and maintains a dynamic hidden state h(t) and cell state C(t); the purpose of the input gate i(t) is to determine which information should be added to the cell state, the forget gate determines which information is discarded from the cell state, and the output gate o(t) determines the update of the hidden state; i(t) = σ(W i ·[f fusion (t), h(t - 1)] + b i ) f(t) = σ(W f ·[f fusion (t), h(t - 1)] + b f ) o)t) = σ(W o ·[f fusion (t), h(t - 1)] + b o ) C(t) = f(t)·C(t - 1) + i(t)·tanh(W c ·[f fusion (t), h(t - 1)] + b C ) h(t) = o(t) · tanh(C(t)) W i ,W f ,W o ,W c is the weight matrix, b i ,b f ,b o ,b C is the bias term; σ is the sigmoid activation function; initialize W from the standard normal distribution i ,W f ,W o ,W c , using the overall distribution with mean 0 and standard deviation : where n is the number of connections between the input and the previous hidden state; Use a fully connected layer to map the hidden state h(t) at the last moment to the intention prediction probability: p k = σ(W k ·h(t) + b k )); (3) The attention mechanism calculates the final predicted intention. First, it calculates the attention weights. After being processed by each LSTM module, the output at each stage is an intention probability vector p k ; A learnable weight vector related to the stage output is used to calculate the attention score: e k = W att ·h t + b att ; Among which W att and b att are parameters to be trained; Use softmax to normalize the scores to weights αk: Then perform final intention probability fusion, according to the attention weight α k for the prediction probabilities p k at each stage and perform weighted summation: Set a threshold τ. When P final ≥ τ, then predict that the getting-up intention is Yes; otherwise, it is No.
7. The multi-modal getting-up intention recognition method adapted to the characteristics of the elderly according to claim 1, characterized in that It further includes step S5: after recognizing the user's getting-up intention, output the recognition result to the user terminal.
8. A multi-modal getting-up intention recognition device adapted to the characteristics of the elderly, characterized in that, It includes: The first data acquisition device is used to acquire the user's millimeter-wave radar data and generate a point cloud, and map the distance information to the point cloud coordinates; The second data acquisition device is used to acquire the user's visual data, identify human key points and object bounding boxes from the user's visual data, and judge the user's getting-up intention according to the overlapping relationship between the human key points and the object bounding boxes; The data fusion module is used to align and fuse the millimeter-wave radar feature vector and the visual feature vector to obtain the fused feature vector; And the getting-up intention recognition module is used to construct a hierarchical temporal attention model HTAM, and recognize the user's getting-up intention through the temporal attention model HTAM.
9. A terminal for recognizing the intention of getting up, characterized in that, It includes a processor and a memory. When the computer program stored in the memory is called and executed by the processor, it implements the multi-modal getting-up intention recognition method adapted to the characteristics of the elderly as described in any one of claims 1-7.
10. A computer-readable medium, characterized in that, The computer-readable medium stores a computer program. When the computer program is called and executed by the computer, it implements the multi-modal getting-up intention recognition method adapted to the characteristics of the elderly as described in any one of claims 1-7.
Citation Information
Patent Citations
A method for monitoring the intention to get up on a smart nursing bed
CN109833045B
A method for identifying human behavioral intentions based on object detection networks and knowledge reasoning
CN114724078B
Millimeter wave radar-based non-contact real-time vital sign monitoring system and method
ZA202305013B
Cited By
Multi-target hand posture estimation method based on millimeter wave radar
CN120523331A