Multi-scene gait monitoring and motion function evaluation method and system

Through a multi-camera collaborative system and a dual verification mechanism, combined with a long short-term memory network to evaluate the gait function of the elderly, the environmental adaptability and individual recognition problems of gait monitoring in multiple scenarios are solved, and accurate motor function assessment and early pathological feature detection are achieved.

CN120605004APending Publication Date: 2025-09-09HEBEI UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510723187.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies have poor environmental adaptability for gait monitoring of the elderly in multiple scenarios, making it difficult to accurately distinguish individuals and effectively evaluate the functional relevance of high-frequency movements, resulting in distorted evaluation results and data confusion.

Method used

A multi-camera collaborative system is used, combined with a Gaussian mixture model and the OpenPose algorithm to extract joint and skeleton key points. Identity matching is performed using a dual verification mechanism of face recognition and skeleton geometric features. The motion function is evaluated through a long short-term memory network, and a risk suppression mechanism is designed for comprehensive scoring.

Benefits of technology

It has achieved accurate gait monitoring and motor function assessment of the elderly in multiple scenarios, improved the accuracy of identity recognition and the reliability of data association, and can detect pathological characteristics early and provide personalized intervention recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120605004A_ABST
    Figure CN120605004A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scene gait monitoring and motion function assessment method and system, a multi-camera cooperation system is deployed in a life scene to accurately identify an identity and extract multi-scene motion features, and an accurate old people gait function assessment scheme is provided through key motion hierarchical modeling, multi-scene data fusion and a risk suppression mechanism. Specifically, 2D human body key points are estimated based on an OpenPose algorithm, and a PoseLifter model is utilized to lift the 2D key points to 3D, so that a three-dimensional motion track of a target is estimated; dynamically adjusting the weights of the face features and the skeleton features during identity recognition according to ambient light; designing an independent long-short-term memory network for different actions to evaluate the motion function score and the fall risk rating of the actions; and designing a risk sensitivity mechanism and fusing multi-scene data to carry out comprehensive scoring. The gait function abnormity of the old people can be found in time, and a basis is provided for health management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and intelligent health monitoring, and in particular to a multi-scenario gait monitoring and motor function assessment method and system. Background Art

[0002] As the global population ages, the health of the elderly is receiving increasing attention. As an important indicator of their health status, the ability to carry out daily activities urgently requires accurate and convenient assessment methods. However, existing methods, which primarily rely on manual observation or wearable sensors, have the following limitations:

[0003] 1. Poor environmental adaptability: Wearable devices are easily disturbed by the activities of multiple people and cannot cover multiple scenarios such as bedrooms, corridors, and activity rooms. The laboratory environment is significantly different from the real scene, resulting in distorted evaluation results.

[0004] 2. Single-minded action recognition: Existing algorithms mostly focus on basic actions such as walking and standing, ignoring the functional relevance of high-frequency actions such as bed-to-chair transfers and avoiding narrow areas.

[0005] 3. Individual confusion problem: When multiple people are engaged in parallel activities, traditional computer vision algorithms have difficulty accurately distinguishing individuals, resulting in incorrect data attribution.

[0006] Therefore, there is an urgent need for a gait monitoring system that can adapt to complex scenarios and support accurate identification of multiple people, so as to provide early warning and personalized intervention support for the decline of motor function in the elderly. Summary of the Invention

[0007] The purpose of this application is to address the technical deficiencies in existing solutions and provide a multi-scenario gait monitoring and motor function assessment method and system. A multi-camera collaborative system is deployed in key life scenarios to accurately identify identities and extract multi-scenario motion features. Through layered modeling of key actions, multi-scenario data fusion, and risk suppression mechanisms, the aforementioned technical issues are addressed, providing an accurate gait function assessment solution for the elderly. This invention can promptly detect gait function abnormalities in the elderly, providing a basis for health management.

[0008] The present invention is achieved through the following technical solutions:

[0009] In a first aspect, the present invention provides a multi-scenario gait monitoring and motor function assessment method, comprising the following steps:

[0010] Step 1: Deploy a multi-camera collaborative monitoring system to collect real-time video data of the elderly in multiple life scenarios. Extract key point information of the elderly's joints and skeleton from the video data and establish the three-dimensional motion trajectory of the key points.

[0011] Step 2: Match and associate the video data and key point 3D motion trajectory data obtained in step 1 with the individual identity information of the elderly;

[0012] Step 3: For different scenarios, extract the motion parameters of the corresponding actions and design an independent long-short-term memory network model for each action. The input of this model is the action parameter sequence, and the output is the motion function score and fall risk rating of the action;

[0013] Step 4: Integrate multi-scenario motion data and dynamically predict the motion parameters of the next monitoring cycle based on the action risk level. The designed method uses the difference between the predicted and actual values ​​of the motion parameters of the next cycle and the motion function score of each action to obtain the comprehensive motion function score of multiple scenarios.

[0014] In the above technical solution, in step 1, a Gaussian mixture model is used to separate moving objects from the captured video, the OpenPose human pose estimation algorithm is applied to extract the joint and skeleton key point information of the elderly, and the PoseLifter model is used to predict the position of the key points in three-dimensional space; the key point positions are tracked in time series to obtain the three-dimensional key point coordinates in each frame of the video, and the Kalman filter is used to smooth the trajectory of the key points.

[0015] In the above technical solution, in step 2, a dual verification mechanism of face recognition and human skeleton geometric features is used to accurately match and associate the video and key point three-dimensional motion trajectory data obtained in step 1 with the individual identity information of the elderly.

[0016] In the above technical solution, step 2 includes the following steps:

[0017] S2.1: Extract facial features from video data and calculate the cosine similarity S between the facial features extracted from the video data and the facial features stored in the identity database face ;

[0018] S2.2: Decompose the joints and skeleton key points extracted from the video data into multiple groups of vector chains, calculate the dynamic length ratio and vector angle of each group of chains, and obtain the skeleton features;

[0019] Calculate the Mahalanobis distance d between the skeleton features in the video data and the sample mean of the skeleton feature sample library of each identity pre-stored in the identity database mah ; Then calculate the Mahalanobis distance d mah and the skeleton feature reference value d of the identity in the identity database base The similarity S skeleton ;

[0020] S2.3: The cosine similarity S of facial features face Similarity S with skeleton featuresskeleton The weights are dynamically set according to the light intensity for fusion to obtain the fused similarity score ID-Score; the identity of the person is determined based on the fused similarity score ID-Score.

[0021] In the above technical solution, multiple sets of skeleton feature samples are pre-entered for each identity in the identity database to construct a skeleton feature sample library for each identity; and for each identity, the Mahalanobis distance between each sample in its skeleton feature sample library and the sample average of the skeleton feature sample library is calculated; then, for each identity, the 95% quantile of all the calculated Mahalanobis distances is taken as the normalized reference value, thereby obtaining the reference value d of the skeleton feature corresponding to each identity. base .

[0022] In the above technical solution, step 3 includes the following steps:

[0023] S3.1: Identify the following specific actions in each scenario: standing balance and sit-to-stand transition in the bedroom; normal walking, emergency stop, standing balance, and turning in the hallway; sit-to-stand transition, standing balance, and avoidance path deviation in the reading room and living room.

[0024] S3.2: Calculate the corresponding action parameters for each action;

[0025] S3.3: Establish a multi-input and multi-output dynamic weighted branch network model. The model input is the motion parameter sequence of each action in three consecutive gait cycles, and the output is two branches: the motor function score of the action and the fall risk rating.

[0026] In the above technical solution, in step S3.2:

[0027] For standing balance, calculate the standard deviation σ of the pelvic center coordinates p =(σ x ,σ y ,σ z ), calculate the 95% confidence ellipse area based on the covariance matrix of the center of gravity trajectory;

[0028] For sit-stand transfers, calculate the angle change, hip flexion and extension moments, and the ratio of the vertical displacement of the center of mass of the trunk to the total displacement using inverse dynamics.

[0029] For normal walking movements, the correlation coefficient of the left and right lower limb joint angle curves and the standard deviation of the step length for 10 consecutive steps were calculated;

[0030] For emergency stopping actions, calculate braking distance and angular velocity mutation;

[0031] For turning movements, the rotation radius, the integrated value of the trunk angular velocity, and the phase difference of the shoulder-hip angle change were calculated;

[0032] For evasive path deviation actions, the path deviation degree, speed adjustment rate, and the rate of change of the angle between the actual trajectory and the preset straight line are calculated.

[0033] In the above technical solution, step S3.3 includes the following steps:

[0034] S3.3.1: The multi-input feature processing layer is designed with six independent input branches, corresponding to the six action types: standing balance, sit-stand transition, normal walking, emergency stop, turn action, and avoidance path deviation action. The input of each branch is a sequence of action parameters, and the sequence length is unified through dynamic time warping.

[0035] S3.3.2: The shared-private feature encoding layer is designed to consist of two parts: a shared feature extractor and an action-specific feature extractor. The shared feature extractor consists of fully connected layers and is responsible for extracting shared features of actions. The action-specific feature extractor is designed for each action with a separate bidirectional long short-term memory network that incorporates an attention mechanism and adds dilated convolution to capture long-range dependencies. The action-specific feature extractor is responsible for learning the specific features of each action.

[0036] S3.3.3: The dynamic weight module receives shared features and specific features after the shared-private feature encoding layer. It compresses the shared features into a vector using global average pooling, integrates the spatial information of the entire feature map, and preserves the overall pattern of the features. It uses a multi-layer perceptron to map the global features into a 32-dimensional weight vector. It has two fully connected layers: the first layer converts the C-dimensional vector into a 16-dimensional vector, and the second layer converts the 16-dimensional vector into a 32-dimensional vector. It then calculates the weights of the features of each action branch and fuses the shared features with the specific features.

[0037] S3.3.4: The dual-task output layer is designed as two output branches: motion scoring branch and risk rating branch; the motion scoring branch is composed of a fully connected layer, which outputs the score of the action, and designs a dynamic structural similarity loss function and an action-specific regularization loss; the main body of the risk rating branch is composed of a three-layer cascade classifier, which divides the risk into five levels of 0-4 through the Softmax activation function, and designs a dynamic focus cross entropy loss as the loss function of the risk rating branch, combining JS divergence and KL divergence to constrain the semantic consistency of the motion scoring and risk rating branches.

[0038] In the above technical solution, step 4 includes the following steps:

[0039] S4.1: Design a dynamic risk sensitivity prediction model to dynamically adjust the prediction strategy based on the risk level of the current action and output the motion parameters for the next posture detection cycle;

[0040] S4.2: Calculate the predicted value of motion parameters for the next monitoring period The difference MSE from the true value y;

[0041] S4.3: Combine multi-scenario data and historical data to perform a comprehensive motor function score:

[0042]

[0043] Among them, S i represents the motion score of the i-th action, λ total is the adjustment parameter, represents the risk rating weight of the i-th action.

[0044] In the above technical solution, step S4.1 includes the following steps:

[0045] S4.1.1: The model is designed as a dual-input architecture. The two inputs are the concatenation of the average motion parameter sequences of different actions in the current cycle and the risk rating vector of different actions obtained in step 3.

[0046] S4.1.2: After the input layer, we perform spatiotemporal feature extraction on the input motion parameter sequence. We first use a one-dimensional convolutional layer to capture local temporal patterns, then use a Transformer encoder to extract global dependencies and output a fixed-length feature vector F. We then decouple the unified feature F into six parts based on the action: F1, F2, L, and F6.

[0047] S4.1.3: Design a risk suppression module to adjust the weights of different action parameters in feature fusion using a linear or exponential decay function for the risk rating of each action, and at the same time, Calculate the weights in segments; use the weights to calculate the fusion features;

[0048] S4.1.4: Establish independent parameter prediction networks for the six actions, use the fused features as input, and use five fully connected layers and ReLU activation functions to mine the unique patterns of the actions and output predicted values.

[0049] In a second aspect, the present invention further discloses a system for implementing the above-mentioned multi-scenario gait monitoring and motor function assessment method, the system comprising the following parts:

[0050] The monitoring and data collection terminal is used to collect real-time video data of the elderly in multiple life scenarios;

[0051] The computer processing end is used to complete the computational evaluation of the system, including establishing the three-dimensional motion trajectory of key points, identity matching and association, calculating motion parameters, calculating the independent score and fall risk rating of each action, and calculating the comprehensive motion function score in multiple scenarios;

[0052] The computer storage terminal is used to store the independent score and risk level of each action in each scene, the comprehensive score data of multiple scenes, and the predicted and actual values ​​of the motion parameters of each action in the local file corresponding to the identity, so that they can be called up at any time later;

[0053] The computer display terminal is used to display the scene-score-time data curve, the historical data curve of the multi-scene comprehensive score, and the time data curve of the motion parameter predicted value and the actual value, so as to intuitively show the character's gait function status and remind relevant personnel to pay attention to functional decline in time.

[0054] The advantages and beneficial effects of the present invention are:

[0055] This invention is applicable to daily living environments and can conduct gait assessment on the elderly in real scenarios such as homes and communities in a contactless manner. The system is easy to deploy and can achieve long-term continuous monitoring without interfering with the elderly's lives, making it suitable for home-based care and community health management.

[0056] The present invention significantly improves the accuracy of identity recognition and anti-camouflage ability through dual verification of face recognition and skeleton features, accurately associates action data with individual files, and improves the reliability of data association and the applicability of application scenarios.

[0057] The present invention integrates the spatiotemporal characteristics of different scenarios and activity types, dynamically assigns action weights, and can calculate the comprehensive deviation of gait in real time, thereby detecting early pathological characteristics in advance and facilitating early intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0059] Figure 1 It is a flow chart of the multi-scenario gait monitoring and motor function assessment method of the present invention.

[0060] Figure 2 It is a schematic diagram of 8 groups of skeleton vector chains in the present invention.

[0061] Figure 3It is the model framework for predicting the motion parameter sequence of the next monitoring period in the present invention. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0063] Example 1

[0064] This embodiment proposes a multi-scenario gait monitoring and motion function assessment method, deploying a multi-camera collaborative system in life scenarios to accurately identify identities and extract multi-scenario motion features, and providing an accurate gait function assessment solution for the elderly through key action hierarchical modeling, multi-scenario data fusion and risk suppression mechanism. Specifically, based on the OpenPose algorithm, 2D human body key points are estimated, and the PoseLifter model is used to enhance the 2D key points to 3D, thereby estimating the three-dimensional motion trajectory of the target; the weights of facial features and skeleton features in identity recognition are dynamically adjusted according to the ambient lighting; independent long and short-term memory networks are designed for different actions to assess the motion function score and fall risk rating of the action; and a risk sensitivity mechanism is designed to fuse multi-scenario data for comprehensive scoring. The method includes the following steps:

[0065] Step 1: Deploy a multi-camera collaborative monitoring system to collect real-time video data of the elderly in multiple life scenarios, extract key point information of the elderly's joints and skeleton from the video data, and establish the three-dimensional motion trajectory of the key points.

[0066] Exemplarily, step 1 includes the following steps:

[0067] S1.1: The monitoring system is deployed in multiple locations throughout the nursing home, including corridors, bedrooms, reading rooms, and living rooms. A wide-angle camera group (field of view ≥ 120°) is installed on the ceiling of the nursing home, and fisheye cameras (resolution ≥ 4K) are deployed in blind spots such as corridor corners and bedroom doorways. This eliminates monitoring blind spots, establishes cross-coverage areas, and captures real-time motion video data of the elderly.

[0068] S1.2: Use a Gaussian mixture model to separate moving objects from the captured video, apply the OpenPose human pose estimation algorithm to extract key point information such as joints and skeletons of the elderly, and use the PoseLifter model to predict the position of key points in 3D space.

[0069] S1.3: Track the key point positions in time series, obtain the three-dimensional key point coordinates in each frame of video, and use Kalman filtering to smooth the trajectory of the key points.

[0070] Step 2: Using a dual verification mechanism of face recognition and human skeleton geometric features, the video and key point three-dimensional motion trajectory data obtained in step 1 are accurately matched and associated with the individual identity information of the elderly.

[0071] Exemplarily, step 2 includes the following steps:

[0072] S2.1: First, facial features are extracted from the video data. Specifically, this embodiment uses a geometric topology network to take the Delaunay triangulation of the 68 facial key points in the collected video data as the topological structure, and calculates the side length and angle of each triangle to form a 120×6=720-dimensional feature vector, thereby obtaining facial features.

[0073] Calculate the cosine similarity S between the facial features extracted from the video data and the facial features of each identity pre-stored in the identity database face :

[0074]

[0075] in, They represent the facial features extracted from the video data and the facial features stored in the identity database respectively.

[0076] S2.2: First, it should be noted that in addition to pre-storing facial feature data, the identity database also pre-enters multiple sets of skeleton feature samples for each identity, constructing a skeleton feature sample library for each identity. Furthermore, for each identity, the Mahalanobis distance between each sample in its skeleton feature sample library and the sample average of the skeleton feature sample library is calculated using the following formula:

[0077]

[0078] in, Represents the feature vector of the i-th sample in the skeleton feature sample library of a certain identity, μ ref Represents the mean of the feature vectors of all samples in the skeleton feature sample library of this identity (i.e., the sample average value), and the symbol T represents the matrix transpose. Represents the inverse matrix of the sample library feature covariance matrix;

[0079] Then, for each identity, all the Mahalanobis distances calculated were taken as the 95% quantile as the normalized benchmark value, thereby obtaining the benchmark value d of the skeleton feature corresponding to each identity. base ;

[0080] d base =quantile(d ref ,0.95)

[0081] Among them, dref The set of all Mahalanobis distances representing an identity.

[0082] Then, for this step, first obtain the skeleton features in the video data. Specifically, in this embodiment, the joints and skeleton key points extracted from the video data are decomposed into 8 groups of vector chains, namely left shoulder-left elbow-left wrist, right shoulder-right elbow-right wrist, left hip-left knee-left ankle, right hip-right knee-right ankle, neck-left shoulder-left hip, neck-right shoulder-right hip, head-left shoulder-left hip, head-right shoulder-right hip, as shown in the attached figure. Figure 2 As shown, the dynamic length ratio L and vector angle θ of each group of chains are calculated c , thus obtaining a 2×8=16-dimensional skeleton feature vector, that is, the skeleton feature:

[0083]

[0084] in, represents a vector chain, and represents a sub-vector of a vector chain, is the reference vector in the resting standing state.

[0085] Then, the Mahalanobis distance d between the skeleton features in the video data and the sample mean of the skeleton feature sample library of each identity pre-stored in the identity database is calculated. mah ; Then calculate the Mahalanobis distance d mah and the skeleton feature reference value d of the identity in the identity database base The similarity S skeleton :

[0086]

[0087] Where λ is the attenuation coefficient, which is 3.

[0088] S2.3: The cosine similarity S of facial features face Similarity S with skeleton features skeleton The fusion process dynamically sets weights based on illumination intensity, generating a fused similarity score (ID-Score). This score is then used to determine the person's identity. The ID-Score ranges from [0 to 1], with higher values ​​indicating a higher fusion similarity between the person in the video and a known identity in the identity database. A fusion similarity threshold is set; if the ID-Score exceeds the threshold, the identities are considered a match.

[0089] The calculation formula of ID-Score is: ID-Score = ω face ·S face +ω skeleton ·S skeleton ;

[0090] Among them, ω face and ω skeleton They represent the weights of facial features and skeleton features respectively, and the calculation formula is as follows:

[0091] ω skeleton =1-ω face

[0092] Among them, I light Indicates the ambient light intensity, ranging from 0 to 100 lux.

[0093] Step 3: For different scenarios, extract the motion parameters of the corresponding actions and design an independent long short-term memory network (LSTM) for each action. The model input is the action parameter sequence, and the output is the motion function score and fall risk rating of the action.

[0094] Exemplarily, step 3 includes the following steps:

[0095] S3.1: Identify the following specific actions in each scenario: standing balance and sit-to-stand transition in the bedroom scenario; normal walking, emergency stop, standing balance, and turning in the corridor scenario; sit-to-stand transition, standing balance, and avoidance path deviation in the reading room and living room.

[0096] S3.2: Calculate the corresponding action parameters for each action.

[0097] ① For standing balance movements, calculate the standard deviation σ of the pelvic center coordinates p =(σ x ,σ y ,σ z ), calculate the 95% confidence ellipse area based on the covariance matrix of the center of gravity trajectory;

[0098] Furthermore, for standing balance actions in public activity areas (corridors, reading rooms, and living rooms), weighting coefficients are set based on the number of people around and the intensity of the activity;

[0099] Activity intensity is calculated as follows:

[0100]

[0101] Among them, U represents the number of people in the detection area, and the value range is 0-20 (20 is used when there are more than 20 people). is the velocity component of the i-th person in the x and y directions, and T is the sampling time;

[0102] The weighting coefficient is calculated as follows:

[0103]

[0104] Among them, α public is the number of people weight coefficient, take 0.6, β public is the activity intensity weight coefficient, which is 0.4, I max This is the highest activity intensity in history.

[0105] The weighted coefficient is used to adjust the value of the action parameter in the common area. For example, the original calculation formula of the standard deviation of the pelvic center coordinate in the x direction is:

[0106]

[0107] The corrected calculation formula is:

[0108]

[0109] ② For sit-stand transfers, calculate the angle change, the hip flexion and extension torque (unit: N·m) through inverse dynamics, and the ratio of the vertical displacement of the trunk center of mass to the total displacement;

[0110] ③ For normal walking movements, the correlation coefficient of the left and right lower limb joint angle curves (gait symmetry) and the standard deviation of the step length of 10 consecutive steps (variability) were calculated;

[0111] ④ For emergency stopping, calculate the braking distance (center of gravity displacement from maximum speed to standstill) and angular velocity mutation (maximum value of the first-order derivative of ankle joint angular velocity);

[0112] ⑤ For turning movements, calculate the rotation radius, the integral value of the trunk angular velocity, and the phase difference of the shoulder-hip angle change;

[0113] ⑥ For evasive path deviation actions, calculate the path deviation degree and speed adjustment rate, as well as the rate of change of the angle between the actual trajectory and the preset straight line;

[0114] The path deviation is measured by the degree of vertical distance dispersion between the actual trajectory point and the preset straight line:

[0115]

[0116] Among them, kx i +b indicates the preset straight line equation, y i represents the y-coordinate of the actual trajectory point, and T represents the number of sampling points;

[0117] The speed adjustment rate is expressed as the magnitude of the speed change during the avoidance process:

[0118]

[0119] Among them, v max 、v min Indicates the maximum and minimum speeds during the avoidance phase, vbase Indicates the baseline speed, such as normal walking speed;

[0120] The angle change rate is used to describe the dynamic change of the angle between the actual trajectory direction and the preset straight line direction:

[0121]

[0122] Among them, v x,t 、v y,t represents the velocity component at time t, θ t It represents the angle between the current trajectory direction and the preset straight line, and Δt is the sampling interval;

[0123] right To perform smoothing:

[0124]

[0125] S3.3: Establish a multi-input and multi-output dynamic weighted branch network model. The model input is the motion parameter sequence of each action in three consecutive gait cycles, a total of five inputs, and the output is two branches: the motor function score of the action and the fall risk rating.

[0126] S3.3.1: The multi-input feature processing layer is designed with six independent input branches, corresponding to six action types (standing balance, sit-stand transition, normal walking, emergency stop, turn action, and evasive path deviation action). The input of each branch is a sequence of action parameters, and the sequence length is unified to T = 50 through dynamic time warping (DTW);

[0127] S3.3.2: The shared-private feature encoding layer is designed as a shared feature extractor and an action-specific feature extractor. The shared feature extractor consists of a fully connected layer and is responsible for extracting the general motion features of the action (also called shared features). The action-specific feature extractor is designed for each action independently, introduces an attention mechanism and a bidirectional long short-term memory network, and adds a dilated convolution (the dilation rate is set to 2) to capture long-range dependencies. The action-specific feature extractor is responsible for learning the specific features of each action. Where i represents the i-th specific branch, n represents the number of branches, which is 6;

[0128] S3.3.3: The dynamic weight module receives the shared feature x after the shared-private feature encoding layer. shared and specific characteristics

[0129] Use global average pooling to transform x shared Compress to vector Integrate the spatial information of the entire feature map and retain the overall pattern of the features; use a multi-layer perceptron to map the global features into a 32-dimensional weight vector. There are two fully connected layers. The first layer converts the C-dimensional vector into 16 dimensions, and the second layer converts the 16-dimensional vector into 32 dimensions:

[0130] f transformed =MLP(f global )

[0131] Then calculate the weight α of each action branch feature:

[0132]

[0133] Where ω is the trained 6×32 weight matrix, b is the 6-dimensional bias vector, and τ is the temperature coefficient, which controls the sensitivity of the weight to the action category and is set to 0.55.

[0134] Fusion of shared features with specific features:

[0135]

[0136] S3.3.4: The dual-task output layer is designed to have two output branches: the motion scoring branch and the risk rating branch.

[0137] The motion scoring branch consists of a fully connected layer, which outputs the score of the action and designs the dynamic structure similarity loss function L score and action-specific regularization loss L reg :

[0138] L score =β1·MSE(y pred ,y true )+β2·(1-SSIM(y pred ,y true ))

[0139]

[0140] Among them, β1 and β2 are weight coefficients, y pred and y true Represents the predicted value and true value of the data, MSE(y pred ,y true ) represents the mean square error, SSIM(y pred ,y true ) represents the structural similarity between the predicted value and the true value, γ is the regularization coefficient, is the variance, which indicates the fluctuation of the prediction value of the i-th branch. The calculation formula is as follows:

[0141]

[0142] in, represents the mean of the predicted value and the true value, represents the variance, Indicates covariance, c1 and c2 are constants used to prevent the denominator from being 0, c1=(0.01L y ) 2 ,c2=(0.03L y ) 2 , L y is the data dynamic range.

[0143] The main body of the risk rating branch consists of a three-layer cascade classifier. The Softmax activation function is used to classify risks into five levels from 0 to 4. The dynamic focal cross entropy loss is designed as the loss function of the risk rating branch:

[0144]

[0145] Among them, p true,c and p pred,c is the true probability and predicted probability of risk type c, γ c It is associated with the actual risk level, with higher risks increasing penalties, and M represents the number of risk level categories;

[0146] Combining JS divergence and KL divergence to constrain the semantic consistency of the two branches of motion scoring and risk rating:

[0147] L align =JS(p score ,p risk )+KL(p score ‖ ‖p risk ).

[0148] Step 4: Integrate multi-scenario motion data and dynamically predict the motion parameters of the next monitoring cycle based on the action risk level obtained in step 3; use the difference between the predicted value and the actual value of the motion parameters of the next cycle and the motion function score of each action to obtain the comprehensive motion function score of multiple scenarios.

[0149] Exemplarily, step 4 includes the following steps:

[0150] S4.1: Design a risk sensitivity dynamic prediction model to dynamically adjust the prediction strategy based on the risk level of the current action and output the motion parameters for the next posture detection cycle.

[0151] Specifically, S4.1 includes the following steps:

[0152] S4.1.1: The model is designed as a dual-input architecture. The two inputs are the concatenation of the average motion parameter sequences of different actions in the current cycle and the risk rating vector of different actions obtained in step 3.

[0153] S4.1.2: After the input layer, we perform spatiotemporal feature extraction on the input motion parameter sequence. We first use a one-dimensional convolutional layer to capture local temporal patterns, then use a Transformer encoder to extract global dependencies and output a fixed-length feature vector F. We then decouple the unified feature F into six parts based on the action: F1, F2, L, and F6.

[0154] S4.1.3: Design a risk suppression module, and use a linear or exponential decay function to adjust the weights of different action parameters in subsequent feature fusion based on the risk rating of each action. The higher the risk, the lower the weight of the action. At the same time, according to the probability of risk rating Perform segmented calculations;

[0155] Specifically, for actions with higher risk ratings, the smaller their role in parameter prediction is set, that is, the action parameters for the next monitoring period are predicted as much as possible based on the current healthiest conditions. The greater the difference between the actual value and the predicted value, the lower the comprehensive score.

[0156] The segmented calculation strategy of risk rating weight is as follows: when the risk rating probability of an action is When the risk rating probability is When , the exponential decay is suppressed, and the risk rating weight The calculation formula is as follows:

[0157]

[0158] Among them, p i represents the action risk rating, is the variance coefficient, regulating the influence of variance, σ i represents the standard deviation of historical forecast errors, τ p Represents the temperature coefficient, which is used to balance the weight distribution, is the suppression coefficient, which dominates the suppression strength and is calculated as follows:

[0159]

[0160] Among them, γ0 is the basic coefficient, γ1 is the risk sensitivity increment, and θ is the risk threshold. When γ1 exceeds θ, the suppression strength is strengthened;

[0161] Use weights to calculate fusion features:

[0162]

[0163] S4.1.4: Set up independent parameter prediction networks for the six actions and combine the fusion features F fusedAs input, it passes through 5 layers of fully connected layers and ReLU activation function to mine the unique patterns of actions and output the predicted value.

[0164] S4.2: Calculate the predicted value of motion parameters for the next monitoring period The difference from the true value y is calculated using the mean square error formula:

[0165]

[0166] Where k is the number of motion parameters.

[0167] S4.3: Based on the above, a comprehensive motor function score combining multi-scenario data and historical data is implemented:

[0168]

[0169] Among them, S i represents the motion score of the i-th action, λ total It is an adjustment parameter used to balance the impact of prediction error on the comprehensive score, and is set to 0.3.

[0170] Finally, preferably, step 5 is also included: the computer interactive terminal displays the "scene-score-time" data curve, the historical data curve of the multi-scene comprehensive score and the "motion parameter predicted value and actual value" time data curve.

[0171] Example 2

[0172] This embodiment provides a system for implementing the above-mentioned multi-scenario gait monitoring and motor function assessment method. The system includes the following parts:

[0173] The monitoring and data collection terminal is used to collect real-time video data of the elderly in multiple life scenarios;

[0174] The computer processing end is used to complete the computational evaluation of the system, including establishing the three-dimensional motion trajectory of key points, identity matching and association, calculating motion parameters, calculating the independent score and fall risk rating of each action, and calculating the comprehensive motion function score in multiple scenarios;

[0175] The computer storage terminal is used to store the independent score and risk level of each action in each scene, the comprehensive score data of multiple scenes, and the predicted and actual values ​​of the motion parameters of each action in the local file corresponding to the identity, so that they can be called up at any time later;

[0176] The computer display terminal is used to display the scene-score-time data curve, the historical data curve of the multi-scene comprehensive score, and the time data curve of the motion parameter predicted value and the actual value, so as to intuitively show the character's gait function status and remind relevant personnel to pay attention to functional decline in time.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-scenario gait monitoring and motor function assessment method, characterized in that: The following steps are involved: Step 1: Deploy a multi-camera collaborative monitoring system to collect real-time video data of the elderly in multiple life scenarios. Extract key point information of the elderly's joints and skeleton from the video data and establish the three-dimensional motion trajectory of the key points. Step 2: Match and associate the video data and key point 3D motion trajectory data obtained in step 1 with the individual identity information of the elderly; Step 3: For different scenarios, extract the motion parameters of the corresponding actions and design an independent long-short-term memory network model for each action. The input of this model is the action parameter sequence, and the output is the motion function score and fall risk rating of the action; Step 4: Integrate multi-scenario motion data and dynamically predict the motion parameters of the next monitoring cycle based on the action risk level. The designed method uses the difference between the predicted and actual values ​​of the motion parameters of the next cycle and the motion function score of each action to obtain the comprehensive motion function score of multiple scenarios.

2. The multi-scenario gait monitoring and motor function assessment method according to claim 1, characterized in that: In step 1, a Gaussian mixture model is used to separate moving objects from the captured video. The OpenPose human pose estimation algorithm is applied to extract the joint and skeleton key point information of the elderly. The PoseLifter model is used to predict the position of the key points in three-dimensional space. The key point positions are tracked in time series to obtain the three-dimensional key point coordinates in each frame of the video, and the Kalman filter is used to smooth the trajectory of the key points.

3. The multi-scenario gait monitoring and motor function assessment method according to claim 1, characterized in that: In step 2, a dual verification mechanism of face recognition and human skeleton geometric features is used to accurately match and associate the video and key point three-dimensional motion trajectory data obtained in step 1 with the individual identity information of the elderly.

4. The multi-scenario gait monitoring and motor function assessment method according to claim 3, characterized in that: Step 2 includes the following steps: S2.1: Extract facial features from video data and calculate the cosine similarity S between the facial features extracted from the video data and the facial features stored in the identity database face ; S2.2: Decompose the joints and skeleton key points extracted from the video data into multiple groups of vector chains, calculate the dynamic length ratio and vector angle of each group of chains, and obtain the skeleton features; Calculate the Mahalanobis distance d between the skeleton features in the video data and the sample mean of the skeleton feature sample library of each identity pre-stored in the identity database mah ; Then calculate the Mahalanobis distance d mah and the skeleton feature reference value d of the identity in the identity database base The similarity S skeleton ; Multiple sets of skeleton feature samples are pre-entered into the identity database for each identity, and a skeleton feature sample library for each identity is constructed. Furthermore, for each identity, the Mahalanobis distance between each sample in its skeleton feature sample library and the sample average value of the skeleton feature sample library is calculated. Then, for each identity, the 95% quantile of all the calculated Mahalanobis distances is taken as the normalized benchmark value, thereby obtaining the benchmark value d of the skeleton feature corresponding to each identity. base ; S2.3: The cosine similarity S of facial features face Similarity S with skeleton features skeleton The weights are dynamically set according to the light intensity for fusion to obtain the fused similarity score ID-Score; the identity of the person is determined based on the fused similarity score ID-Score.

5. The multi-scenario gait monitoring and motor function assessment method according to claim 1, characterized in that: Step 3 includes the following steps: S3.1: Identify the following specific actions in each scenario: standing balance and sit-to-stand transition in the bedroom; normal walking, emergency stop, standing balance, and turning in the hallway; sit-to-stand transition, standing balance, and avoidance path deviation in the reading room and living room. S3.2: Calculate the corresponding action parameters for each action; S3.3: Establish a multi-input and multi-output dynamic weighted branch network model. The model input is the motion parameter sequence of each action in three consecutive gait cycles, and the output is two branches: the motor function score of the action and the fall risk rating.

6. The multi-scenario gait monitoring and motor function assessment method according to claim 5, characterized in that: In step S3.2: For standing balance, calculate the standard deviation σ of the pelvic center coordinates p =(σ x ,σ y ,σ z ), calculate the 95% confidence ellipse area based on the covariance matrix of the center of gravity trajectory; For sit-stand transfers, calculate the angle change, hip flexion and extension moments, and the ratio of the vertical displacement of the center of mass of the trunk to the total displacement using inverse dynamics. For normal walking movements, the correlation coefficient of the left and right lower limb joint angle curves and the standard deviation of the step length for 10 consecutive steps were calculated; For emergency stopping actions, calculate braking distance and angular velocity mutation; For turning movements, the rotation radius, the integrated value of the trunk angular velocity, and the phase difference of the shoulder-hip angle change were calculated; For evasive path deviation actions, the path deviation degree, speed adjustment rate, and the rate of change of the angle between the actual trajectory and the preset straight line are calculated.

7. The multi-scenario gait monitoring and motor function assessment method according to claim 5, characterized in that: Step S3.3 includes the following steps: S3.3.1: The multi-input feature processing layer is designed with six independent input branches, corresponding to the six action types: standing balance, sit-stand transition, normal walking, emergency stop, turn action, and avoidance path deviation action. The input of each branch is a sequence of action parameters, and the sequence length is unified through dynamic time warping. S3.3.2: The shared-private feature encoding layer is designed to consist of two parts: a shared feature extractor and an action-specific feature extractor. The shared feature extractor consists of fully connected layers and is responsible for extracting shared features of actions. The action-specific feature extractor is designed for each action with a separate bidirectional long short-term memory network that incorporates an attention mechanism and adds dilated convolution to capture long-range dependencies. The action-specific feature extractor is responsible for learning the specific features of each action. S3.3.3: The dynamic weight module receives shared features and specific features after the shared-private feature encoding layer. It compresses the shared features into a vector using global average pooling, integrates the spatial information of the entire feature map, and preserves the overall pattern of the features. It uses a multi-layer perceptron to map the global features into a 32-dimensional weight vector. It has two fully connected layers: the first layer converts the C-dimensional vector into a 16-dimensional vector, and the second layer converts the 16-dimensional vector into a 32-dimensional vector. It then calculates the weights of the features of each action branch and fuses the shared features with the specific features. S3.3.4: The dual-task output layer is designed as two output branches: motion scoring branch and risk rating branch; the motion scoring branch is composed of a fully connected layer, which outputs the score of the action, and designs a dynamic structural similarity loss function and an action-specific regularization loss; the main body of the risk rating branch is composed of a three-layer cascade classifier, which divides the risk into five levels of 0-4 through the Softmax activation function, and designs a dynamic focus cross entropy loss as the loss function of the risk rating branch, combining JS divergence and KL divergence to constrain the semantic consistency of the motion scoring and risk rating branches.

8. The multi-scenario gait monitoring and motor function assessment method according to claim 1, characterized in that: Step 4 includes the following steps: S4.1: Design a dynamic risk sensitivity prediction model to dynamically adjust the prediction strategy based on the risk level of the current action and output the motion parameters for the next posture detection cycle; S4.2: Calculate the predicted value of motion parameters for the next monitoring period The difference MSE from the true value y; S4.3: Combine multi-scenario data and historical data to perform a comprehensive motor function score: Among them, S i represents the motion score of the i-th action, λ total is the adjustment parameter, represents the risk rating weight of the i-th action.

9. The multi-scenario gait monitoring and motor function assessment method according to claim 8, characterized in that: Step S4.1 includes the following steps: S4.1.1: The model is designed as a dual-input architecture. The two inputs are the concatenation of the average motion parameter sequences of different actions in the current cycle and the risk rating vector of different actions obtained in step 3. S4.1.2: After the input layer, we perform spatiotemporal feature extraction on the input motion parameter sequence. We first use a one-dimensional convolutional layer to capture local temporal patterns, then use a Transformer encoder to extract global dependencies and output a fixed-length feature vector F. We then decouple the unified feature F into six parts based on the action: F1, F2, L, and F6. S4.1.3: Design a risk suppression module to adjust the weights of different action parameters in feature fusion using a linear or exponential decay function for the risk rating of each action, and at the same time, Calculate the weights in segments; use the weights to calculate the fusion features; S4.1.4: Establish independent parameter prediction networks for the six actions, use the fused features as input, and use five fully connected layers and ReLU activation functions to mine the unique patterns of the actions and output predicted values.

10. A multi-scenario gait monitoring and motor function assessment, characterized by: For implementing the method described in claim 1, the system comprises: The monitoring and data collection terminal is used to collect real-time video data of the elderly in multiple life scenarios; The computer processing end is used to complete the computational evaluation of the system, including establishing the three-dimensional motion trajectory of key points, identity matching and association, calculating motion parameters, calculating the independent score and fall risk rating of each action, and calculating the comprehensive motion function score in multiple scenarios; The computer storage terminal is used to store the independent score and risk level of each action in each scene, the comprehensive score data of multiple scenes, and the predicted and actual values ​​of the motion parameters of each action in the local file corresponding to the identity, so that they can be called up at any time later; The computer display terminal is used to display the scene-score-time data curve, the historical data curve of multi-scene comprehensive scores, and the time data curve of motion parameter predicted values ​​and actual values.

Citation Information

Cited By

  • Ankle joint exoskeleton auxiliary strategy generation method and device based on personalized search

    CN121313163A

  • Ankle exoskeleton assistance strategy generation method and device based on personalized search

    CN121313163B

  • Underground coal mine intelligent supervision method based on data fusion

    CN122135302A