Gait recognition method, system and equipment applied to exercise therapy
By employing techniques such as neural frequent differential equations and multi-scale convolution kernels, the problems of incomplete gait data and environmental interference have been solved, achieving high-precision, low-latency gait recognition and meeting the personalized exercise therapy needs of patients with depression.
Patent Information
- Application Number
- CN202510990532.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Existing gait recognition methods are affected by camera angle, lighting changes, and environmental occlusion in real-world scenarios, resulting in incomplete data that is difficult to apply to different patients. Furthermore, traditional methods lack robustness in response to gait changes, affecting recognition accuracy and the effectiveness of personalized movement interventions.
Data completion is performed using neural frequent differential equations, gait features are extracted by combining multi-scale convolution kernels and adaptive attention mechanisms, and recognition accuracy and robustness are improved by using depthwise separable convolution and multiple loss function optimization strategies.
To improve the accuracy and robustness of gait recognition in environments with incomplete data and complex conditions, reduce computational complexity, and enable real-time feedback for personalized motion intervention.
Smart Images

Figure CN120877376A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of biometric recognition, and more specifically, relates to a gait recognition method, system and device for use in exercise therapy. Background Technology
[0002] As societal attention to mental health issues continues to grow, intelligent health monitoring and intervention methods based on biometric recognition technology are attracting increasing attention from academia and industry. Gait recognition, as an important branch of biometrics, shows broad application prospects in fields such as medical rehabilitation and mental health monitoring due to its advantages of being non-contact, individual-specific, and requiring low-level data acquisition equipment. Psychological and biomedical research indicates that the gait characteristics of patients with depression often differ significantly from those of healthy individuals, such as reduced stride length, decreased walking speed, and impaired limb coordination. Accurate analysis of gait data can assist psychologists in assessing the progression of a patient's condition and providing personalized exercise intervention plans.
[0003] In the fields of intelligent rehabilitation and humanoid robots, gait recognition technology is gradually becoming one of the key technologies for improving the intelligence and personalization of exercise therapy for depression. General-purpose humanoid robots can analyze the movement patterns of patients with depression in real time based on gait recognition technology, providing accurate movement feedback and assisting in the development of personalized rehabilitation training programs. For example, during rehabilitation training, the robot can identify the patient's gait characteristics, determine whether their movement status meets the expected goals, and adjust the training intensity through voice, visual, or motion feedback to optimize the rehabilitation effect. Furthermore, in home or medical institution environments, intelligent robots with gait recognition capabilities can monitor patients' movement habits over a long period and provide warnings when abnormal gait patterns are detected, thereby improving the effectiveness and safety of exercise therapy.
[0004] However, existing methods still face some challenges in this recognition task: (1) In real-world scenarios, the collected gait data may be incomplete due to factors such as camera angle, lighting changes, and environmental occlusion, which affects the accuracy of recognition; (2) There are significant differences in the motor abilities and gait characteristics of different patients, making it difficult to apply fixed gait assessment criteria to all individuals; (3) During the rehabilitation process, patients may experience gait changes due to factors such as emotional fluctuations and fatigue, and traditional gait recognition methods are not robust enough in dealing with these dynamic changes. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this application is to provide a gait recognition method, system and device for use in exercise therapy, which aims to solve the problem of low accuracy of existing gait recognition methods.
[0006] To achieve the above objectives, in a first aspect, this application provides a gait recognition method applied to exercise therapy, comprising: Acquire the user's gait frame sequence; the gait frame sequence includes multiple gait frames arranged in time, and each gait frame includes position data of multiple key points set on the user's torso; The temporal dependency of the gait frame sequence is captured by the constant differential equation of the gait, and the missing data in the gait frame sequence is filled in to obtain the filled gait frame sequence. A multi-scale convolutional kernel combined with an adaptive attention mechanism is used to extract gait features at different scales from the completed gait frame sequence. A depthwise separable convolution is also used to extract gait features from the completed gait frame sequence. In the multi-scale convolutional kernel, the large-size convolutional kernel is used to capture global spatiotemporal features, and the small-size convolutional kernel is used to capture local detail features. The adaptive attention mechanism is used to adjust the attention to different time gait frames and different key points in gait frames. The extracted gait features are input into a pre-trained convolutional neural network (CNN) to identify the user's gait; the loss function used in the training process of the convolutional neural network considers the loss of gait recognition accuracy, computational complexity, and computational latency.
[0007] It should be noted that the above gait recognition method is applicable to exercise therapy. By accurately analyzing gait data, it can help provide users with personalized exercise intervention plans and achieve exercise therapy.
[0008] In one possible implementation, noise terms are introduced into the constant differential equation to simulate random disturbances in the user's environment.
[0009] In one possible implementation, the adaptive attention mechanism is used to adjust the attention given to different gait frames at different times and different key points within the gait frames, including: The adaptive attention mechanism is used to maintain high attention to important gait frames; the important gait frames include: time frames carrying important motion information of the user and frames with small changes between adjacent time points.
[0010] In one possible implementation, the time frame carrying important motion information of the user is a gait frame that plays a decisive role in identifying gait features; the frame with minute changes is a gait frame that reflects key adjustments in the gait.
[0011] In one possible implementation, the multi-scale convolution kernel is:
[0012] in, Indicates input data, For convolution kernel, This represents the convolution result. i and j Indicates the input feature points, Indicates the index of the convolution kernel; The adaptive attention mechanism is as follows:
[0013] in, y This represents the characteristics of the output after weighted averaging. Indicates the first The importance weights of each feature It is the first One characteristic, The total number of features;
[0014] in, It is the first Attention weights for each feature point This represents the gait feature function, used to calculate the importance of each feature point. It is the length of the gait sequence.
[0015] In one possible implementation, depthwise separable convolution is used to extract gait features from the completed gait frame sequence, including: Gait features are extracted from gait frame sequences using depthwise convolution and pointwise convolution.
[0016] In one possible implementation, the extracted gait features are further processed before being input into a pre-trained convolutional neural network, including the following steps: An adaptive attention mechanism is used to enhance the extracted gait features; The adaptive attention mechanism dynamically adjusts its weights based on changes in the user's environment to counteract the effects of external noise. The objective function used to dynamically adjust the weights of the adaptive attention mechanism is:
[0017] Where C represents the loss function of the entire optimization problem, and its objective is to minimize it. These are enhanced gait features. It's a real label. It is a regularization parameter. It is attention weight. It is the length of the gait sequence; T represents the total number of feature points.
[0018] In one possible implementation, the loss function used in the training process of the convolutional neural network is... for:
[0019] in, This indicates a loss in gait recognition accuracy. This represents the computational complexity loss. Indicates delay loss, These are the weights of each loss function.
[0020] Specifically, by adjusting the weights of different loss functions, multiple aspects such as accuracy, complexity, and latency can be balanced to obtain a more comprehensive and efficient model.
[0021] Secondly, this application provides a gait recognition system for use in exercise therapy, comprising: The gait frame sequence acquisition module is used to acquire the user's gait frame sequence; the gait frame sequence includes multiple gait frames arranged in time, and each gait frame includes position data of multiple key points set on the user's torso; The gait frame sequence completion module is used to capture the time dependency of the gait frame sequence through a neural ordinary differential equation, and to complete the missing data in the gait frame sequence to obtain the completed gait frame sequence. The gait feature extraction module is used to extract gait features at different scales from the completed gait frame sequence by using multi-scale convolutional kernels combined with an adaptive attention mechanism, and to extract gait features from the completed gait frame sequence by using depthwise separable convolution. Among the multi-scale convolutional kernels, large-size convolutional kernels are used to capture global spatiotemporal features, while small-size convolutional kernels are used to capture local detail features. The adaptive attention mechanism is used to adjust the attention to different time gait frames and different key points in gait frames. The gait recognition module is used to input the extracted gait features into a pre-trained convolutional neural network to recognize the user's gait; the loss function used in the training process of the convolutional neural network considers the loss of gait recognition accuracy, the loss of computational complexity, and the loss of computational delay.
[0022] Thirdly, this application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation thereof.
[0023] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to perform the method described in the first aspect or any possible implementation thereof.
[0024] Fifthly, this application provides a computer program product that, when run on a processor, causes the processor to perform the method described in the first aspect or any possible implementation thereof.
[0025] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art: This application provides a gait recognition method, system, and device for exercise therapy, proposing a gait joint trajectory prediction method based on stochastic neural differential equations. By introducing stochastic neural differential equations, it is possible to predict missing frames and supplement the original gait trajectories when data is incomplete, utilizing the dynamic characteristics of time series. This technique not only improves the completeness of gait features but also enables the model to maintain high recognition accuracy even with insufficient frames, significantly enhancing the robustness of the gait recognition system.
[0026] This application provides a gait recognition method, system, and device for exercise therapy. Through multi-scale convolutional kernels and an adaptive attention mechanism, the model's ability to extract gait temporal features is enhanced. Specifically, the module dynamically adjusts the level of attention given to gait features by adaptively weighting keyframes and subtle changes in the gait sequence, thereby capturing detailed information within the gait sequence. This method not only improves the ability to recognize global gait patterns but also enhances the model's resistance to noise and interference, enabling the system to maintain high recognition accuracy even in complex environments.
[0027] This application provides a gait recognition method, system, and device for exercise therapy. The proposed gait recognition method fully considers real-time performance and computational efficiency, significantly reducing computational complexity through a multiple loss function optimization strategy and depthwise separable convolution technology. While ensuring high recognition accuracy, it reduces hardware resource requirements, achieving low computational cost and low latency. Furthermore, the use of layer normalization and adaptive attention mechanisms effectively improves the system's robustness, enabling stable operation of gait recognition even in complex environments, ensuring real-time monitoring and providing accurate motion feedback. Attached Figure Description
[0028] Figure 1 This is one of the flowcharts illustrating the gait recognition method for exercise therapy provided in the embodiments of this application; Figure 2 This is a second schematic flowchart of the gait recognition method for exercise therapy provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of a gait recognition system for exercise therapy provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0030] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.
[0031] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0032] The embodiments of this application are described below with reference to the accompanying drawings.
[0033] To address the aforementioned issues, this application proposes a gait joint trajectory prediction and recognition method based on Neural Ordinary Differential Equations (SODEs) and Taylor series extensions, and applies it to a general humanoid robot-assisted exercise therapy for depression. This method first uses SODEs to model and predict the motion trajectory of the patient's gait joints, effectively solving the problem of incomplete data collection. Specifically, by introducing a noise-based stochastic SODE model, this method can capture and model the uncertainty in gait data, thereby improving recognition accuracy under conditions of camera angle, lighting changes, or environmental occlusion. By introducing keypoint extraction error penalties and insufficient frame rate penalties into the model's loss function, this method can effectively reduce errors caused by environmental factors and improve robustness in complex environments.
[0034] Figure 1 This is one of the flowcharts illustrating a gait recognition method for exercise therapy provided in this application embodiment; as follows: Figure 1 As shown, it includes the following steps: Step S101: Obtain the user's gait frame sequence; the gait frame sequence includes multiple gait frames arranged in time, and each gait frame includes position data of multiple key points set on the user's torso; Step S102: Capture the time dependency of the gait frame sequence using a normal differential equation to complete the missing data in the gait frame sequence and obtain the completed gait frame sequence. Step S103: A multi-scale convolutional kernel combined with an adaptive attention mechanism is used to extract gait features at different scales in the completed gait frame sequence, and a depthwise separable convolution is used to extract gait features in the completed gait frame sequence; wherein, in the multi-scale convolutional kernel, the large-size convolutional kernel is used to capture global spatiotemporal features, and the small-size convolutional kernel is used to capture local detail features, and the adaptive attention mechanism is used to adjust the attention to different time gait frames and different key points in gait frames; It is understandable that the multi-scale convolutional kernels and adaptive attention mechanism mentioned above can be understood as a convolutional augmentation module.
[0035] Step S104: The extracted gait features are input into a pre-trained convolutional neural network to identify the user's gait; the loss function used in the training process of the convolutional neural network considers the loss of gait recognition accuracy, the loss of computational complexity, and the loss of computational delay.
[0036] It should be noted that gait recognition in exercise therapy for depression is often limited by environmental factors such as camera acquisition angle, lighting changes, and occlusion, resulting in insufficient gait data frames and affecting the completeness and accuracy of gait features. Due to the lack of gait data, traditional gait recognition methods often face performance degradation. To solve this problem, the key technology lies in effectively predicting and supplementing the limited gait frame sequences to obtain a more complete sequence of gait motion features.
[0037] This application proposes a gait joint trajectory prediction method based on stochastic neural network differential equations and Taylor series expansions. By introducing stochastic neural network differential equations, we can utilize the dynamic characteristics of time series data to predict missing frames and supplement the original gait trajectories even with incomplete data. This technique not only improves the completeness of gait features but also enables the model to maintain high recognition accuracy even with insufficient frames, significantly enhancing the robustness of the gait recognition system.
[0038] Furthermore, gait characteristics of patients with depression exhibit significant individual differences, and may display different movement patterns under different situations (such as mood swings, fatigue, or changes in the external environment). In addition, changes in clothing (such as heavy clothing) and walking styles (such as dragging steps or walking with the head down) may introduce noise interference, affecting the accuracy of gait recognition. Therefore, designing a feature extraction module that can capture both global gait movement patterns and local gait details is one of the core issues in optimizing gait recognition models.
[0039] To address this challenge, this application designs a convolutional amplification module based on an attention mechanism. Through multi-scale convolutional kernels and an adaptive attention mechanism, it enhances the model's ability to extract gait temporal features. Specifically, the module dynamically adjusts the level of attention given to gait features by adaptively weighting keyframes and subtle changes in the gait sequence, thereby capturing detailed information within the gait sequence. This method not only improves the ability to recognize global gait patterns but also enhances the model's robustness against noise and interference, enabling the system to maintain high recognition accuracy even in complex environments.
[0040] Furthermore, in practical applications of exercise therapy for depression, general-purpose humanoid robots need to monitor patients' gait changes in real time and provide timely motion feedback. Therefore, gait recognition systems must not only have high recognition accuracy but also meet the requirements of low computational cost, low latency, and high robustness. In addition, in complex environments such as homes, hospitals, or rehabilitation centers, patients' gait may be affected by factors such as spatial layout, obstacles, and dynamic crowds. Therefore, designing a gait recognition method that combines computational efficiency and stability to ensure the real-time monitoring capability of humanoid robots in different environments is an important direction for optimizing gait feedback systems.
[0041] The gait recognition method proposed in this application fully considers real-time performance and computational efficiency. Through a multiple loss function optimization strategy and depthwise separable convolution technology, it significantly reduces computational complexity. While ensuring high recognition accuracy, it reduces hardware resource requirements, achieving low computational cost and low latency. Furthermore, by employing techniques such as layer normalization and adaptive attention mechanisms, the robustness of the system is effectively improved, enabling stable operation of gait recognition even in complex environments. This ensures that general-purpose humanoid robots can perform real-time monitoring and provide accurate motion feedback.
[0042] In summary, this application proposes a gait key point trajectory prediction and recognition method based on neural frequent differential equations, and applies it to general humanoid robot-assisted exercise therapy for depression. This method constructs a prediction-recognition joint optimization framework, fully utilizes neural frequent differential equations to model and predict gait key point trajectories, and combines attention-enhanced temporal convolutional networks for feature extraction. Furthermore, this application specifically optimizes the gait feedback mechanism to address the gait characteristics and exercise rehabilitation needs of patients with depression, ensuring that the robot can perceive and adjust exercise intervention strategies in real time. This application mainly addresses the following three key issues: ① Insufficient gait data frames: During the rehabilitation process, due to factors such as slow patient movement, limited camera angles, or obstruction, gait data is often incomplete, affecting the accuracy of gait assessment.
[0043] ② The problem of dynamic changes in gait characteristics: Depressed patients may exhibit abnormal gait such as slow walking, unstable stride, and forward leaning. These characteristics vary from person to person, and traditional gait recognition methods are difficult to model accurately.
[0044] ③ Computational efficiency and real-time feedback issues: General-purpose humanoid robots need to complete gait monitoring with low latency and provide adaptive motion feedback. Therefore, the model needs to have strong computational efficiency while ensuring high accuracy.
[0045] To address the aforementioned issues, this application first employs a trajectory prediction model based on neural constant differential equations, using existing gait keypoint trajectories to reasonably interpolate missing frames, and then combines time-series Transformer modeling to complete the gait sequence. Furthermore, we integrate a gait reconstruction method that combines motion constraints and biomechanical characteristics to ensure the authenticity of the completed data and improve the accuracy of gait recognition when data frames are insufficient. A detailed flowchart is shown below. Figure 2 As shown.
[0046] Secondly, this application constructs a gait feature extraction network based on a multi-scale spatiotemporal attention mechanism, combining the modeling of local gait micro-motion features (such as foot swing and trunk balance) and global gait patterns (such as walking rhythm and stride changes) to enhance the model's ability to express gait dynamics. Simultaneously, an adaptive weighting mechanism is introduced to automatically adjust the contribution weights of different gait features to the final recognition result, thereby improving the model's generalization ability and robustness under different walking states.
[0047] Finally, this application employs the lightweight deep learning model EfficientNet combined with knowledge distillation model compression technology to reduce computational costs while maintaining recognition accuracy. Furthermore, we introduce an Adaptive Gait Alignment (AGA) mechanism, enabling the robot to automatically adjust recognition parameters based on the gait rhythm of different patients, ensuring stability in dynamic environments. By combining edge computing and cloud-based collaborative optimization, we further enhance the humanoid robot's real-time processing capabilities for gait data, meeting the personalized needs of rehabilitation training for patients with depression.
[0048] like Figure 2 As shown, the above gait recognition method includes the following steps: Step S1: Gait sequence completion and recognition method based on Neural ODEs.
[0049] This application proposes an innovative gait recognition method specifically designed for exercise therapy for patients with depression. This method combines cutting-edge technologies such as Neural ODEs (Neural ODEs), Stochastic ODEs (SODEs), multi-scale convolutional kernels, attention mechanisms, and depthwise separable convolution techniques. It addresses the challenges of gait recognition by constructing a multi-layered prediction and recognition framework, especially in scenarios with incomplete gait data, dynamically changing features, and limited computational resources. The gait recognition method in this application not only significantly improves recognition accuracy and robustness but also optimizes real-time performance and computational efficiency, enabling stable operation even in complex environments and meeting the personalized needs of rehabilitation treatment for patients with depression.
[0050] Step S1.1: Handling missing gait data and trajectory prediction In practical applications of exercise therapy for depression, gait data acquisition is often affected by various factors, such as camera angle, occlusion, and changes in lighting. These factors frequently result in insufficient gait data frames, thus affecting the completeness of gait features and the accuracy of recognition. Traditional gait recognition methods perform poorly when faced with missing data. This application effectively solves this problem by employing Neural ODEs, supplementing missing gait frames through dynamic modeling of gait key point trajectories.
[0051] Neuron ordinary differential equations (ODEs) are a deep learning framework capable of handling dynamic sequential data. They model the changes in time series by solving ordinary differential equations (ODEs). In gait data, the trajectories of gait keypoints can be modeled as a time series. Neuron ODEs can capture the temporal dependencies of gait data, thereby predicting and filling in missing frames. Its core formula is:
[0052] in, Indicates the key points of gait in time state, It is a neural network used to describe the dynamic evolution of gait key points. For model parameters, This represents the initial state, i.e., the state of the gait keypoints in the Neural ODE at time t=0. In this way, future gait keypoints can be predicted step-by-step based on existing data, thus completing missing data frames. For completing missing gait data, we use the following discretization formula for prediction:
[0053] in, This represents the time step. Through this discretization process, we can progressively predict the gait trajectory and fill in the missing key points.
[0054] Step S1.2: Stochastic God Ordinary Differential Equation (SODE) To improve robustness under incomplete data conditions, this application further introduces stochastic neural ordinary differential equations (SODEs). SODEs, by introducing noise terms into the neural ordinary differential equations, can model the randomness of the environment and improve the model's adaptability. This method enables the model to maintain high recognition accuracy even when faced with noisy or missing data. The formula for SODEs is as follows:
[0055] in, This represents the additive noise term, which the model introduces to simulate random disturbances in the environment. This indicates that the noise term follows a Gaussian distribution with zero mean and unit covariance. In this way, we can improve the model's robustness to incomplete data and external environmental noise, thereby ensuring the accuracy of gait recognition.
[0056] Step S1.3: Multi-scale convolutional kernels and attention mechanisms To improve gait recognition accuracy, this application employs multi-scale convolutional kernels and an attention mechanism. Multi-scale convolutional kernels can extract features from gait data at different scales, thus helping the model capture subtle changes in gait. The attention mechanism dynamically adjusts the model's focus based on the importance of different features, improving the ability to recognize key gait features. The formula for the multi-scale convolutional kernel is as follows:
[0057] in, Indicates input data, For convolution kernel, This is the result of convolution. Convolutional kernels of different scales help the model better capture the changes in gait data at different scales. For the attention mechanism, the model calculates the importance weights for each feature. The weighted average yields the final output:
[0058] in, Indicates the first The importance weights of each feature It is the first One characteristic, This represents the total number of features. The attention mechanism enables the model to adaptively focus on key features in gait data, thereby improving recognition accuracy.
[0059] Step S1.4: Depthwise separable convolution technique To optimize computational efficiency, this application employs depthwise separable convolution. Depthwise separable convolution decomposes traditional convolution into two steps: depthwise convolution and pointwise convolution. This separable convolution method significantly reduces computational cost and improves computational efficiency. The formula for depthwise convolution is:
[0060] in, It is the output of depthwise convolution. It is the first input One channel, It is the corresponding convolution kernel. M This represents the total number of channels. Depthwise convolution convolves each input channel independently, while pointwise convolution combines the results in the final step. The formula for pointwise convolution is:
[0061] in, y It is the output of pointwise convolution. For depthwise convolution m Output of each channel, These are the weights for pointwise convolution.
[0062] By introducing techniques such as Neural ODEs, Stochastic ODEs (SODEs), multi-scale convolutional kernels, attention mechanisms, and depthwise separable convolution, all extracted gait features are preprocessed in step S2 and then input into a convolutional neural network for gait recognition. This application provides an innovative gait recognition method. This method not only supplements gait data gaps caused by factors such as camera angle, occlusion, and lighting changes, but also improves recognition accuracy and robustness under incomplete data conditions. Through this method, stable gait recognition can be achieved in complex environments, meeting the personalized needs of exercise therapy for patients with depression.
[0063] Step S2: Design of Convolutional Amplification Module Based on Attention Mechanism This application proposes an innovative attention-based convolutional amplification module specifically designed to address individual differences and environmental interference in gait recognition of patients with depression. Gait characteristics of patients with depression may exhibit different movement patterns under different contexts (such as mood swings, fatigue, or changes in the external environment). Furthermore, changes in clothing (such as heavy clothing) and walking styles (such as dragging steps or walking with the head down) can introduce noise, leading to a decrease in the accuracy of gait recognition. Therefore, designing a feature extraction module that can capture both global gait movement patterns and local gait details is crucial for optimizing the gait recognition model.
[0064] To address this challenge, this application designs a feature extraction module that combines multi-scale convolutional kernels and an adaptive attention mechanism. This module enhances the model's ability to extract temporal features of gait, capturing keyframes and subtle changes in the gait sequence, thereby improving gait recognition accuracy and enhancing its robustness against noise and interference.
[0065] Step S2.1: Application of multi-scale convolution kernels Multi-scale convolutional kernels are used in this application to extract features at different scales from gait sequences. By using convolutional kernels of different sizes, the model can extract detailed information from gait data at multiple scales, capturing both global and local patterns of gait motion. The formula is expressed as:
[0066] in, The convolution operation is in the 1st... Output at each scale It is the input gait sequence data. It is the first Convolutional kernels of various scales, and These represent the height and width of the convolution kernel, respectively. In this way, the convolution kernel can capture multi-level features of gait data at different scales.
[0067] Step S2.2: Adaptive Attention Mechanism The adaptive attention mechanism is one of the key technologies in this application, aiming to dynamically adjust the model's focus based on important features in gait data. Through adaptive weighting, the model can give higher attention to keyframes and subtle changes in the gait sequence at each time step, thereby better recognizing gait patterns. This method enhances the model's robustness to noise and interference. The mathematical expression of the attention mechanism is:
[0068] in, It is the first Attention weight at each moment, It is a gait feature function used to calculate the importance of features at each time step. This refers to the length of the gait sequence. In this way, the model can adaptively weight gait features, thereby dynamically adjusting the level of attention given to keyframes.
[0069] It should be noted that steps S2.1 and S2.2 above correspond to step S1.3.
[0070] Step S2.3: Feature Enhancement and Improvement of Anti-interference Capability By combining multi-scale convolutional kernels and adaptive attention mechanisms, this module effectively enhances gait feature extraction capabilities, especially when facing noise interference in complex environments. The model extracts global gait patterns while also considering local gait details, improving the ability to recognize different gait patterns. The enhancement of gait features can be expressed by the following formula:
[0071] in, This represents the enhanced gait features. The weights are calculated using an adaptive attention mechanism. It is the 1st in the gait sequence The convolutional output at each time step. By using a weighted summation method, the model can improve its ability to capture details while maintaining global gait recognition accuracy.
[0072] Step S2.4: Overall Optimization and Noise Robustness In the final optimization process, this application designs a noise robustness enhancement strategy based on adaptive learning. By dynamically adjusting the attention weights, the model can effectively counteract the influence of external noise according to changes in the environment, enabling gait recognition to maintain high accuracy in complex environments. Noise suppression can be optimized using the following objective function:
[0073] in, Indicates the optimization objective. These are enhanced gait features. It's a real label. It is a regularization parameter. This refers to attention weights. The objective function enhances robustness to noise by minimizing the difference between features and labels.
[0074] The attention-based convolutional amplification module in this application, by combining multi-scale convolutional kernels and an adaptive attention mechanism, can effectively improve the accuracy and robustness of gait recognition in patients with depression. This method can capture global motion patterns and local details in gait sequences, while improving the model's resistance to noise and interference, enabling the gait recognition system to maintain high recognition accuracy even in complex environments.
[0075] Step S3: A Gait Recognition Method with Low Computational Cost and High Robustness In practical applications of exercise therapy for depression, general-purpose humanoid robots need to monitor patients' gait changes in real time and provide timely motion feedback. To achieve this goal, gait recognition systems must not only have high recognition accuracy but also meet the requirements of low computational cost, low latency, and high robustness. However, in complex environments such as homes, hospitals, or rehabilitation centers, patients' gait may be affected by factors such as spatial layout, obstacles, and dynamic crowds. Therefore, designing a gait recognition method that combines computational efficiency and stability to ensure the real-time monitoring capability of humanoid robots in different environments is the core issue in optimizing gait feedback systems.
[0076] This application proposes a gait recognition method that significantly reduces computational complexity and optimizes recognition accuracy and system robustness by combining a multiple loss function optimization strategy, depthwise separable convolution, layer normalization, and adaptive attention mechanism. While maintaining high recognition accuracy, it reduces hardware resource requirements, achieving low computational cost and low latency, ensuring that general-purpose humanoid robots can operate stably in complex environments and provide accurate motion feedback.
[0077] Step S3.1: Optimization Strategy for Multiple Loss Functions To balance computational efficiency and recognition accuracy, this application employs a multiple loss function optimization strategy. This strategy introduces multiple loss functions to comprehensively consider both recognition accuracy and computational efficiency, thereby optimizing the overall performance of the model. Specifically, the loss functions include gait recognition accuracy loss, computational resource loss, and latency loss. The formula for the comprehensive loss function is:
[0078] in, This indicates a loss of recognition accuracy. This represents the computational complexity loss. Indicates delay loss, These are the weights of each loss function. By optimizing this combined loss function, the model can find a balance between accuracy, computational complexity, and latency.
[0079] Step S3.2: Depthwise separable convolution technique To reduce computational complexity, this application employs depthwise separable convolution. Depthwise separable convolution divides the traditional convolution operation into two steps: depthwise convolution and pointwise convolution, thus significantly reducing computational load. In gait recognition, depthwise separable convolution helps reduce computational resource requirements while maintaining high recognition accuracy. Its formula is as follows: (1) Depthwise convolution:
[0080] in, The output of depthwise convolution. For the input of the first m One channel, is the corresponding convolution kernel.
[0081] (2) Pointwise convolution:
[0082] in, For depthwise convolution m Output of each channel, These are the weights for pointwise convolution. Depthwise separable convolution significantly reduces computational complexity and reliance on hardware resources, enabling gait recognition to operate efficiently in environments with low computational resources.
[0083] Step S3.3: Layer Normalization To further improve the robustness and stability of the system, this application employs layer normalization technology. Layer normalization reduces the variability between different time-state features by normalizing the output of each layer, thereby enhancing the model's adaptability to disturbances in complex environments. Its mathematical formula is as follows:
[0084] in, It is the first input One characteristic, and These are the mean and standard deviation of the feature, respectively. and These are learnable scaling and offset parameters. These are the normalized features. In this way, layer normalization can reduce fluctuations during model training and improve the stability and robustness of gait recognition.
[0085] Step S3.4: Adaptive Attention Mechanism To enhance the ability to focus on gait features, this application introduces an adaptive attention mechanism. By calculating the importance weights of gait features, the model can dynamically adjust the degree of attention given to different gait features, thereby improving recognition accuracy. The mathematical expression of the adaptive attention mechanism is:
[0086] in, Indicates the first Attention weights for each gait feature, It is a function of gait features, calculating the importance of each feature. This refers to the length of the gait sequence. In this way, the model can adaptively adjust its focus on key gait features, enhancing the system's robustness and recognition accuracy.
[0087] Step S3.5: System Optimization and Real-time Monitoring Capabilities The gait recognition method presented in this application, while maintaining high accuracy, significantly improves the computational efficiency and stability of the system by combining a multiple loss function optimization strategy, depthwise separable convolution technology, layer normalization, and adaptive attention mechanism. By reducing computational complexity and latency, this application ensures that a general-purpose humanoid robot can monitor patients' gait changes in real time in complex environments and provide accurate motion feedback. The optimized system maintains high robustness in different environments (such as home, hospital, or rehabilitation center), enabling the gait recognition system to operate stably.
[0088] By combining multiple loss function optimization strategies, depthwise separable convolution techniques, layer normalization, and adaptive attention mechanisms, this application provides a low-computational-cost, highly robust gait recognition method. This method not only ensures high-precision gait recognition but also effectively improves real-time monitoring capabilities, operates stably in complex environments, and ensures that general-purpose humanoid robots can provide timely and accurate motion feedback.
[0089] This application proposes a gait recognition method based on neural ordinary differential equations and convolutional neural networks. It significantly improves the accuracy and robustness of gait recognition by combining dynamic modeling and feature extraction. To verify the effectiveness of this application, we adopted multiple evaluation metrics, such as gait recognition accuracy, real-time response capability, and noise resistance in complex environments. Through comparative experiments with multiple benchmark models, we demonstrate the advantages of this application in various scenarios.
[0090] In the experiments, the gait recognition model used in this application is referred to as the GaitODE model. To verify the effectiveness of this model, existing deep learning-based gait recognition models, such as LSTM-GaitNet and Pose-CNN, were selected as benchmark models for comparison. Comparison with these benchmark models demonstrates the significant improvement in recognition accuracy and robustness of the proposed GaitODE model.
[0091] (1) Comparison of benchmark models In skeleton-based gait recognition methods, Pose-CNN and the model proposed in this application achieve similar recognition accuracies in normal walking (NW) mode, demonstrating the importance of gait feature extraction from skeleton key points for gait recognition. However, the Pose-CNN model only performs temporal processing on skeleton features and fails to effectively address missing key points and noise interference in gait sequences. Therefore, its accuracy is significantly lower than that of the GaitODE model proposed in this application in complex environments (such as gait recognition while wearing a raincoat and fast walking scenarios).
[0092] For example, in scenarios involving raincoat walking (RW) and fast walking (FW), Pose-CNN experiences a significant drop in accuracy when processing noisy and occluded data. In contrast, the GaitODE model effectively addresses these challenges by introducing neural frequent differential equations to predict and model the trajectory of key skeleton points. Specifically, the GaitODE model can significantly improve gait recognition accuracy when there are insufficient frames by predicting missing frames.
[0093] (2) Gait sequence completion based on the regular differential equation of the nervous system Unlike the LSTM-GaitNet model, the GaitODE model in this application uses neural ordinary differential equations to predict the trajectory of missing frames, thus maintaining a high recognition accuracy even with partially missing data. Although LSTM-GaitNet uses an LSTM network for temporal modeling, it fails to adequately address the issue of missing keypoints in gait data, resulting in poor prediction performance in complex scenes.
[0094] This application improves the model's ability to recognize subtle motion changes by introducing a convolutional amplification module that not only focuses on the overall motion pattern but also weights local details when extracting gait temporal features. Through the convolutional amplification module, GaitODE can dynamically adjust feature weights, making the model more robust in complex environments.
[0095] (3) Practical application verification To verify the effectiveness of this application in practical applications, a simulated sports training environment was constructed, comprising three mobile robots equipped with high-definition cameras and fifteen test personnel. The test scenarios included standard, challenging, and extreme scenarios, simulating common environmental changes in sports training.
[0096] ① Standard scenario (unobstructed, good lighting): In this scenario, the recognition accuracy of the GaitODE model is very close to that of the benchmark model, at 95.6% (GaitODE) and 94.3% (Pose-CNN), respectively.
[0097] ② Challenging scenarios (partial occlusion, changing lighting): In this environment, the GaitODE model of this application has an accuracy of 89.5%, which is significantly higher than the baseline model (Pose-CNN) of 80.4%.
[0098] ③ Extreme scenarios (severe occlusion, insufficient lighting, frame loss): In this scenario, the GaitODE model still maintains a high recognition accuracy of 85.2%, while the accuracy of the benchmark model Pose-CNN is 74.1%, demonstrating the superiority of this application in complex environments.
[0099] Furthermore, in the multi-robot collaborative mode, the overall recognition accuracy of the system in this application is improved by 7.5%, which proves the superiority of the GaitODE model when multiple robots work together and meets the requirements of real-time gait recognition.
[0100] The system was comprehensively evaluated from multiple camera viewpoints. The viewpoints involved in the experiment (0° to 180°, ten in total) refer to the camera's shooting angle relative to the front of the pedestrian being tested. 0° indicates the camera is directly in front of the pedestrian, 90° is a side view, 180° indicates a shot from behind, and the others, such as 18°, 36°, 54°, 72°, 126°, 144°, and 162°, are intermediate angles. These angle settings comprehensively evaluate the system's robustness and generalization ability in cross-view gait recognition tasks.
[0101] The gait recognition accuracy is shown in Table 1 below. Gallery NM#1-4 refers to the standard gait library used for training or comparison, consisting of normal walking (numbered 1-4), serving as a "library" or template reference during the recognition process. Probe indicates the source of the test samples. NM#5-6 represents normal walking test samples, numbered 5 and 6, where the person is walking naturally during the test. BG#1-2 represents gait test samples when the person is carrying an object (Bag), numbered 1 and 2. CL#1-2 represents gait test samples after the person has changed clothes, numbered 1 and 2. mean represents the average recognition accuracy for all angles on this test set (NM#5-6, BG#1-2, CL#1-2), used as a metric for overall performance. Table 1 shows that GaitODE outperforms existing methods in recognition accuracy at most angles. It also performs well on three typical test sets: normal gait (NM), bag-wearing gait (BG), and clothing-changing gait (CL), demonstrating the practical value of the proposed method in complex and variable training scenarios.
[0102] Table 1
[0103] (4) Summary of technical advantages By comparing with benchmark models such as LSTM-GaitNet and Pose-CNN, the GaitODE model in this application demonstrates significant advantages in multiple complex environments. Specifically: ① Improved recognition accuracy when there are insufficient frames: The missing frames are predicted by using the constant differential equation of the neural network, which avoids the decrease in accuracy caused by missing data in traditional methods.
[0104] ② Robustness to noise interference and environmental changes: Through the convolutional amplification module and adaptive feature weighting, GaitODE can maintain high accuracy under the influence of noise and occlusion.
[0105] ③ Real-time performance and low latency: In environments with high real-time requirements, the model in this application can meet the real-time gait recognition requirements with a processing latency of less than 45ms.
[0106] In summary, the gait recognition method based on neural ordinary differential equations proposed in this application has demonstrated its superiority in accuracy, robustness, and real-time performance in multiple experimental scenarios. It can effectively address key issues in gait recognition, provide important technical support for the application of exercise therapy for depression, and promote the widespread application of gait recognition technology in other complex environments.
[0107] Figure 3This is a schematic diagram of the structure of a gait recognition system for exercise therapy provided in an embodiment of this application; The gait frame sequence acquisition module 310 is used to acquire the user's gait frame sequence; the gait frame sequence includes multiple gait frames arranged in time, and each gait frame includes position data of multiple key points set on the user's torso; The gait frame sequence completion module 320 is used to capture the time dependency of the gait frame sequence through a neural ordinary differential equation, and to complete the missing data in the gait frame sequence to obtain the completed gait frame sequence. The gait feature extraction module 330 is used to extract gait features at different scales in the completed gait frame sequence by using multi-scale convolutional kernels combined with an adaptive attention mechanism, and to extract gait features in the completed gait frame sequence by using depthwise separable convolution; wherein, in the multi-scale convolutional kernels, large-size convolutional kernels are used to capture global spatiotemporal features, small-size convolutional kernels are used to capture local detail features, and the adaptive attention mechanism is used to adjust the attention to different time gait frames and different key points in gait frames; The gait recognition module 340 is used to input the extracted gait features into a pre-trained convolutional neural network to recognize the user's gait; the loss function used in the training process of the convolutional neural network considers the loss of gait recognition accuracy, the loss of computational complexity, and the loss of computational delay.
[0108] It should be understood that the above system is used to execute the methods in the above embodiments. The corresponding program modules in the system are similar in implementation principle and technical effect to those described in the above methods. The working process of the system can be referred to the corresponding process in the above methods, and will not be repeated here.
[0109] Based on the methods in the above embodiments, this application provides an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call logical instructions in the memory 430 to execute the methods in the above embodiments.
[0110] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0111] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0112] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0113] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.
[0114] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0115] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0116] It is understood that the various numerical designations used in the embodiments of this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application.
[0117] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A gait recognition method applied to exercise therapy, characterized in that, include: Acquire the user's gait frame sequence; the gait frame sequence includes multiple gait frames arranged in time, and each gait frame includes position data of multiple key points set on the user's torso; The temporal dependency of the gait frame sequence is captured by the constant differential equation of the gait, and the missing data in the gait frame sequence is filled in to obtain the filled gait frame sequence. A multi-scale convolutional kernel combined with an adaptive attention mechanism is used to extract gait features at different scales from the completed gait frame sequence. A depthwise separable convolution is also used to extract gait features from the completed gait frame sequence. In the multi-scale convolutional kernel, the large-size convolutional kernel is used to capture global spatiotemporal features, and the small-size convolutional kernel is used to capture local detail features. The adaptive attention mechanism is used to adjust the attention to different time gait frames and different key points in gait frames. The extracted gait features are input into a pre-trained convolutional neural network to identify the user's gait; the loss function used in the training process of the convolutional neural network considers the loss of gait recognition accuracy, the loss of computational complexity, and the loss of computational delay.
2. The method according to claim 1, characterized in that, The aforementioned constant differential equation introduces a noise term to simulate random disturbances in the user's environment.
3. The method according to claim 1, characterized in that, The adaptive attention mechanism is used to adjust the attention given to different gait frames at different times and different key points within the gait frames, including: The adaptive attention mechanism is used to maintain high attention to important gait frames; the important gait frames include: time frames carrying important motion information of the user and frames with small changes between adjacent time points.
4. The method according to claim 3, characterized in that, The time frame carrying important motion information of the user is the gait frame that plays a decisive role in recognizing gait features; the frame with minute changes is the gait frame that reflects key adjustments in gait.
5. The method according to any one of claims 1 to 4, characterized in that, The multi-scale convolution kernel is: in, Indicates input data, For convolution kernel, This represents the convolution result. i and j Indicates the input feature points, Indicates the index of the convolution kernel; The adaptive attention mechanism is as follows: in, y This represents the characteristics of the output after weighted averaging. Indicates the first The importance weights of each feature It is the first One characteristic, The total number of features; in, It is the first Attention weights for each feature point This represents the gait feature function, used to calculate the importance of each feature point. It is the length of the gait sequence.
6. The method according to claim 1, characterized in that, Gait features of the completed gait frame sequence are extracted using depthwise separable convolution, including: Gait features are extracted from gait frame sequences using depthwise convolution and pointwise convolution.
7. The method according to any one of claims 1 to 4, characterized in that, Before the extracted gait features are input into a pre-trained convolutional neural network, the following steps are also included: An adaptive attention mechanism is used to enhance the extracted gait features; The adaptive attention mechanism dynamically adjusts its weights based on changes in the user's environment to counteract the effects of external noise. The objective function used to dynamically adjust the weights of the adaptive attention mechanism is: Where C represents the loss function of the entire optimization problem. These are enhanced gait features. It's a real label. It is a regularization parameter. It is attention weight. It is the length of the gait sequence; T represents the total number of feature points.
8. The method according to any one of claims 1 to 4, characterized in that, The loss function used in the training process of the convolutional neural network for: in, This indicates a loss in gait recognition accuracy. This represents the computational complexity loss. Indicates delay loss, These are the weights of each loss function.
9. A gait recognition system for use in exercise therapy, characterized in that, include: The gait frame sequence acquisition module is used to acquire the user's gait frame sequence; the gait frame sequence includes multiple gait frames arranged in time, and each gait frame includes position data of multiple key points set on the user's torso; The gait frame sequence completion module is used to capture the time dependency of the gait frame sequence through a neural ordinary differential equation, and to complete the missing data in the gait frame sequence to obtain the completed gait frame sequence. The gait feature extraction module is used to extract gait features at different scales from the completed gait frame sequence by using multi-scale convolutional kernels combined with an adaptive attention mechanism, and to extract gait features from the completed gait frame sequence by using depthwise separable convolution. Among the multi-scale convolutional kernels, large-size convolutional kernels are used to capture global spatiotemporal features, while small-size convolutional kernels are used to capture local detail features. The adaptive attention mechanism is used to adjust the attention to different time gait frames and different key points in gait frames. The gait recognition module is used to input the extracted gait features into a pre-trained convolutional neural network to recognize the user's gait; the loss function used in the training process of the convolutional neural network considers the loss of gait recognition accuracy, the loss of computational complexity, and the loss of computational delay.
10. An electronic device, characterized in that, include: At least one memory for storing computer programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-8.