A parkinson hand rehabilitation training guidance method and system based on visual features and agents

By constructing a two-dimensional feature assessment system based on U-Net, HRNet, and ST-GCN, and combining it with an agent based on the Transformer architecture to generate personalized multimodal feedback, the problems of inaccurate assessment and unsuitable feedback in hand rehabilitation training for Parkinson's disease were solved, significantly improving training effectiveness and patient compliance.

CN122369795APending Publication Date: 2026-07-10DALIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN UNIV
Filing Date
2026-03-25
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Current technologies for hand rehabilitation training in Parkinson's disease rely on single joint angle errors to assess movement accuracy, ignore the impact of hand position deviation, fail to adjust feedback strategies according to the patient's disease stage, rely heavily on single-modal feedback leading to poor adaptability, and lack the integration of personalized parameters, resulting in low training efficiency.

Method used

We use computer vision models based on U-Net and HRNet to extract key anatomical points of the hand, combine ST-GCN to eliminate coordinate drift, construct a two-dimensional feature fusion evaluation of Euclidean distance and joint angle difference, generate personalized multimodal feedback commands through a medical agent with Transformer architecture, dynamically adjust modal priority, and construct a closed-loop control loop.

Benefits of technology

It enables accurate differentiation between Parkinson's disease tremor and essential tremor, improves the accuracy and compliance of rehabilitation training, adapts to the needs of different rehabilitation stages, and enhances training efficiency and personalized adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369795A_ABST
    Figure CN122369795A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of intelligent agents and computer vision technology, specifically to a method and system for guiding Parkinson's disease hand rehabilitation training based on visual features and intelligent agents. It integrates hand movement videos captured by high frame rate event cameras, physiological sensor data, and prior medical knowledge to construct a closed-loop rehabilitation guidance system centered on a medical intelligent agent. This system utilizes deep neural networks to extract spatiotemporal features of key hand points, combines Euclidean distance error and joint angle error to construct a two-dimensional comprehensive evaluation index, and dynamically adjusts the evaluation weights according to the patient's UPDRS stage. Through a medical intelligent agent based on the Transformer architecture, combined with a multi-objective reinforcement learning strategy, personalized, multimodal rehabilitation feedback instructions are generated. This invention achieves precise quantitative assessment and personalized guidance of hand movements in Parkinson's disease patients, effectively distinguishing between Parkinson's tremor and essential tremor, and significantly improving the accuracy, compliance, and individual adaptability of rehabilitation training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent agents and computer vision technology, specifically to a method and system for guiding Parkinson's hand rehabilitation training based on visual features and intelligent agents. Background Technology

[0002] Parkinson's disease is a common neurodegenerative disease of the central nervous system in middle-aged and elderly people. Its core symptoms include resting tremor, rigidity, bradykinesia, and postural and gait disturbances. Resting tremor of the hands is a typical and specific symptom that severely affects patients' daily living abilities and motor function. Hand rehabilitation training is a core means of improving motor function in Parkinson's disease patients, and computer vision-based hand movement analysis technology provides a new direction for the quantitative guidance of rehabilitation training.

[0003] Current technologies primarily extract anatomical key points of the hand using hand key point detection models such as OpenPose and MediaPipeHands, using joint angle error as a single evaluation indicator to judge movement accuracy, and combining this with a static rule engine to generate standardized rehabilitation feedback instructions. However, this type of technology has significant limitations in clinical applications: Relying solely on joint angle errors to assess movement accuracy ignores the impact of hand position deviation on the quantification of Parkinson's disease tremor, leading to a bias in the quantification of movement accuracy, especially with low accuracy in assessing late-stage patients. The feedback generated by the static rule engine does not adjust the assessment and feedback strategy according to the differences in motor control ability of patients at different stages of the disease (early / middle / late stage). The fixed strategy is difficult to take into account the rehabilitation needs of different stages, resulting in low training efficiency. It relies heavily on single-modal feedback, such as voice or vision, which can easily lead to patient distraction or sensory fatigue. It does not combine the immediacy and accuracy of tactile feedback, and lacks a dynamic coordination mechanism for modal priority, resulting in poor adaptability. The system does not fully integrate personalized parameters such as patient age, medical history, and physiological data, resulting in low personalization of feedback instructions. Furthermore, it fails to form a closed-loop control system of perception-assessment-decision-feedback-optimization, making it difficult to continuously optimize the rehabilitation training effect. Summary of the Invention

[0004] The purpose of this invention is to propose a method and system for guiding Parkinson's hand rehabilitation training based on visual features and intelligent agents, which can adapt to the differences in motor ability of patients at different rehabilitation stages and accurately distinguish between different types of tremor, such as Parkinson's tremor and essential tremor.

[0005] According to a first aspect of the present disclosure, a method for guiding Parkinson's hand rehabilitation training based on visual features and intelligent agents is provided, comprising the following steps: The video stream of the patient's hand rehabilitation movements was acquired. The hand region (ROI) of each frame of the video stream was extracted using a computer vision model based on U-Net semantic segmentation. Then, the three-dimensional coordinates of the anatomical key points of the hand were obtained based on HRNet deep neural network regression. The key points include the fingertips, proximal interphalangeal joints, distal interphalangeal joints, metacarpophalangeal joints, and palmar points of the thumb, index finger, middle finger, ring finger, and little finger. A graph convolutional network (ST-GCN) incorporating spatiotemporal context information is used to perform temporal alignment and spatial topology modeling on the hand keypoint sequence in the video stream, eliminating coordinate drift caused by camera shake or minor displacement of the patient's hand, and outputting a corrected three-dimensional coordinate sequence of hand keypoints. Based on the corrected three-dimensional coordinate spatial relationship of key hand points, the Euclidean distance difference between the key points and the standard target position, and the angle difference between the actual angle and the standard angle of adjacent joints are obtained, forming a two-dimensional feature index. The Euclidean distance difference reflects the degree of hand position deviation, and the joint angle difference reflects the hand's ability to control the angle of movement. Then, according to the dynamic allocation of weights based on the patient's Parkinson's disease rehabilitation stage, the two-dimensional feature index is integrated into a comprehensive evaluation index. This index can distinguish between Parkinson's disease tremor and essential tremor. The comprehensive evaluation indicators are compared in real time with the standard action features defined by the doctor's prior knowledge to generate feature differences; The feature difference is combined with the patient's real-time physiological data to form a state vector, which is then input into a large medical intelligent agent model built on the Transformer architecture. The large medical intelligent agent model combines the patient's personalized parameters to generate personalized rehabilitation action feedback instructions for the patient's current state, guiding the patient to adjust the hand movement trajectory and minimize the feature difference.

[0006] In one embodiment, before obtaining the two-dimensional feature indicators, the changes in the patient's hand rehabilitation movements are judged through an action triggering mechanism. After triggering feature extraction, the Euclidean distance difference and angle difference are obtained. The action triggering condition is: in, Indicates the type of hand action event; and These are the Euclidean distances between key hand points at two consecutive time points; and These are the values ​​of the angles of adjacent joints at two consecutive moments. For time intervals; T The threshold for action change; For decision functions; Feature extraction is triggered when any of the following conditions are met: (1) The rate of change of Euclidean distance at key points exceeds the threshold, and the rate of change of joint angle exceeds the threshold, i.e.: (2) Triggered by compound conditions, i.e.: in, , and These are the trigger thresholds for Euclidean distance, angle, and composite changes, respectively; the judgment function f matches the action event type using an action pattern library defined by the doctor's prior knowledge. The types of movement events include standard Parkinson's rehabilitation movement categories such as clenching a fist, extending fingers, and rotating the wrist.

[0007] In one embodiment, a graph convolutional network incorporating spatiotemporal context information constructs an inter-joint topology graph to learn the spatial adjacency relationships and temporal motion trends of hand keypoints in adjacent frames of the video stream, and outputs a corrected 3D coordinate sequence. ,in Indicates the first A key point at a moment The coordinates.

[0008] In one embodiment, the two-dimensional feature indicators are integrated into a comprehensive evaluation indicator, specifically as follows: For any key point Based on the corrected three-dimensional coordinates, its position relative to the standard target position The difference in Euclidean distance is: Where k represents the three-dimensional coordinate axes (x, y, z) and the target position. Provided by a standard action model defined by the doctor's prior knowledge; For adjacent joint j, based on the corrected 3D coordinates, its actual angle Compared to standard angle The difference is: in, and The two vectors that constitute the joint are obtained from the coordinates of the key points; Euclidean distance difference and joint angle difference The following formula is used to integrate the comprehensive evaluation indicators: in The total number of key points. The total number of joints, weight and satisfy It will automatically adjust according to the patient's current recovery stage, with the following specific rules: In the early stages, the UPDRS score is <20, such as The assessment focuses on angle control capabilities. In the mid-term, the UPDRS score is 20-3, such as Equilibrium position and angle assessment; In the late stage, the UPDRS score is >35, such as Focus on positional accuracy assessment.

[0009] In one embodiment, personalized rehabilitation action feedback instructions are generated based on the patient's current state, specifically by: incorporating comprehensive assessment indicators... Difference with features Combined with the patient's real-time physiological data, a state vector is formed. ,in For real-time heart rate, This represents the current average joint angle. To standardize the scoring of Parkinson's disease rating scales; Define the set of executable actions for a large medical agent model based on the Transformer architecture. Each action It corresponds to a type of rehabilitation movement feedback instruction, including visually guided movements, voice prompt movements, tactile feedback movements, and training duration adjustment movements; Design a multi-objective reward function for a large medical agent model for policy update optimization: in: The smaller the feature difference, the higher the reward; After the patient performs the action A positive reward is given when the decline exceeds the threshold τ; Increased heart rate is considered an indicator of fatigue, and the penalty value is directly proportional to the heart rate. ,w , The weighting coefficients were optimized through clinical trials. The medical intelligent agent large model dynamically updates its strategy based on deep Q-network (DQN) or PPO algorithm. : Where γ∈(0,1) is the discount factor. The length of the training period; In the decision-making process of the large-scale medical intelligent agent model, personalized patient parameters are dynamically injected: P=[Age,Disease_Stage,Previous_Responses], Where: Age represents the patient's age, influencing the selection of action complexity; Disease_Stage represents the stage of Parkinson's disease (early / middle / late stage), determining the difficulty gradient of the action; Previous_Responses represents the historical action feedback effects, used to adjust the weights of the reward function. , , ; The medical intelligent agent model combines state vectors and personalized parameters to generate multimodal personalized rehabilitation action feedback instructions. If the feedback instructions include a combination of multiple modalities, the priority of different modalities is coordinated through preset rules, and visual modality attention weights are used. Visual modality attention weights , Dynamically allocate the intensity of each mode, and satisfy the following conditions: .

[0010] In one embodiment, the priority of different modalities is coordinated by: setting a preset feature difference threshold. The rehabilitation training scenario is divided into emergency adjustment scenario and stable training scenario; if If the scenario is determined to be an emergency adjustment, the haptic feedback mode will be triggered first; if the step... If the training scenario is determined to be stable, the visual guidance mode will be used first. The visual guidance mode is a low-intrusive feedback mode.

[0011] In one embodiment, a large medical intelligent agent model built on the Transformer architecture is fine-tuned for Parkinson's disease rehabilitation using LoRA low-rank adaptation technology. The specific construction process is as follows: A general pre-trained Transformer architecture model was selected as the baseline model, and the pre-trained weights of the baseline model were frozen. , For the input feature dimension, Output feature dimension; In the self-attention mechanism layer of the baseline Transformer architecture, a trainable side-path low-rank decomposition matrix is ​​injected. and , where rank This makes the forward propagation of the model become , For the input vector, ; A dedicated instruction fine-tuning dataset for Parkinson's rehabilitation was constructed, which includes rehabilitation medicine guidelines, expert hand rehabilitation movement correction corpus, and rehabilitation log data of historical Parkinson's patients. The baseline model with the injected low-rank decomposition matrix is ​​trained on the dedicated instruction fine-tuning dataset. During the training process, only the parameters of the low-rank decomposition matrices A and B are updated, so that the general pre-trained large model learns the pathological features of Parkinson's disease and the logic of hand rehabilitation training, forming a large medical intelligent agent model with professional capabilities in the field of Parkinson's rehabilitation.

[0012] According to a second aspect of the present disclosure, a Parkinson's hand rehabilitation training guidance system based on visual features and an intelligent agent is provided, comprising: The video stream acquisition and hand ROI extraction module acquires video streams of the patient's hand rehabilitation movements. It uses a computer vision model based on U-Net semantic segmentation to extract the hand region (ROI) from each frame of the video stream. Then, it uses HRNet deep neural network regression to obtain the three-dimensional coordinates of the anatomical key points of the hand. The key points include the fingertips, proximal interphalangeal joints, distal interphalangeal joints, metacarpophalangeal joints, and palmar points of the thumb, index finger, middle finger, ring finger, and little finger. The key point temporal correction and spatial modeling module uses a graph convolutional network (ST-GCN) that combines spatiotemporal context information to perform temporal alignment and spatial topology modeling on the hand key point sequence in the video stream, eliminating coordinate drift caused by camera shake or slight displacement of the patient's hand, and outputting the corrected three-dimensional coordinate sequence of hand key points. The dual-dimensional feature calculation and dynamic fusion assessment module, based on the corrected three-dimensional coordinate space relationship of key hand points, obtains the Euclidean distance difference between the key points and the standard target position, and the angle difference between the actual angle and the standard angle of adjacent joints, forming dual-dimensional feature indicators. The Euclidean distance difference reflects the degree of hand position deviation, and the joint angle difference reflects the hand's ability to control the angle of movement. Then, according to the dynamic weight allocation based on the patient's Parkinson's disease rehabilitation stage, the dual-dimensional feature indicators are integrated into a comprehensive assessment indicator. This indicator can distinguish between Parkinson's disease tremor and essential tremor. The dual-dimensional feature calculation and dynamic fusion evaluation module compares the comprehensive evaluation index with the standard action features defined by the doctor's prior knowledge in real time to generate feature difference values. The medical intelligent agent decision-making and personalized feedback module combines the feature difference with the patient's real-time physiological data into a state vector, which is then input into a large medical intelligent agent model built on the Transformer architecture. The large medical intelligent agent model combines the patient's personalized parameters to generate personalized rehabilitation action feedback instructions for the patient's current state, guiding the patient to adjust the hand movement trajectory and minimize the feature difference.

[0013] According to a third aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the memory, wherein the processor executes the program to implement the aforementioned Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents.

[0014] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the aforementioned method for guiding Parkinson's hand rehabilitation training based on visual features and an intelligent agent.

[0015] The advantages of the above technical solutions adopted in this invention compared with the prior art are as follows: 1. A dual-dimensional feature fusion assessment system based on Euclidean distance difference and joint angle difference was constructed. Combined with ST-GCN spatiotemporal modeling to eliminate coordinate drift, it effectively made up for the defects of single-dimensional assessment. It can accurately quantify hand position offset and angle control deviation, and greatly improve the accuracy of distinguishing different tremor types such as Parkinson's disease tremor and essential tremor, providing a reliable quantitative basis for clinical rehabilitation training.

[0016] 2. The rehabilitation stages are divided according to the patient's UPDRS score, and the weights of the dual-dimensional feature fusion are dynamically adjusted. The early stage focuses on the assessment of angle control ability, the middle stage on the assessment of balance position and angle, and the late stage on the assessment of position accuracy. This achieves the matching of rehabilitation assessment strategy with the patient's motor control ability and avoids the problem of low training efficiency caused by fixed strategies.

[0017] 3. A large-scale medical intelligent agent model is constructed based on the Transformer architecture and LoRA fine-tuning technology. It integrates real-time physiological data and generates multimodal personalized feedback instructions through reinforcement learning. Combined with a modal coordination mechanism that prioritizes tactile feedback in emergency scenarios and visual guidance in stable scenarios, it reduces invasiveness while ensuring the immediacy of feedback, significantly improving the patient's training task completion rate and enhancing rehabilitation training compliance.

[0018] 4. A closed-loop control circuit of "perception-assessment-decision-feedback-optimization" is constructed, which feeds back the difference in motion characteristics to the medical intelligent agent in real time, dynamically updates the rehabilitation strategy and training duration, and continuously guides patients to optimize their hand movement trajectory, thereby improving the accuracy of motion assessment and effectively improving the hand motor function of Parkinson's disease patients. It has high clinical application value and industrialization prospects. Attached Figure Description

[0019] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.

[0020] Figure 1 A flowchart for a Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents; Figure 2 Flowchart for hand keypoint detection and spatiotemporal modeling; Figure 3 This is a schematic diagram of a two-dimensional feature fusion model; Figure 4 Architecture diagram of a personalized rehabilitation training intelligent guidance system. Detailed Implementation

[0021] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0023] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0024] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems according to various embodiments of this disclosure. It should be noted that each block in a flowchart or block diagram may represent a module, segment, or portion of code, which may include one or more executable instructions for implementing the logical functions specified in the various embodiments. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, may be implemented using a dedicated hardware-based system that performs the specified functions or operations, or using a combination of dedicated hardware and computer instructions.

[0025] Example 1: This embodiment provides a Parkinson's disease hand rehabilitation training guidance method based on visual features and intelligent agents. It aims to integrate hand movement videos captured by high frame rate event cameras, physiological sensor data, and prior medical knowledge to construct a closed-loop rehabilitation guidance system centered on a medical intelligent agent. This method utilizes deep neural networks to extract spatiotemporal features of key hand points, combines Euclidean distance error (EDE) and joint angle error (JAE) to construct a two-dimensional comprehensive evaluation index, Score(t), and dynamically adjusts the evaluation weights according to the patient's UPDRS stage. Furthermore, through a medical intelligent agent based on a Transformer architecture, combined with a multi-objective reinforcement learning strategy, personalized, multimodal rehabilitation feedback instructions are generated. Simultaneously, the priority and intensity of tactile, visual, and verbal feedback are adaptively coordinated based on changes in the feature difference ΔScore(t), ultimately forming an intelligent rehabilitation closed loop of "perception—evaluation—decision—feedback—optimization," significantly improving the accuracy, compliance, and individual adaptability of rehabilitation training for Parkinson's disease patients. The method includes the following steps: S1. Acquire a video stream of the patient's hand rehabilitation movements, use a computer vision model based on U-Net semantic segmentation to extract the hand region (ROI) of each frame of the video stream, and then use HRNet deep neural network regression to obtain the three-dimensional coordinates of 21 anatomical key points of the hand. The key points include the fingertips, proximal interphalangeal joints, distal interphalangeal joints, metacarpophalangeal joints and palmar points of the thumb, index finger, middle finger, ring finger and little finger. Specifically, a high frame rate event camera (such as PropheseeGen4) is used to capture video streams of the patient's hand rehabilitation movements. The video resolution is 1920×1080 pixels, and the frame rate is 30fps, used to capture minute hand tremors and continuous movement trajectories. The event camera records hand movement events in real time using an asynchronous event-driven mode, reducing data redundancy and improving capture accuracy in dynamic scenes. A computer vision model is used to extract the region of interest (ROI) of the hand from each frame of the video stream. The computer vision model is a deep learning-based image segmentation network that accurately identifies the hand region using U-Net semantic segmentation technology.

[0026] Construct hand action trigger conditions ,in , , These are the trigger thresholds for Euclidean distance, angle, and composite features, respectively. Feature extraction is triggered when any of the following conditions are met: or This triggering mechanism monitors hand movements in real time via an event camera, ensuring that keypoint detection is initiated only when a valid action occurs, reducing redundant computation. The event camera continuously triggers the subject's hand tremor sequence to acquire a video stream.

[0027] S2. A graph convolutional network (ST-GCN) incorporating spatiotemporal context information is used to perform temporal alignment and spatial topology modeling on the hand keypoint sequence in the video stream, and outputs a corrected three-dimensional coordinate sequence of the hand keypoints; Specifically, a graph convolutional network (ST-GCN) is used to construct an inter-joint topology graph, learn the spatial adjacency relationships and temporal motion trends of key points in adjacent frames, and output a corrected 3D coordinate sequence. This eliminates coordinate drift caused by camera shake or slight displacement of the patient's hand.

[0028] S3. Based on the corrected three-dimensional coordinate space relationship of the key points of the hand, the Euclidean distance difference between the key points and the standard target position, and the angle difference between the actual angle and the standard angle of the adjacent joint are obtained to form a two-dimensional feature index. The Euclidean distance difference reflects the degree of hand position deviation, and the joint angle difference reflects the hand's ability to control the angle of movement. Then, according to the dynamic allocation of weights based on the patient's Parkinson's disease rehabilitation stage, the two-dimensional feature index is integrated into a comprehensive evaluation index. This index can distinguish between Parkinson's disease tremor and essential tremor. Specifically, based on the three-dimensional coordinate space relationship of key hand points, the Euclidean distance difference between any key point and the target position is obtained. , among which the target location The standard action model is provided by the doctor's prior knowledge: Where k represents the three-dimensional coordinate axes (x, y, z); Obtain the angle difference between adjacent joints, and then perform operations on adjacent joints. j Define its actual angle Compared to standard angle The difference is: in, and The two vectors that constitute the joint are obtained from the coordinates of the key points.

[0029] Existing technologies often use single joint angle error (JAE) to assess movement accuracy, but neglect the impact of hand position deviation (EDE) on the quantification of tremor in Parkinson's disease. Experiments show that assessment methods relying solely on JAE have an accuracy rate of only 72.3% in advanced patients, while a two-dimensional fusion model incorporating EDE can improve the accuracy to 87.6%. This invention utilizes Euclidean distance error... Joint angle error The joint modeling is as follows: in, The total number of key points. The total number of joints, weight and satisfy .

[0030] Patients are categorized into stages based on the Unified Parkinson's Disease Rating Scale (UPDRS), with weights dynamically adjusted. Specifically, the early stage (UPDRS score <20) is classified as... Emphasis is placed on angle control ability; mid-term (UPDRS score 20-35) is 5. Balance position and angle; late stage (UPDRS score > 35) is Emphasis is placed on positional accuracy; S4. The comprehensive evaluation indicators are compared with the standard action features defined by the doctor's prior knowledge in real time to generate feature differences; Specifically, the dual-dimensional feature fusion index Standard action characteristics defined by the doctor's prior knowledge Perform real-time comparison and generate feature differences. ; S5. The feature difference is combined with the patient's real-time physiological data to form a state vector, which is then input into a large medical intelligent agent model built on the Transformer architecture. The large medical intelligent agent model combines the patient's personalized parameters to generate personalized rehabilitation action feedback instructions for the patient's current state, guiding the patient to adjust the hand movement trajectory and minimize the feature difference.

[0031] A large-scale medical model based on LoRA (Low-Rank Adaptation) fine-tuning is constructed as the core of the agent. A general pre-trained Transformer architecture model (such as LLaMA) is selected as the baseline. A dedicated instruction fine-tuning dataset for Parkinson's rehabilitation is constructed, which includes rehabilitation medicine guidelines, expert correction corpora of movements, and various rehabilitation log data from historical patients. LoRA technology is used to efficiently fine-tune the parameters of the base model.

[0032] Specifically, freeze the pre-trained weights of the baseline model. Inject a trainable side-path low-rank factorization matrix into the Transformer's self-attention layer. and (where rank) This makes forward propagation computation become By training on the dedicated instruction fine-tuning dataset, updating only the parameters of matrices A and B, the general-purpose large model learns Parkinson's pathological features and rehabilitation logic, generating a medical intelligent agent with domain-specific knowledge. This agent models the long-range temporal dependency between historical hand keypoint sequences and candidate feedback instructions through a self-attention mechanism to generate personalized rehabilitation feedback strategies adapted to the patient's current state.

[0033] Fusion index of dynamically calculated two-dimensional features Combined with the patient's real-time physiological data (including heart rate HR(t), muscle tone, and UPDRS score) to form a state vector. ; Define the set of executable actions for a medical intelligent agent. Each action It corresponds to a type of rehabilitation action feedback instruction, including visual guidance actions (such as AR interface overlay of target outline), voice prompt actions (such as please bend your thumb inward 15°), tactile feedback actions (such as smart gloves applying directional vibration), and training duration adjustment actions (such as extending or shortening the current training cycle). Constructing a multi-objective reward function ,in , , Optimized through clinical trials. , (when When the drop exceeds the threshold τ). ; Medical intelligent agents dynamically update strategies based on deep Q-networks (DQN) or PPO algorithms: in, This is a discount factor used to balance long-term and short-term returns, with a value range of 0.95 ≤ ≤0.99, T This refers to the length of the training cycle. It ensures that the medical AI agent prioritizes the current motor feedback during rehabilitation training while also considering long-term rehabilitation goals.

[0034] Personalized parameters of patients The state vector S(t) is fused with features through an attention mechanism. The specific steps are as follows: Will and Input gated recurrent unit (GRU) to generate age-stage embedding vector ; Will Extracting historical feedback features using a convolutional neural network (CNN) ; Through attention weight , Dynamic adjustment and Contribution level: in, This is integrated into medical intelligent agents to enable the generation of personalized rehabilitation strategies. Among these, The patient's age is used to assess the age-related correlation between motor ability and rehabilitation needs; The stages of Parkinson's disease (early / middle / late stage) serve as a key basis for dynamic weight allocation; This records historical action feedback effects, including patient response data to verbal, visual, and tactile commands. This parameter set is correlated with the current state vector. The joint modeling drives the medical intelligent agent to generate rehabilitation action feedback instructions adapted to the individual characteristics of patients, ensuring the clinical rationality and personalized adaptability of the strategy output.

[0035] When the generated feedback instructions contain a combination of multimodalities, the medical agent is based on feature differences. and preset threshold The relationship between dynamic coordination modal priorities; if If the scenario is determined to be an emergency adjustment, haptic feedback will be activated first (high priority, response latency <100ms); if If the scenario is deemed a stable training scenario, visual guidance (low-intrusive, AR interface overlaid with target outline) will be prioritized. Dynamically adjust multimodal feedback weights based on attention mechanism , , The weight allocation is determined by the feature difference. Driven by real-time physiological data of patients, and meeting normalization constraints. Specifically, the softmax function is used to optimize the modal priority vector. Mapping is performed to output the dynamic contribution coefficients of each modality, ensuring that the feedback intensity is adapted to the patient's current rehabilitation needs and environmental disturbances; according to The convergence speed is dynamically adjusted to change the current training duration, adaptively extending or shortening the training cycle to ensure a balance between training efficiency and patient comfort. This is based on the feature difference. Based on the patient's physiological data, personalized instructions are generated, including visual guidance (an AR interface overlaid with a virtual outline of the target hand posture), voice prompts (please bend your thumb inward 15°), and tactile feedback (a smart glove applies a slight vibration to prompt you to adjust the direction). In addition, by monitoring the coordinates of key points of the hand in real time, the feature difference ΔScore(t) is fed back to the medical intelligent agent to form a closed-loop control loop, continuously optimize the rehabilitation training effect, minimize the feature difference ΔScore(t), and guide the patient to adjust the hand movement trajectory. Through the above steps, this invention achieves accurate assessment and personalized guidance of hand movements in Parkinson's disease patients, effectively distinguishes between different types of tremor such as Parkinson's disease tremor and essential tremor, provides quantitative guidance for rehabilitation training, and significantly improves the effectiveness of rehabilitation training and patient compliance.

[0036] Example 2: This embodiment provides a Parkinson's hand rehabilitation training guidance system based on visual features and intelligent agents, including: The video stream acquisition and hand ROI extraction module acquires video streams of the patient's hand rehabilitation movements. It uses a computer vision model based on U-Net semantic segmentation to extract the hand region (ROI) from each frame of the video stream. Then, it uses HRNet deep neural network regression to obtain the three-dimensional coordinates of the anatomical key points of the hand. The key points include the fingertips, proximal interphalangeal joints, distal interphalangeal joints, metacarpophalangeal joints, and palmar points of the thumb, index finger, middle finger, ring finger, and little finger. The key point temporal correction and spatial modeling module uses a graph convolutional network (ST-GCN) that combines spatiotemporal context information to perform temporal alignment and spatial topology modeling on the hand key point sequence in the video stream, eliminating coordinate drift caused by camera shake or slight displacement of the patient's hand, and outputting the corrected three-dimensional coordinate sequence of hand key points. The dual-dimensional feature calculation and dynamic fusion assessment module, based on the corrected three-dimensional coordinate space relationship of key hand points, obtains the Euclidean distance difference between the key points and the standard target position, and the angle difference between the actual angle and the standard angle of adjacent joints, forming dual-dimensional feature indicators. The Euclidean distance difference reflects the degree of hand position deviation, and the joint angle difference reflects the hand's ability to control the angle of movement. Then, according to the dynamic weight allocation based on the patient's Parkinson's disease rehabilitation stage, the dual-dimensional feature indicators are integrated into a comprehensive assessment indicator. This indicator can distinguish between Parkinson's disease tremor and essential tremor. The dual-dimensional feature calculation and dynamic fusion evaluation module compares the comprehensive evaluation index with the standard action features defined by the doctor's prior knowledge in real time to generate feature difference values. The medical intelligent agent decision-making and personalized feedback module combines the feature difference with the patient's real-time physiological data into a state vector, which is then input into a large medical intelligent agent model built on the Transformer architecture. The large medical intelligent agent model combines the patient's personalized parameters to generate personalized rehabilitation action feedback instructions for the patient's current state, guiding the patient to adjust the hand movement trajectory and minimize the feature difference.

[0037] The above modules can be deployed on the same device or distributed devices; the division of modules is only a functional logic description and does not limit the specific physical boundaries or implementation order.

[0038] Example 3: An electronic device is provided for running the aforementioned "A Parkinson's Hand Rehabilitation Training Guidance Method Based on Visual Features and Intelligent Agents". The electronic device includes: a processor, a memory, and optional communication interfaces / display devices / input devices, etc.; the memory stores a computer program that can run on the processor, and when the processor executes the program, it implements steps S1 to S5 of the method described in Embodiment 1, specifically including but not limited to: S1. Acquire a video stream of the patient's hand rehabilitation movements, use a computer vision model based on U-Net semantic segmentation to extract the hand region (ROI) of each frame of the video stream, and then use HRNet deep neural network regression to obtain the three-dimensional coordinates of the anatomical key points of the hand. The key points include the fingertips, proximal interphalangeal joints, distal interphalangeal joints, metacarpophalangeal joints and palmar points of the thumb, index finger, middle finger, ring finger and little finger. S2. A graph convolutional network (ST-GCN) incorporating spatiotemporal context information is used to perform temporal alignment and spatial topology modeling on the hand key point sequence in the video stream, eliminating coordinate drift caused by camera shake or slight displacement of the patient's hand, and outputting a corrected three-dimensional coordinate sequence of hand key points. S3. Based on the corrected three-dimensional coordinate space relationship of the key points of the hand, the Euclidean distance difference between the key points and the standard target position, and the angle difference between the actual angle and the standard angle of the adjacent joint are obtained to form a two-dimensional feature index. The Euclidean distance difference reflects the degree of hand position deviation, and the joint angle difference reflects the hand's ability to control the angle of movement. Then, according to the dynamic allocation of weights based on the patient's Parkinson's disease rehabilitation stage, the two-dimensional feature index is integrated into a comprehensive evaluation index. This index can distinguish between Parkinson's disease tremor and essential tremor. S4. The comprehensive evaluation indicators are compared with the standard action features defined by the doctor's prior knowledge in real time to generate feature differences; S5. The feature difference is combined with the patient's real-time physiological data to form a state vector, which is then input into a large medical intelligent agent model built on the Transformer architecture. The large medical intelligent agent model combines the patient's personalized parameters to generate personalized rehabilitation action feedback instructions for the patient's current state, guiding the patient to adjust the hand movement trajectory and minimize the feature difference.

[0039] The electronic device hardware can be one of a server, personal computer, workstation, industrial controller, edge computing device, or mobile terminal; the processor can be a general-purpose CPU, GPU, NPU, FPGA, or a combination thereof; the memory can be RAM, ROM, flash memory, or disk array. The device can interact with local / remote data storage (acquiring observation data and outputting inversion results) through a communication interface. The above hardware configuration does not constitute a limitation of the present invention.

[0040] Example 4: A computer-readable storage medium storing a computer program, which, when run on a processor of an electronic device, causes the program to execute the method steps S1 to S5 described in Embodiment 1; the storage medium may be a disk, optical disk, flash memory, solid-state drive, read-only memory, random access memory, or any combination of the above media.

[0041] Application examples This embodiment uses Ubuntu 20.04 as the development environment, PyTorch 1.12 as the deep learning framework, and Python 3.8 as the development language. It employs the Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents of this invention to complete personalized rehabilitation training guidance for patients with early to late-stage Parkinson's disease. The steps include: Step 1: Deploy the PropheseeGen4 event camera in front of the rehabilitation training table, set the resolution to 1920×1080, and the frame rate to 30fps, to capture video streams of the patient performing standard actions such as "thumb to finger" and "clenching and unfolding the fist"; simultaneously connect to the smart glove (integrating IMU and vibration motor) and wrist-worn heart rate sensor to obtain real-time data. and .

[0042] Step 2: Segment the hand ROI using a pre-trained U-Net model, then use HRNet to extract the 3D coordinates of 21 keypoints; perform spatiotemporal alignment of the keypoint sequence of 30 consecutive frames using ST-GCN, and output the corrected coordinates. .

[0043] Step 3: Follow the standard movement model defined by the doctor. Calculate EDE With JAE According to the formula Fusion assessment score; α and β are automatically set based on the patient's UPDRS score (early stage: 0.3 / 0.7; mid-stage: 0.5 / 0.5; late stage: 0.7 / 0.3).

[0044] Step 4: Construct the state vector Initializing the medical intelligent agent: A Transformer model with 7B parameters is selected as the base, and LoRA technology is used for domain-specific fine-tuning. Setting LoRA parameters: rank scaling factor Dropout=0.1, for and The module was adapted. Fine-tuning was performed using 3000 Q&A data points from Parkinson's hand rehabilitation experts, trained for 5 epochs, until the model converged to a state possessing prior knowledge of rehabilitation medicine. The reward function was set to... , .

[0045] Step 5: Generate multimodal feedback instruction set A, if > (Set to 0.35), then the directional vibration of the smart glove will be triggered first (response delay <80ms); otherwise, the outline of the target hand will be superimposed on the AR glasses; simultaneously through Dynamic allocation , , , .

[0046] The system updates Pt′ every 200ms and recalculates. And feed back to the intelligent agent to form a closed-loop control; when for 5 consecutive seconds When the value is less than 0.1, the training duration of the current action will be automatically extended by 10%.

[0047] To verify the efficacy, a 4-week controlled trial was conducted in 30 Parkinson's patients (UPDRS score 12–42). The experimental group used the method of this invention, while the control group used a traditional single visual feedback system. The results showed that the accuracy of motion assessment in the experimental group reached 87.6%, significantly higher than the 72.3% in the control group (p<0.01); patient training compliance improved by 34.2% (based on task completion rate and subjective questionnaire); and the average response time to tactile feedback in late-stage patients (UPDRS>35) was 92ms, meeting the needs for emergency adjustments.

[0048] Table 1. Performance comparison of the method of the present invention at different rehabilitation stages: The above embodiments demonstrate that the present invention effectively solves the core problems in Parkinson's rehabilitation training, such as inaccurate assessment, delayed feedback, and large individual differences, through multimodal feature fusion, personalized strategies driven by medical intelligent agents, and adaptive multimodal feedback mechanisms, and has significant clinical application value.

[0049] Those skilled in the art will understand that the modules or steps described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, which can then be stored in a storage device for execution by a computer device. Alternatively, they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. This disclosure is not limited to any particular combination of hardware and software.

[0050] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0051] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.

Claims

1. A method for guiding Parkinson's hand rehabilitation training based on visual features and intelligent agents, characterized in that, Includes the following steps: The video stream of the patient's hand rehabilitation movements was acquired. The hand region of each frame of the video stream was extracted using a computer vision model based on U-Net semantic segmentation. Then, the three-dimensional coordinates of the anatomical key points of the hand were obtained by regression based on HRNet deep neural network. A graph convolutional network incorporating spatiotemporal context information is used to perform temporal alignment and spatial topology modeling on the hand keypoint sequence in the video stream, and outputs a corrected three-dimensional coordinate sequence of the hand keypoints. Based on the corrected three-dimensional coordinate spatial relationship of key hand points, the Euclidean distance difference between the key points and the standard target position, and the angle difference between the actual angle and the standard angle of adjacent joints are obtained, forming a two-dimensional feature index. The Euclidean distance difference reflects the degree of hand position deviation, and the joint angle difference reflects the hand's ability to control the angle of movement. Then, according to the dynamic allocation of weights based on the patient's Parkinson's disease rehabilitation stage, the two-dimensional feature index is integrated into a comprehensive evaluation index. This index can distinguish between Parkinson's disease tremor and essential tremor. The comprehensive evaluation indicators are compared in real time with the standard action features defined by the doctor's prior knowledge to generate feature differences; The feature difference is combined with the patient's real-time physiological data to form a state vector, which is then input into a large medical intelligent agent model built on the Transformer architecture. The large medical intelligent agent model combines the patient's personalized parameters to generate personalized rehabilitation action feedback instructions for the patient's current state, guiding the patient to adjust the hand movement trajectory and minimize the feature difference.

2. The Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents according to claim 1, characterized in that, Before obtaining the two-dimensional feature indicators, the changes in the patient's hand rehabilitation movements are judged through an action triggering mechanism. After triggering feature extraction, the Euclidean distance difference and angle difference are obtained. The action triggering condition is as follows: in, Indicates the type of hand action event; and These are the Euclidean distances between key hand points at two consecutive time points; and These are the values ​​of the angles of adjacent joints at two consecutive moments. For time intervals; T The threshold for action change; For decision functions; Feature extraction is triggered when any of the following conditions are met: (1) The rate of change of Euclidean distance at key points exceeds the threshold, and the rate of change of joint angle exceeds the threshold, i.e.: (2) Triggered by compound conditions, i.e.: in, , and These are the trigger thresholds for Euclidean distance, angle, and composite changes, respectively; the judgment function f matches the action event type using an action pattern library defined by the doctor's prior knowledge. The types of movement events include standard Parkinson's rehabilitation movement categories such as clenching a fist, extending fingers, and rotating the wrist.

3. The Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents according to claim 1, characterized in that, Graph convolutional networks that incorporate spatiotemporal contextual information learn the spatial adjacency relationships and temporal motion trends of hand keypoints in adjacent frames of a video stream by constructing an inter-joint topology graph, and output a corrected 3D coordinate sequence. ,in Indicates the first A key point at a moment The coordinates.

4. The Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents according to claim 3, characterized in that, The two-dimensional feature indicators are integrated into a comprehensive evaluation indicator, specifically: For any key point Based on the corrected three-dimensional coordinates, its position relative to the standard target position The difference in Euclidean distance is: Where k represents the three-dimensional coordinate axes (x, y, z) and the target position. Provided by a standard action model defined by the doctor's prior knowledge; For adjacent joint j, based on the corrected 3D coordinates, its actual angle Compared to standard angle The difference is: in, and The two vectors that constitute the joint are obtained from the coordinates of the key points; Euclidean distance difference and joint angle difference The following formula is used to integrate the comprehensive evaluation indicators: in The total number of key points. The total number of joints, weight and satisfy It will automatically adjust according to the patient's current recovery stage, with the following specific rules: In the early stages, the UPDRS score is <20, focusing on the assessment of angle control ability; In the mid-term stage, the UPDRS score is 20-3, and the balance position and angle are assessed. In the later stages, when the UPDRS score is >35, the focus is on assessing positional accuracy.

5. The Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents according to claim 1, characterized in that, Generate personalized rehabilitation action feedback instructions based on the patient's current state, specifically by incorporating comprehensive assessment indicators. Difference with features Combined with the patient's real-time physiological data, a state vector is formed. ,in For real-time heart rate, This represents the current average joint angle. To standardize the scoring of Parkinson's disease rating scales; Define the set of executable actions for a large medical agent model based on the Transformer architecture. Each action It corresponds to a type of rehabilitation movement feedback instruction, including visually guided movements, voice prompt movements, tactile feedback movements, and training duration adjustment movements; Design a multi-objective reward function for a large medical agent model for policy update optimization: in: The smaller the feature difference, the higher the reward; After the patient performs the action A positive reward is given when the decline exceeds the threshold τ; Increased heart rate is considered an indicator of fatigue, and the penalty value is directly proportional to the heart rate. ,w , The weighting coefficients were optimized through clinical trials. Large-scale medical intelligent agent models dynamically update strategies based on deep Q-networks or PPO algorithms. : Where γ∈(0,1) is the discount factor. The length of the training period; In the decision-making process of the large-scale medical intelligent agent model, personalized patient parameters are dynamically injected: P=[Age,Disease_Stage,Previous_Responses], Where: Age represents the patient's age, influencing the selection of action complexity; Disease_Stage represents the stage of Parkinson's disease, determining the difficulty gradient of the action; Previous_Responses represents the historical action feedback effects, used to adjust the weights of the reward function. , , ; The medical intelligent agent model combines state vectors and personalized parameters to generate multimodal personalized rehabilitation action feedback instructions. If the feedback instructions include a combination of multiple modalities, the priority of different modalities is coordinated through preset rules, and visual modality attention weights are used. Visual modality attention weights , Dynamically allocate the intensity of each mode, and satisfy the following conditions: .

6. The Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents according to claim 5, characterized in that, The priority method for coordinating different modalities is as follows: preset a critical threshold for feature differences. The rehabilitation training scenario is divided into emergency adjustment scenario and stable training scenario; if If the scenario is determined to be an emergency adjustment, the haptic feedback mode will be triggered first; if the step... If the training scenario is determined to be stable, the visual guidance mode will be used first. The visual guidance mode is a low-intrusive feedback mode.

7. The Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents according to claim 1, characterized in that, A large medical intelligent agent model based on the Transformer architecture was used to fine-tune the model for Parkinson's disease rehabilitation through LoRA low-rank adaptation technology. The specific construction process is as follows: A general pre-trained Transformer architecture model was selected as the baseline model, and the pre-trained weights of the baseline model were frozen. , For the input feature dimension, Output feature dimension; In the self-attention mechanism layer of the baseline Transformer architecture, a trainable side-path low-rank decomposition matrix is ​​injected. and , where rank This makes the forward propagation of the model become , For the input vector, ; A dedicated instruction fine-tuning dataset for Parkinson's rehabilitation was constructed, which includes rehabilitation medicine guidelines, expert hand rehabilitation movement correction corpus, and rehabilitation log data of historical Parkinson's patients. The baseline model with the injected low-rank decomposition matrix is ​​trained on the dedicated instruction fine-tuning dataset. During the training process, only the parameters of the low-rank decomposition matrices A and B are updated, so that the general pre-trained large model learns the pathological features of Parkinson's disease and the logic of hand rehabilitation training, forming a large medical intelligent agent model with professional capabilities in the field of Parkinson's rehabilitation.

8. A Parkinson's hand rehabilitation training guidance system based on visual features and intelligent agents, characterized in that, include: The video stream acquisition and hand ROI extraction module acquires video streams of patients' hand rehabilitation movements, uses a computer vision model based on U-Net semantic segmentation to extract the hand region from each frame of the video stream, and then uses HRNet deep neural network regression to obtain the three-dimensional coordinates of the anatomical key points of the hand. The key point temporal correction and spatial modeling module uses a graph convolutional network that combines spatiotemporal context information to perform temporal alignment and spatial topology modeling on the hand key point sequence in the video stream, and outputs the corrected hand key point three-dimensional coordinate sequence. The dual-dimensional feature calculation and dynamic fusion assessment module, based on the corrected three-dimensional coordinate space relationship of key hand points, obtains the Euclidean distance difference between the key points and the standard target position, and the angle difference between the actual angle and the standard angle of adjacent joints, forming dual-dimensional feature indicators. The Euclidean distance difference reflects the degree of hand position deviation, and the joint angle difference reflects the hand's ability to control the angle of movement. Then, according to the dynamic weight allocation based on the patient's Parkinson's disease rehabilitation stage, the dual-dimensional feature indicators are integrated into a comprehensive assessment indicator. This indicator can distinguish between Parkinson's disease tremor and essential tremor. The dual-dimensional feature calculation and dynamic fusion evaluation module compares the comprehensive evaluation index with the standard action features defined by the doctor's prior knowledge in real time to generate feature difference values. The medical intelligent agent decision-making and personalized feedback module combines the feature difference with the patient's real-time physiological data into a state vector, which is then input into a large medical intelligent agent model built on the Transformer architecture. The large medical intelligent agent model combines the patient's personalized parameters to generate personalized rehabilitation action feedback instructions for the patient's current state, guiding the patient to adjust the hand movement trajectory and minimize the feature difference.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running thereon, characterized in that, When the processor executes the program, it implements the Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the Parkinson's hand rehabilitation training guidance method based on visual features and intelligent agents as described in any one of claims 1-7.