Multi-agent-based stroke automatic evaluation and rehabilitation management method and system
Through video comparison and biomechanical analysis of multi-agent systems, a personalized rehabilitation plan is generated, which solves the problems of environmental sensitivity and privacy protection in the existing technology, and realizes the intelligence and safety of stroke rehabilitation management. Patients can conduct efficient rehabilitation training at home.
Patent Information
- Application Number
- CN202510379517.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
The existing rehabilitation exercise identification methods are insufficient in terms of environmental conditions, privacy protection, computing resource consumption and data privacy leakage, and it is difficult to achieve efficient, safe and personalized stroke rehabilitation management.
A multi-agent system is adopted, including a patient video comparison expert model, a rehabilitation evaluation expert model and a rehabilitation exercise recommendation expert model. Through video data processing, biomechanical analysis and reinforcement learning, a personalized rehabilitation plan is generated to achieve action feature extraction, abnormality detection and safe and gradual rehabilitation training.
The real-time, security and user-friendliness of stroke rehabilitation management have been improved, and patients can conduct intelligent rehabilitation training at home, reducing their dependence on environmental conditions and privacy risks.
Smart Images

Figure CN120299612A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rehabilitation exercises, and in particular to a multi-agent based automatic stroke assessment and rehabilitation management method and system. Background Art
[0002] Currently, in the field of rehabilitation exercise recognition, the combination of sensors and computer vision technology is an emerging topic. Compared with traditional rehabilitation exercise recognition methods, inertial measurement unit (IMU) sensors and computer vision (CV) have higher accuracy, objectivity, and personalization, can provide real-time feedback and adjustment, and can better improve the training effect of patients. At the same time, remote monitoring and intelligent management also make the rehabilitation process more efficient and convenient. However, in some cases, such as scenes with insufficient light, overlapping of multiple objects, complex background, etc., or due to users' concerns about privacy, the applicability of computer vision is greatly reduced, and the visual recognition of human skeletons consumes a large amount of computer resources and is difficult to be deployed lightweight to edge devices.
[0003] Traditional rehabilitation exercise recognition methods mostly rely on doctors' observations and manual measurements, emphasizing qualitative evaluation and manual records. These methods can help understand the overall motor ability of patients, but they usually lack accurate quantitative data and may be affected by the experience and subjective factors of operators. With the development of technology, more modern methods based on sensors and intelligent devices have gradually replaced these traditional means, providing more accurate and real-time motion analysis and feedback.
[0004] Current rehabilitation exercise recognition methods include computer vision, wearable sensors, inertial motion capture, remote monitoring and evaluation, etc. Although these technologies are effective to a certain extent, they generally have some problems, including being sensitive to environmental conditions, methods based on optical sensors being difficult to achieve ideal results in strong light or poor lighting environments, the frequent wearing of wearable devices being relatively inconvenient, possibly infringing on privacy, and the real-time performance of vision technology still having problems, etc.
[0005] The invention patent application with the application publication number CN119400364A discloses a rehabilitation training management system and method for stroke patients based on an Internet platform. The rehabilitation training management system for stroke patients based on an Internet platform includes: a training plan generation module, a training time dynamic adjustment module, and a training intensity dynamic adjustment module. The disadvantages of this method are that there are risks in data privacy: sensitive health data of patients is centrally stored and transmitted, there is a hidden danger of privacy leakage. High dependence on labeled data: The action evaluation model requires a large amount of labeled video data, the data acquisition cost is high, and the dependence on labeled data is high. Lack of flexibility: The centralized architecture in this patent is difficult to quickly expand new functions or adapt to diverse requirements. Summary of the Invention
[0006] To solve the above technical problems, a method and system for automatic stroke assessment and rehabilitation management based on multi-agent are proposed in the present invention. By introducing the currently popular multi-agent module, the stroke rehabilitation management is made intelligent, enabling patients to carry out rehabilitation training at home.
[0007] The first object of the present invention is to provide a method for automatic stroke assessment and rehabilitation management based on multi-agent, including obtaining patient video data and performing preprocessing. It is characterized by further comprising the following steps:
[0008] Step 1: Training the first agent as a patient video comparison expert model;
[0009] Step 2: Training the second agent as a rehabilitation assessment expert model;
[0010] Step 3: Training the third agent as a rehabilitation exercise recommendation expert model;
[0011] Step 4: Establishing a collaborative connection among the three expert models. The patient video comparison expert model completes action feature extraction and anomaly detection. The rehabilitation assessment expert model completes the quantitative assessment of the degree of motor function recovery. The rehabilitation exercise recommendation expert model completes the generation of a safe and progressive rehabilitation plan;
[0012] Step 5: Outputting the rehabilitation plan.
[0013] Preferably, the obtaining of patient video data and performing preprocessing includes the following sub-steps:
[0014] Step 01: Collecting video by recording the rehabilitation action video of stroke patients through the camera of a mobile device;
[0015] Step 02: Using H.264 encoding to unify the format and adjust the resolution of the collected data to complete data standardization;
[0016] Step 03: Performing preprocessing on the standardized data.
[0017] In any of the above solutions, preferably, the preprocessing includes denoising, clipping, and temporal alignment.
[0018] In any of the above solutions, preferably, Step 1 includes the following sub-steps:
[0019] Step 11: Using 3D-CNN to extract the spatio-temporal dynamic features of the patient video;
[0020] Step 12: Enhancing the weight of key action frames through an attention mechanism;
[0021] Step 13: Comparing the feature similarity of different patient videos through cosine similarity to complete modal fusion;
[0022] Step 14: Complete the training of the patient video comparison expert model based on the Triplet Loss function.
[0023] In any of the above solutions, preferably, step 11 includes splitting the video clip into a continuous frame sequence, each clip containing T frames and having a resolution of H×W. The convolutional kernel slides in three dimensions: space (H×W) and time (H×W), while extracting spatial features and temporal dynamics.
[0024] In any of the above solutions, preferably, step 12 includes identifying the frames crucial for rehabilitation assessment through temporal attention and focusing on the significant regions of joint movement through spatial attention. When significant regions of movement appear in the patient video, the weight of the key action frames is enhanced.
[0025] In any of the above solutions, preferably, step 13 includes normalizing the feature vectors extracted by 3D-CNN, performing similarity measurement, mapping the video features and physiological indicators to the same embedding space, and completing modality fusion by assigning different weights to different features.
[0026] In any of the above solutions, preferably, the Triplet Loss function consists of three parts: an anchor A, a positive sample P, and a negative sample N. Among them, the anchor A is a certain reference sample, the positive sample P is a sample belonging to the same category as the anchor A, and the negative sample N is a sample of a different category from the anchor.
[0027] In any of the above solutions, preferably, the Triplet Loss function L is defined as
[0028] L = max(d(A, P) - d(A, N) + margin, 0)
[0029] where d(A, P) is the distance between the anchor and the positive sample, d(A, N) is the distance between the anchor and the negative sample, and margin is a preset margin.
[0030] In any of the above solutions, preferably, step 2 includes the following sub-steps:
[0031] Step 21: Extract the temporal data of skeletal joint points through Openpose;
[0032] Step 22: Calculate biomechanical indicators through motion features to complete the analysis of the motion trajectory;
[0033] Step 23: Adopt feature splicing technology to fuse the RGB features and skeletal pose features of the video to complete modality fusion;
[0034] Step 24: Complete the training of the rehabilitation assessment expert model based on the temporal classification model Transformer.
[0035] Preferably, in any of the above solutions, step 21 includes using VGGNet as the backbone network to extract image features, generating joint heatmaps, and connecting the joints into a complete skeleton through graph optimization.
[0036] Preferably, in any of the above solutions, step 22 includes the following sub-steps:
[0037] Step 221: Extract kinematic features;
[0038] Step 222: Calculate joint angles;
[0039] Step 223: Complete the analysis of the motion trajectory by comparing the differences in joint angles between the affected side and the healthy side and calculating the variances of each axis of joint acceleration.
[0040] Preferably, in any of the above solutions, step 222 includes calculating the joint angle θ using the vector dot product formula through the coordinates of three points of the shoulder, elbow, and wrist.
[0041]
[0042] where represents the vector formed by the coordinates of the shoulder and elbow points, represents the vector formed by the coordinates of the elbow and wrist points, and represent the distances of the corresponding coordinates respectively.
[0043] Preferably, in any of the above solutions, step 23 includes first normalizing the RGB features and the bone pose features, then processing the RGB video with 3D-CNN and processing the bone time series data with LSTM respectively, and finally splicing the high-level features to complete modality fusion.
[0044] Preferably, in any of the above solutions, the calculation formula for the difference in joint angles between the affected side and the healthy side is
[0045] Δθ(t) = θ 患侧 (t) - θ 健侧 (t)
[0046] where θ 患侧 (t) and θ 健侧 (t) are the joint angle values at the same time point t.
[0047] Preferably, in any of the above solutions, the calculation formula for the variance of each axis of joint acceleration is
[0048]
[0049] where represents the sample variance, N represents the sample size, and a x (t) represents the data of the x-axis acceleration, and a y (t) represents the data of the y-axis acceleration, and a z (t) represents the data of the z-axis acceleration, and t represents the sample serial number. represents the sample mean, that is, the average value of all data points.
[0050] Preferably, in any of the above solutions, the step 24 includes the following sub-steps:
[0051] Step 241: Complete the Transformer time series modeling, and use the self-attention mechanism to calculate the correlation of features at different time steps;
[0052] Step 242: Use positional encoding to make up for the time series perception defect of the Transformer;
[0053] Step 243: Input the fused multi-modal feature sequence into the Transformer based on the time series classification model to complete the training of the rehabilitation evaluation expert model.
[0054] Preferably, in any of the above solutions, the step 3 includes the following sub-steps:
[0055] Step 31: Generate an adapted training plan for the rehabilitation exercise recommendation agent based on the reinforcement learning PPO algorithm and the knowledge graph network;
[0056] Step 32: Train the rehabilitation recommendation expert model;
[0057] Step 33: Combine the patient's historical data, evaluation results, and medical knowledge base to complete personalized recommendations for the patient;
[0058] Step 34: Update the recommendation strategy in real time according to the rehabilitation progress to complete dynamic adjustment.
[0059] Preferably, in any of the above solutions, the step 31 includes completing the policy optimization goal based on the reinforcement learning PPO algorithm, while pursuing the maximum expected return, restricting the amplitude of policy update.
[0060] Preferably, in any of the above solutions, the step 31 further includes using the knowledge graph network to model medical knowledge, and retrieving the knowledge graph in real time according to the patient's status, restricting the action space, and generating an adapted training plan for the rehabilitation exercise recommendation agent.
[0061] Preferably, in any of the above solutions, the step 32 includes training the rehabilitation recommendation expert model by inputting the multi-modal state vector based on the training plan.
[0062] Preferably, in any of the above solutions, step 33 includes querying the knowledge graph network to retrieve recommended actions related to the patient's diagnosis and excluding the patient's contraindicated actions, and outputting personalized recommendations for the patient based on the policy optimization network of the PPO algorithm.
[0063] Preferably, in any of the above solutions, step 34 includes fine-tuning the policy network according to the newly collected patient training effect data every week, preferentially reusing the trajectory data with significant or abnormal effects, and monitoring the rehabilitation progress throughout the process.
[0064] Preferably, in any of the above solutions, step 4 includes the following sub-steps:
[0065] Step 41: The patient uploads a recorded video, and the skeletal time-series data is obtained through the patient video comparison expert model, and then the result is transmitted to the rehabilitation evaluation expert model;
[0066] Step 42: The rehabilitation evaluation expert model evaluates the patient's motor function according to the obtained result and the Fugl-Meyer assessment scale;
[0067] Step 43: Transmit the motor function score result to the rehabilitation exercise recommendation expert model to generate a rehabilitation recommendation plan for the patient;
[0068] Step 44: The patient performs data backhaul;
[0069] Step 45: Repeat steps 41 to 44 to complete the collaborative connection between the three expert models and generate a safe and progressive rehabilitation plan.
[0070] The second object of the present invention is to provide a multi-agent-based automatic stroke assessment and rehabilitation management system, including a mobile device and a server, characterized in that it further includes
[0071] Multi-agent
[0072] The mobile device is used to obtain videos and pictures during the execution of rehabilitation action tasks and establish a connection with the server;
[0073] The server is responsible for communicating with the multi-agent, uploading the patient data uploaded from the mobile device end to the multi-agent
[0074] The multi-agent model includes a first agent, a second agent and a third agent
[0075] The first agent is trained as a patient video comparison expert model;
[0076] The second agent is trained as a rehabilitation evaluation expert model;
[0077] The third agent is trained as a rehabilitation exercise recommendation expert model;
[0078] The system uses the method described in the first objective for multi-agent based automatic stroke assessment and rehabilitation management.
[0079] Preferably, the network architecture used by the patient video comparison expert model is 3D-ResNet50 + LSTM time series modeling.
[0080] In any of the above solutions, preferably, the training method of the patient video comparison expert model includes the following sub-steps:
[0081] Step A1: Collect the patient-recorded video and the standard action video, and perform preprocessing operations of frame alignment and normalization on the data;
[0082] Step A2: Use 3D-ResNet50 for feature extraction to extract the spatio-temporal joint features of the video;
[0083] Step A3: Pass the extracted features through the spatio-temporal attention module, respectively through the temporal attention mechanism and the spatial attention mechanism, analyze the video and make predictions on the data;
[0084] Step A4: Via the contrast learning head, compare the processed video with the standard rehabilitation video;
[0085] Step A5: Use the DTW distance in the time series features to calculate the difference between the patient's actions and the standard template.
[0086] In any of the above solutions, preferably, the network architecture used by the rehabilitation assessment expert model is a multi-modal Transformer.
[0087] In any of the above solutions, preferably, the training method of the rehabilitation assessment expert model includes the following sub-steps:
[0088] Step B1: Align the video features extracted from the patient video across modalities, and at the same time align the clinical assessment data Fugl-Meyer score table and the knowledge of 3 chief rehabilitation physicians across modalities;
[0089] Step B2: Pass the features and data after cross-modal alignment through the hierarchical assessment head, perform coarse-grained classification and fine-grained scoring respectively, and perform weighted aggregation scoring on them, and comprehensively make a decision to output the assessment level;
[0090] Step B3: Conduct a preliminary assessment of the patient, and then further predict the patient's motor function assessment according to the Fugl-Meyer score table.
[0091] In any of the above schemes, it is preferred that the cross-modal alignment strategy mainly shortens the semantic distance between the patient video features and the expert knowledge embedding vector through contrast loss.
[0092] In any of the above schemes, it is preferred that the model architecture used by the rehabilitation exercise recommendation expert model is a reinforcement learning PPO algorithm combined with a knowledge graph.
[0093] In any of the above solutions, preferably, the training method of the rehabilitation exercise recommendation expert model comprises the following sub-steps:
[0094] Step C1: Obtain the patient's current status according to the rehabilitation assessment expert;
[0095] Step C2: The current state is constructed through the knowledge graph (and the reinforcement learning PPO algorithm to obtain action recommendations;
[0096] Step C3: Reward calculation is performed using the designed reward function;
[0097] Step C4: The motor relearning plan of rehabilitation training is changed according to the changes in the patient's motor function assessment through the strategy optimization network.
[0098] The present invention proposes a multi-agent-based automatic stroke assessment and rehabilitation management method and system, which integrates stroke assessment and rehabilitation and applies multi-agent to stroke assessment and rehabilitation to improve real-time performance, safety and user-friendliness.
[0099] OpenPose is an open source real-time multi-person pose estimation library.
[0100] VGGNet is a deep learning model proposed by the Visual Geometry Group of the University of Oxford.
[0101] Transformer is a deep learning model architecture that has revolutionized the field of natural language processing (NLP).
[0102] LSTM is a special type of recurrent neural network.
[0103] The Fugl-Meyer test is a standardized assessment tool widely used in the field of clinical rehabilitation. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] Figure 1 It is a flow chart of a preferred embodiment of the multi-agent based automatic stroke assessment and rehabilitation management method according to the present invention.
[0105] Figure 2 It is a schematic diagram of the composition of a preferred embodiment of the multi-agent-based automatic stroke assessment and rehabilitation management system according to the present invention.
[0106] Figure 3 Schematic diagram of the agent composition of a preferred embodiment of the multi-agent based stroke automatic evaluation and rehabilitation management system according to the present invention.
[0107] Figure 4 Schematic diagram of the technical implementation of a preferred embodiment of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention.
[0108] Figure 5 Schematic diagram of the multi-agent training of a preferred embodiment of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention.
[0109] Figure 6 Schematic diagram of the process of the patient video comparison expert training of a preferred embodiment of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention.
[0110] Figure 7 Schematic diagram of the process of the rehabilitation evaluation expert training of a preferred embodiment of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention.
[0111] Figure 8 Schematic diagram of the process of the rehabilitation exercise recommendation expert training of a preferred embodiment of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention.
[0112] Figure 9 Schematic diagram of an embodiment of the rehabilitation process with a good evaluation level of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention.
[0113] Figure 10 Schematic diagram of an embodiment of the rehabilitation process with a medium evaluation level of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention.
[0114] Figure 11 Schematic diagram of an embodiment of the rehabilitation process with a poor evaluation level of the multi-agent based stroke automatic evaluation and rehabilitation management method according to the present invention. Detailed implementation manners
[0115] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0116] Embodiment 1
[0117] As Figure 1 shown, a multi-agent based stroke automatic evaluation and rehabilitation management method performs step 1000 to obtain patient video data and perform preprocessing, including the following sub-steps:
[0118] Execute step 1010 to collect video by recording the rehabilitation movement video of the stroke patient through the camera of the mobile device.
[0119] Execute step 1020 to unify the format and adjust the resolution of the collected data using H.264 encoding to complete data standardization.
[0120] Execute step 1030 to preprocess the standardized data, and the preprocessing includes denoising, clipping, and temporal alignment.
[0121] Execute step 1100 to train the first agent as a patient video comparison expert model, and step 1 includes the following sub-steps:
[0122] Execute step 1110 to extract the spatio-temporal dynamic features of the patient video using 3D-CNN, including splitting the video segment into a continuous frame sequence, each segment containing T frames and H×W resolution, and the convolutional kernel sliding in three dimensions of space (H×W) and time (H×W) to extract spatial features and temporal dynamics simultaneously.
[0123] Execute step 1120 to enhance the weight of key action frames through the attention mechanism, including identifying the frames crucial for rehabilitation assessment through temporal attention and focusing on the significantly moving regions of joint movements through spatial attention, and enhancing the weight of key action frames when significantly moving regions appear in the patient video.
[0124] Execute step 1130 to compare the feature similarity of different patient videos through cosine similarity to complete modality fusion, including normalizing the feature vectors extracted by 3D-CNN, performing similarity measurement, mapping the video features and physiological indicators to the same embedding space, and completing modality fusion by assigning different weights to different features.
[0125] Execute step 1140 to complete the training of the patient video comparison expert model based on the Triplet Loss function. The Triplet Loss function consists of three parts: anchor A, positive sample P, and negative sample N. Among them, anchor A is a certain reference sample, positive sample P is a sample belonging to the same category as anchor A, and negative sample N is a sample of a different category from the anchor. The Triplet Loss function L is defined as
[0126] L = max(d(A,P) - d(A,N) + margin, 0)
[0127] where d(A,P) is the distance between the anchor and the positive sample, d(A,N) is the distance between the anchor and the negative sample, and margin is a preset margin.
[0128] Execute step 1200 to train the second agent as a rehabilitation assessment expert model, including the following sub-steps:
[0129] Execute step 1210 to extract the temporal data of skeletal joint points through Openpose, including using VGGNet as the backbone network to extract image features, generating joint point heatmaps, and connecting the joint points into complete skeletons through graph optimization.
[0130] Execute step 1220 to calculate biomechanical indices through motion features and complete the analysis of the motion trajectory, including the following sub-steps:
[0131] Execute step 1221 to extract kinematic features.
[0132] Execute step 1222 to calculate joint angles, including calculating the joint angle θ through the coordinates of three points of the shoulder, elbow, and wrist using the vector dot product formula.
[0133]
[0134] Among them, represents the vector composed of the coordinates of the shoulder and elbow points, represents the vector composed of the coordinates of the elbow and wrist points, and represent the distances of the corresponding coordinates respectively.
[0135] Execute step 1223 to complete the analysis of the motion trajectory by comparing the differences in joint angles between the affected side and the healthy side and calculating the variances of each axis of joint acceleration. The calculation formula for the difference in joint angles between the affected side and the healthy side is
[0136] Δθ(t) = θ 患侧 (t) - θ 健侧 (t)
[0137] Among them, θ 患侧 (t) and θ 健侧 (t) are the joint angle values at the same time point t.
[0138] Calculation of the variance of joint acceleration:
[0139] If the acceleration is a three-dimensional vector (such as a x , a y , a z ), the variances of each axis can be calculated respectively:
[0140]
[0141] Among them, represents the sample variance, N represents the number of samples, a x (t) represents the data of the x-axis acceleration, a y (t) represents the data of the y-axis acceleration, a z(t) represents the data of the z-axis acceleration, and t represents the sample serial number. represents the sample mean, that is, the average value of all data points. The denominator is the sample size minus one. Using N - 1 is for unbiased estimation to correct the small sample error and make the calculation result closer to the population variance.
[0142] Execute step 1230, adopt the feature splicing technology to fuse the RGB features and skeletal pose features of the video, and complete the modality fusion, including first normalizing the RGB features and skeletal pose features, then processing the RGB video with 3D - CNN respectively, processing the skeletal temporal data with LSTM, and finally splicing the high - level features to complete the modality fusion.
[0143] Execute step 1240, complete the training of the rehabilitation assessment expert model based on the temporal classification model Transformer, including the following sub - steps:
[0144] Execute step 1240, complete the Transformer temporal modeling, and calculate the correlation of features at different time steps by using the self - attention mechanism;
[0145] Execute step 1240, use positional encoding to make up for the temporal perception defect of Transformer;
[0146] Execute step 1243, input the fused multi - modality feature sequence into the rehabilitation assessment expert model trained based on the temporal classification model Transformer.
[0147] Execute step 1300, train the third agent as the rehabilitation exercise recommendation expert model, including the following sub - steps:
[0148] Execute step 1310, generate an adapted training plan for the rehabilitation exercise recommendation agent based on the reinforcement learning PPO algorithm and the knowledge graph network, including completing the policy optimization goal based on the reinforcement learning PPO algorithm, while pursuing the maximization of the expected return, restricting the amplitude of policy update; using the knowledge graph network to model medical knowledge, and retrieving the knowledge graph in real - time according to the patient's state to constrain the action space and generate an adapted training plan for the rehabilitation exercise recommendation agent.
[0149] Execute step 1320, train the rehabilitation recommendation expert model, including inputting the multi - modality state vector based on the training plan to complete the training of the rehabilitation recommendation expert model.
[0150] Execute step 1330, complete personalized recommendation for the patient by combining the patient's historical data, evaluation results and medical knowledge base, including querying the knowledge graph network to retrieve the recommended actions related to the patient's diagnosis and excluding the patient's taboo actions, and outputting the patient's personalized recommendation based on the policy optimization network of the PPO algorithm.
[0151] Execute step 1340, and dynamically adjust the recommendation strategy in real time according to the rehabilitation progress, including fine-tuning the policy network weekly based on the newly collected patient training effect data, preferentially reusing the trajectory data with significant or abnormal effects, and monitoring the rehabilitation progress throughout the process.
[0152] Execute step 1400 to establish a collaborative connection between three expert models. The patient video comparison expert model completes action feature extraction and anomaly detection. The rehabilitation evaluation expert model completes the quantitative evaluation of the degree of motor function recovery. The rehabilitation exercise recommendation expert model completes the generation of a safe and progressive rehabilitation plan, including the following sub-steps:
[0153] Execute step 1410. The patient uploads a recorded video, and the skeletal time series data is obtained through the patient video comparison expert model, and then the result is transmitted to the rehabilitation evaluation expert model;
[0154] Execute step 1420. The rehabilitation evaluation expert model evaluates the patient's motor function based on the obtained result and the Fugl-Meyer assessment scale;
[0155] Execute step 1430, and transmit the motor function score result to the rehabilitation exercise recommendation expert model to generate a patient rehabilitation recommendation plan;
[0156] Execute step 1440, and the patient performs data backhaul;
[0157] Execute step 1450, and repeat steps 1410 to 1440 to complete the collaborative connection between the three expert models and generate a safe and progressive rehabilitation plan
[0158] Execute step 1500 to output the rehabilitation plan.
[0159] Embodiment 2
[0160] A multi-agent-based automatic stroke assessment and stroke rehabilitation self-management system, including a mobile device, a server, and multiple agents.
[0161] The mobile device is used to obtain videos and pictures during the execution of rehabilitation action tasks and establish a connection with the server;
[0162] The server is responsible for communicating with the multiple agents and uploading the patient data uploaded from the mobile device side to the multiple agents.
[0163] The multi-agent model includes a first agent, a second agent, and a third agent,
[0164] The first agent is trained as a patient video comparison expert model, the network architecture used by the patient video comparison expert model is 3D-ResNet50+LSTM time series modeling, and the training method of the patient video comparison expert model includes the following sub-steps:
[0165] Step A1: Collect patient recorded videos and standard action videos, and perform preprocessing operations of frame alignment and normalization on the data;
[0166] Step A2: Use 3D-ResNet50 to perform feature extraction and extract the spatiotemporal joint features of the video;
[0167] Step A3: The extracted features are subjected to the temporal attention mechanism and the spatial attention mechanism through the spatiotemporal attention module to analyze the video and predict the data;
[0168] Step A4: Compare the processed video with the standard rehabilitation video via a contrast learning head;
[0169] Step A5: Calculate the difference between the patient's action and the standard template using the DTW distance in the time series features.
[0170] The second agent is trained as a rehabilitation assessment expert model, the network architecture used by the rehabilitation assessment expert model is a multimodal Transformer, and the training method of the rehabilitation assessment expert model includes the following sub-steps:
[0171] Step B1: Cross-modal alignment of the video features extracted from the patient video, and cross-modal alignment of the clinical assessment data Fugl-Meyer score sheet and the knowledge of three chief physicians of the rehabilitation department. The cross-modal alignment strategy mainly uses contrast loss to narrow the semantic distance between the patient video features and the expert knowledge embedding vector;
[0172] Step B2: The cross-modal aligned features and data are passed through a hierarchical evaluation head for coarse-grained classification and fine-grained scoring, and weighted summary scores are performed to output the evaluation level through comprehensive decision-making;
[0173] Step B3: After conducting a preliminary assessment of the patient, further predict the patient's motor function assessment based on the Fugl-Meyer scoring table.
[0174] The third agent is trained as a rehabilitation exercise recommendation expert model. The model architecture used by the rehabilitation exercise recommendation expert model is a reinforcement learning PPO algorithm combined with a knowledge graph. The training method of the rehabilitation exercise recommendation expert model includes the following sub-steps:
[0175] Step C1: Obtain the patient's current status according to the rehabilitation assessment expert;
[0176] Step C2: Obtain action recommendations for the currently constructed knowledge graph (and the reinforcement learning PPO algorithm);
[0177] Step C3: Calculate the reward through the designed reward function;
[0178] Step C4: Through the policy optimization network, the motor relearning plan for rehabilitation training is changed according to the changes in the patient's motor function assessment.
[0179] Embodiment III
[0180] The present invention proposes a multi-agent-based automatic stroke assessment and stroke rehabilitation self-management system. The invention introduces the currently popular multi-agent module, making the invention intelligent and enabling patients to perform rehabilitation training at home. The intelligence of this system is reflected in that the multi-agent is trained as a patient video comparison expert, a rehabilitation assessment expert, and a rehabilitation exercise recommendation expert through standard stroke rehabilitation videos, doctors' stroke rehabilitation knowledge, and stroke rehabilitation knowledge obtained from the Internet. The system can intelligently complete the whole process of the patient using a mobile device to upload videos, the mobile device connecting to the Internet to upload the patient's videos to the server, then from the server to the multi-agent, the multi-agent completing the assessment of the patient's motor function, and then recommending appropriate rehabilitation training according to the assessment level.
[0181] As Figure 2 shown, this system includes three components: a mobile device, a server, and a multi-agent. The mobile device is used by patients to upload videos and pictures of their own rehabilitation actions (such as pendulum movement, forward and backward shoulder rotation, wall climbing exercise in shoulder joint activities for shoulder and neck rehabilitation, neck lateral stretch, chin retraction in neck relaxation, cat-camel pose, hip bridge, dead bug pose in core stability activities for waist rehabilitation, supine knee hug roll in lumbar traction, and wall sit, straight leg raise, elastic band resistance knee extension in quadriceps femoris training for knee rehabilitation, and single-leg standing in balance training, etc.), as well as receiving self-management exercise or progress report suggestions and establishing a connection with the server side. The server side is responsible for communicating with the multi-agent and uploading the patient data uploaded from the mobile device side to the multi-agent side. Finally, there is the multi-agent side, as Figure 3As shown in the figure, the multi-agent terminal consists of multiple agents with different knowledge bases, decision-making capabilities, and action capabilities. In this system, the multi-agent terminal is composed of three types of core agents, which respectively correspond to the three major expert functions in the output, and complete the closed-loop from data input to decision output through cooperation. First is knowledge input, and the input content is standard stroke rehabilitation videos, doctors' stroke rehabilitation knowledge, and stroke rehabilitation knowledge obtained from the Internet. Then, through the multi-agent, as the "brain" of the system, it is responsible for knowledge fusion and task distribution. Specifically, knowledge fusion combines doctors' clinical experience with network knowledge to construct a dynamic knowledge graph. Task distribution triggers the corresponding expert agents (comparison, evaluation, recommendation) according to the input data (patient video). Finally, the output function is completed through three expert agent layers. The multi-agent terminal first needs to be trained. By using the obtained standard stroke rehabilitation videos, doctors' stroke rehabilitation knowledge, and stroke rehabilitation knowledge obtained from the Internet, the multi-agent is trained to be an expert for comparing patient videos, a rehabilitation evaluation expert, and a rehabilitation exercise recommendation expert. Then, according to the patient data on the server side, it conducts motor function evaluation and rehabilitation training recommendation for the patient.
[0182] As Figure 4 shown, an overview of the implementation process of the technical solution:
[0183] Step1: Establish an initial data transmission communication between the mobile device and the server as the data transmission bridge for the entire system.
[0184] Step2: As Figure 5 shown, train the multi-agent.
[0185] By using the obtained standard stroke rehabilitation videos, doctors' stroke rehabilitation knowledge, and stroke rehabilitation knowledge obtained from the Internet, the multi-agent is trained to be an expert for comparing patient videos, a rehabilitation evaluation expert, and a rehabilitation exercise recommendation expert.
[0186] Step3: Data preprocessing.
[0187] For the videos uploaded by users using mobile devices and the collected stroke anti-rehabilitation videos, they have different durations according to the performance levels of the patients. At the same time, preprocessing work is carried out on all the collected data.
[0188] Step4: Feature extraction.
[0189] To obtain high-level spatio-temporal features from the input, the system first uses a ResNet3D backbone to extract feature maps (the ResNet3D model is pre-trained on an action recognition dataset). Subsequently, a Transformer encoder is used to capture the latent semantics and global dependencies of the spatio-temporal features output by the backbone. At the same time, the output features of the backbone undergo the following transformation before being input into the Transformer encoder. The dimension of the feature map output by the backbone is C×T×H×W, representing T frames of size C×H×W, where C, H, W, and T represent the channel, height, width, and time dimensions respectively. The spatial dimension is squeezed out through two-dimensional spatial average pooling, changing the feature dimension to C×T, and then it is permuted to T×C. Then each feature vector in the time dimension is projected into an h-dimensional vector. Therefore, the dimension of the output features is T×h, representing T tokens of features of size h. After position encoding the token sequence, it is input into the Transformer encoder.
[0190] Step5: Modal Fusion
[0191] There is asymmetry in the modalities: RGB data is the main modality, while optical flow is a modality derived from RGB data. This step aims to capture this asymmetry and attempts to extract information mainly from RGB features. The RGB and optical flow data are fused in two steps. First, the features output by the two modalities of the Transformer encoder are concatenated and passed to the MLPMixer. The MLPMixer first mixes the concatenated features across the modal dimensions (modal mixing) to generate intermediate features, and further mixes the intermediate features (feature mixing). Then an adapter layer is added to enable parameter-efficient tuning of the Transformer block. Finally, a fully connected layer is used to classify the output of the modal fusion module. This system integrates stroke assessment and rehabilitation, and applies multi-agent to stroke assessment and rehabilitation to improve real-time performance, safety, and user-friendliness.
[0192] Example Four
[0193] The steps of the present invention are as follows:
[0194] Step 1: Obtain patient video data and perform preprocessing;
[0195] In this step, first, a rehabilitation action video of a stroke patient is recorded through the camera of a mobile device for video acquisition, and then the H.264 encoding is used to unify the format and adjust the resolution of the acquired data to complete data standardization. Finally, the preprocessing work is carried out, and the steps are as follows:
[0196] 1. Denoising: Use Gaussian filtering technology to eliminate motion blur.
[0197] 2. Cropping: Based on the human detection box, make the acquired data focus on the patient's action area
[0198] 3. Temporal alignment: Use Dynamic Time Warping (DTW) to synchronize the temporal information of multiple video segments.
[0199] Step 2: Train the first agent as an expert model for patient video comparison;
[0200] In this step, feature extraction is first performed. The main features extracted are the spatio-temporal dynamic features of the patient video using 3D-CNN, and the weights of key action frames are enhanced through an attention mechanism (such as SENet). Then, the feature similarity of different patient videos is compared through cosine similarity to complete modality fusion. Finally, the training of the expert model for patient video comparison is completed based on the Triplet Loss contrastive loss function.
[0201] Step 21: Use 3D-CNN to extract the spatio-temporal dynamic features of the patient video;
[0202] The video segment is sliced into a continuous frame sequence. Each segment contains T frames (time dimension) and an H×W resolution (space dimension). The convolutional kernel slides in three dimensions: space (H×W) and time (T), extracting spatial features (such as joint shape) and temporal dynamics (such as action coherence) simultaneously.
[0203] Step 22: Enhance the weights of key action frames through an attention mechanism;
[0204] Identify the frames crucial for rehabilitation assessment (such as the starting / peak phase of the action) through temporal attention and focus on the significantly moving regions of joint motion (such as the affected arm) through spatial attention. When a significantly moving region appears in the patient video, the weights of the key action frames are enhanced.
[0205] Step 23: Compare the feature similarity of different patient videos through cosine similarity to complete modality fusion;
[0206] First, normalize the feature vectors extracted by 3D-CNN, and then perform similarity measurement. Then, map the video features and physiological indicators to the same embedding space, and complete modality fusion by assigning different weights to different features (e.g., fused feature = 0.7×video feature + 0.3×physiological indicator feature)
[0207] Step 24: Complete the training of the expert model for patient video comparison based on the Triplet Loss contrastive loss function.
[0208] The Triplet Loss function is applicable to tasks that require measuring the similarity between samples. It aims to train the model to learn the similarity relationship between samples, making samples of the same class closer in the feature space and samples of different classes farther apart.
[0209] The Triplet Loss function usually consists of three parts: Anchor (A): A reference sample (such as a face picture). Positive sample (P): A sample belonging to the same category as the anchor (such as a face picture of the same person). Negative sample (N): A sample belonging to a different category from the anchor (such as face pictures of different people).
[0210] The Triplet Loss function is defined as: L = max(d(A, P) - d(A, N) + margin, 0)
[0211] Where: d(A, P): The distance between the anchor and the positive sample (usually the cosine distance). d(A, N): The distance between the anchor and the negative sample. margin: A preset margin used to control the minimum difference between positive and negative samples.
[0212] The training of the patient video comparison expert model is completed by continuously comparing the rehabilitation action video recorded by the patient with the standard rehabilitation action video through the Triplet Loss function.
[0213] Step 3: Train the second agent as a rehabilitation evaluation expert model;
[0214] In this step, feature extraction is first performed. The main work completed is to extract the temporal data of skeletal joint points through Openpose to complete pose estimation and calculate biomechanical indicators such as joint angles, speeds, and accelerations through the motion feature map to complete the analysis of the motion trajectory. Then, the feature splicing technology is used to fuse the RGB features and skeletal pose features of the video to complete modality fusion. Finally, the training of the rehabilitation evaluation expert model is completed using the temporal classification model Transformer.
[0215] Step 31: Extract the temporal data of skeletal joint points through Openpose;
[0216] Use VGGNet (consisting of 13 convolutional layers and 3 fully connected layers) as the backbone network to extract image features, generate joint point heatmaps, and connect the joint points into a complete skeleton through graph optimization.
[0217] Step 32: Calculate biomechanical indicators through the motion feature map and complete the analysis of the motion trajectory;
[0218] First, extract kinematic features, and then perform joint angle calculation. Through the coordinates of the shoulder, elbow, and wrist points, use the vector dot product formula to calculate the angle:
[0219]
[0220] Where, represents the vector composed of the coordinates of the shoulder and elbow points, represents the vector composed of the coordinates of the elbow and wrist points, and are vectors, representing the distances of the corresponding coordinates respectively. Then, the analysis of the motion trajectory is completed by comparing the joint angle differences between the affected side and the healthy side and calculating the variance of the joint acceleration (the larger the variance, the more uncoordinated the movement).
[0221] Step 33: Adopt the feature splicing technology to fuse the RGB features and the skeletal pose features of the video to complete the modality fusion;
[0222] First, perform feature normalization on the RGB features and the skeletal pose features, then process the RGB video with 3D-CNN and process the skeletal time series data with LSTM respectively, and finally splice the high-level features to complete the modality fusion.
[0223] Step 34: Use the Transformer based on the time series classification model to complete the training of the rehabilitation evaluation expert model.
[0224] First, complete the Transformer time series modeling, calculate the correlation of the features at different time steps by using the self-attention mechanism, and then use the position encoding to make up for the time series perception defect of the Transformer. Then, input the fused multi-modal feature sequence into the Transformer based on the time series classification model to complete the training of the rehabilitation evaluation expert model.
[0225] Step 4: Train the third agent as the rehabilitation exercise recommendation expert model;
[0226] In this step, first, generate an adapted training plan for the rehabilitation exercise recommendation agent based on the PPO algorithm of reinforcement learning and the knowledge graph network to complete the training of the rehabilitation recommendation expert model. Then, combine the patient's historical data, evaluation results, and medical knowledge base to complete personalized recommendations for the patient, and update the recommendation strategy in real time according to the rehabilitation progress to complete dynamic adjustment.
[0227] Step 41: Generate an adapted training plan for the rehabilitation exercise recommendation agent based on the PPO algorithm of reinforcement learning and the knowledge graph network;
[0228] Based on the PPO algorithm of reinforcement learning, complete the policy optimization goal, maximize the expected return while restricting the amplitude of policy update to ensure the stability of rehabilitation training. At the same time, use the knowledge graph network to model medical knowledge and retrieve the knowledge graph in real time according to the patient's state to constrain the action space. Generate an adapted training plan for the rehabilitation exercise recommendation agent.
[0229] Step 42: Train the rehabilitation recommendation expert model;
[0230] Based on the training plan generated in Step 41, input the multi-modal state vector (patient state + knowledge graph constraint) to complete the training of the rehabilitation recommendation expert model.
[0231] Step 43: Complete personalized recommendations for the patient by combining the patient's historical data, evaluation results, and medical knowledge base;
[0232] First, query the knowledge graph network to retrieve recommended actions related to the patient's diagnosis and exclude actions contraindicated for the patient. Then, optimize the network based on the PPO algorithm's strategy to output personalized recommendations for the patient.
[0233] Step 44: Update the recommendation strategy in real-time according to the rehabilitation progress to complete dynamic adjustment.
[0234] Fine-tune the policy network weekly based on newly collected patient training effect data, and preferentially reuse trajectory data with significant or abnormal effects (such as training injuries). Monitor the rehabilitation progress throughout the process (for example, if the joint range of motion decreases by more than 10%, trigger a review of the plan. Or adjust the training intensity if there is no progress for three consecutive days) to complete the dynamic adjustment of the rehabilitation strategy.
[0235] Step 5: Establish collaborative connections between the three expert models. The patient video comparison expert extracts action features and detects abnormalities, the rehabilitation evaluation expert completes a quantitative evaluation of the degree of motor function recovery, and the rehabilitation exercise recommendation expert completes the generation of a safe and progressive rehabilitation plan. First, the patient uploads a recorded video, and the patient video comparison expert obtains bone sequence data. Then, the result is passed to the rehabilitation evaluation expert, who evaluates the patient's motor function based on the obtained result and the Fugl-Meyer assessment scale. Then, the motor function score result is passed to the rehabilitation exercise recommendation expert to generate a rehabilitation recommendation plan for the patient. Then, the patient performs data feedback, and the above operations are repeated to complete the collaborative connection between the three expert models. This collaborative mechanism realizes a complete closed-loop from "action capture → function evaluation → plan generation → effect tracking", enabling the rehabilitation treatment to have the ability of self-iteration and optimization, and ultimately improving the rehabilitation efficiency and safety of the patient.
[0236] Step 6: Output a rehabilitation exercise recommendation plan.
[0237] In this step, a personalized training plan (action type, intensity, frequency) is output based on the rehabilitation exercise recommendation expert model.
[0238] The purpose of the present invention is to reduce the frequency of hospital visits for stroke patients with limited mobility and better protect the privacy of patients. The principle of stroke detection in the present invention is to distinguish different human characteristics by using the action skeletons of stroke patients and those of healthy people to determine whether a stroke has occurred. The principle of stroke rehabilitation in the present invention is that when the patient performs rehabilitation actions according to the rehabilitation videos recommended by multiple agents, the recorded video is compared with the standard rehabilitation video to determine whether the patient has fully recovered.
[0239] This system mainly consists of the following parts: mobile devices, servers, and multi-agents. These parts are closely linked in the system. Mobile devices are used by patients to upload videos and pictures when performing rehabilitation action tasks, receive self-management practice or progress report suggestions, and establish connections with the server side. The server side is responsible for communicating with the multi-agents and uploading the patient data uploaded from the mobile device side to the multi-agent side.
[0240] The multi-agent side first needs to be trained. By using the obtained standard stroke rehabilitation videos, doctors' stroke rehabilitation knowledge, and stroke rehabilitation knowledge obtained from the Internet, the multi-agent is trained into an expert for patient video comparison, rehabilitation assessment, and rehabilitation exercise recommendation. Then, based on the patient data from the server side, it conducts kinetic energy assessment and rehabilitation training recommendation for the patient.
[0241] The training methods for each expert are different. The training of the patient video comparison expert requires the use of computer vision technologies such as action recognition, pose estimation, and temporal analysis, and uses standard video data to evaluate whether the patient's actions are standard or detect motor function disorders after a stroke. The training of the rehabilitation assessment expert requires combining clinical data and the patient's historical records and using classification or regression models. The training of the rehabilitation exercise recommendation expert requires the use of reinforcement learning and recommendation system technologies, combined with a medical knowledge graph, and combines the patient's current status and rehabilitation goals to generate personalized exercise plans.
[0242] Before training each expert, multi-modal data preparation is required, including video data such as videos or pictures of patients performing rehabilitation actions recorded or taken using mobile devices and standard stroke rehabilitation videos. Clinical assessment data, the Fugl-Meyer assessment scale (a method for assessing motor function specifically designed for stroke patients). At the same time, a data augmentation strategy is adopted, and Unity3D is used to generate pathological gait simulation videos (covering types such as hemiplegia / spasm). Finally, through feature engineering, mainly focusing on the joint angle covariance matrix in kinematic features, the ground reaction force asymmetry in kinetic features, and using the DTW distance in temporal features to calculate the difference between the patient's actions and the standard template.
[0243] Each expert is trained independently. The detailed training processes of the patient video comparison expert, the rehabilitation assessment expert, and the rehabilitation exercise recommendation expert are as follows:
[0244] Patient Video Comparison Expert: The network architecture used is 3D-ResNet50 + LSTM for temporal modeling. First, collect the videos recorded by the patient and the standard action videos, and perform preprocessing operations such as frame alignment and normalization on the data. Then use 3D-ResNet50 for feature extraction, mainly extracting the spatio-temporal joint features of the video. Then, through the spatio-temporal attention module (the upper part is the spatial attention mechanism, and the lower part is the temporal attention mechanism), the extracted features are respectively processed through the temporal attention mechanism (identifying the key frames in the action and suppressing the interference of redundant frames) and the spatial attention mechanism (focusing on the key body regions and enhancing the local features through Senet channel weighting), further analyzing the video and making predictions on the data. Then, via the contrast learning head, compare the processed video with the standard rehabilitation video, and finally use the DTW distance in the temporal features to calculate the difference between the patient's action and the standard template, as Figure 6 shown.
[0245] Rehabilitation Assessment Expert: The network architecture used is a multi-modal Transformer (combining video features, clinical assessment data Fugl-Meyer score table, and the knowledge of 3 chief rehabilitation physicians). First, align the video features extracted from the patient's video across modalities, and at the same time align the clinical assessment data Fugl-Meyer score table and the knowledge of 3 chief rehabilitation physicians across modalities. This alignment strategy mainly narrows the semantic distance between the patient's video features and the expert knowledge embedding vectors through the contrast loss. Then, pass the features and data that have undergone cross-modal alignment through the hierarchical assessment head, and perform coarse-grained classification (judging the action category) and fine-grained scoring (calculating specific indicators based on the database and rules) respectively, and perform weighted aggregation scoring on them, and comprehensively make a decision to output the assessment level. Through the preliminary assessment of the patient, further predict the patient's motor function assessment according to the Fugl-Meyer score table, as Figure 7 shown.
[0246] Rehabilitation Exercise Recommendation Expert: The model architecture used is the reinforcement learning PPO algorithm + knowledge graph. First, obtain the current state of the patient (such as Fugl-Meyer score, real-time psychological indicators, and action video features) according to the rehabilitation assessment expert. Then pass the current state through the constructed knowledge graph (by integrating clinical guidelines, expert experience, and patient portraits) and the reinforcement learning PPO algorithm to obtain action recommendations, and then perform reward calculation through the designed reward function (mainly considering the patient's completion degree and whether the action conforms to medical rules), and finally pass through the policy optimization network (updating the policy uses the Clip gradient clipping of PPO to prevent overfitting) to make the motor relearning plan of the rehabilitation training change with the change of the patient's motor function assessment, as Figure 8 shown.
[0247] After each expert has completed training, they need to work together
[0248] In this step, the parameters of the video comparison model are first frozen using a joint training strategy, the evaluation and recommendation modules are jointly optimized, and the gradient synchronization of each expert model is achieved through the communication protocol of the message queue. Then, a cross-modal Transformer evaluation model and a recommendation model share a latent feature space, and a multi-task joint loss of comparison loss + evaluation loss + recommendation loss is designed to complete the cooperation mechanism.
[0249] For example, the result of video comparison is used as the input for rehabilitation evaluation, and the evaluation result affects the recommended rehabilitation content. In this system, it is specifically manifested as the integration of doctor's knowledge (authority) and network knowledge (timeliness) in the multi-agent center, forming a dynamically expandable rehabilitation knowledge base. The patient training video and the standard video are analyzed by the patient video comparison expert to generate a quantitative deviation report. The rehabilitation evaluation expert synthesizes the deviation data and the knowledge base to output an evaluation conclusion. The patient recommendation expert generates a personalized plan based on the evaluation result and monitors the patient's execution feedback (such as new video data), forming a closed loop of "execution → evaluation → optimization".
[0250] In addition, during the training process of each expert, there are privacy and security issues with a lot of patient data and medical data used. Therefore, using federated learning is a solution that allows model training without the data leaving the local area. In addition, differential privacy or homomorphic encryption is used to protect patient data.
[0251] As Figure 9 shown in the schematic diagram of the rehabilitation process for stroke patients with a good evaluation grade. First, the stroke patient uses a mobile device to record a video of the actions they perform, and then uploads the video to the server. The server evaluates the action video uploaded by the patient through trained multi-agents. After the multi-agents compare the patient's action skeleton with the standard action skeleton, the patient's motor function is evaluated and the evaluation grade is good. Then, based on the motor function evaluation grade, the multi-agents recommend more difficult rehabilitation exercises for the patient (because the patient's condition is not serious). The patient records a video through the mobile device according to the rehabilitation exercises recommended by the multi-agents, uploads it to the server, and is evaluated and recommended rehabilitation exercises by the trained multi-agents. After repeated rehabilitation exercises many times, until the stroke patient completes the rehabilitation.
[0252] As Figure 10The following is a schematic diagram of the rehabilitation process for stroke patients with a medium evaluation level. First, the stroke patient uses a mobile device to record a video of the actions they perform, and then uploads the video to the server. The server evaluates the action video uploaded by the patient through trained multi-agents. After comparing the patient's action skeleton with the standard action skeleton, the multi-agents conduct a motor function evaluation of the patient and the evaluation level is medium. Then, the multi-agents recommend rehabilitation exercises with moderate difficulty for the patient according to the motor function evaluation level (because the patient's condition is relatively serious). The patient records a video through the mobile device according to the recommended rehabilitation exercises by the multi-agents, uploads it to the server, and is evaluated and recommended rehabilitation exercises by the trained multi-agents. The patient needs to first recover to a good evaluation level by the multi-agents before they can go through multiple repeated rehabilitation exercises until the stroke patient completes the rehabilitation.
[0253] As Figure 11 The following is a schematic diagram of the rehabilitation process for stroke patients with a poor evaluation level. First, the stroke patient uses a mobile device to record a video of the actions they perform, and then uploads the video to the server. The server evaluates the action video uploaded by the patient through trained multi-agents. After comparing the patient's action skeleton with the standard action skeleton, the multi-agents conduct a motor function evaluation of the patient and the evaluation level is poor. Then, the multi-agents recommend rehabilitation exercises with a lower difficulty for the patient according to the motor function evaluation level (because the patient's condition is serious). The patient records a video through the mobile device according to the recommended rehabilitation exercises by the multi-agents, uploads it to the server, and is evaluated and recommended rehabilitation exercises by the trained multi-agents. The patient needs to first recover to a medium evaluation level by the multi-agents, then to a good evaluation level, and finally can go through multiple repeated rehabilitation exercises until the stroke patient completes the rehabilitation.
[0254] To better understand the present invention, the above has been described in detail in conjunction with specific embodiments of the present invention, but it is not a limitation of the present invention. Any simple modification made to the above embodiments based on the technical essence of the present invention still belongs to the scope of the technical solution of the present invention. Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
Claims
1. An automatic stroke assessment and rehabilitation management method based on multi-agent, including obtaining patient video data and performing preprocessing, characterized in that, It further includes the following steps: Step 1: Train the first agent as a patient video comparison expert model; Step 2: Train the second agent as a rehabilitation assessment expert model; Step 3: Train the third agent as a rehabilitation exercise recommendation expert model; Step 4: Establish a collaborative connection among the three expert models. The patient video comparison expert model completes action feature extraction and anomaly detection. The rehabilitation assessment expert model completes the quantitative assessment of the degree of motor function recovery. The rehabilitation exercise recommendation expert model completes the generation of a safe and progressive rehabilitation plan; Step 5: Output the rehabilitation plan.
2. The multi-agent-based automatic stroke assessment and rehabilitation management method according to claim 1, characterized in that The said Step 1 includes the following sub-steps: Step 11: Use 3D-CNN to extract the spatio-temporal dynamic features of the patient video; Step 12: Enhance the weight of key action frames through the attention mechanism; Step 13: Complete modal fusion by comparing the feature similarity of different patient videos through cosine similarity; Step 14: Complete the training of the patient video comparison expert model based on the Triplet Loss function.
3. The multi-agent-based automatic stroke assessment and rehabilitation management method according to claim 2, wherein, The said Step 11 includes splitting the video clip into a continuous frame sequence. Each clip contains T frames and has a resolution of H×W. The convolutional kernel slides in three dimensions: space (H×W) and time (H×W), and simultaneously extracts spatial features and temporal dynamics.
4. The multi-agent-based automatic stroke assessment and rehabilitation management method according to claim 3, wherein The said Step 12 includes identifying the frames crucial for rehabilitation assessment through temporal attention and focusing on the significant regions of joint movement through spatial attention. When significant regions of movement appear in the patient video, the weight of the key action frames is enhanced.
5. The multi-agent-based automatic stroke assessment and rehabilitation management method according to claim 4, wherein The said Step 13 includes normalizing the feature vectors extracted by 3D-CNN, performing similarity measurement, mapping the video features and physiological indicators to the same embedding space, and completing modal fusion by assigning different weights to different features.
6. The multi-agent based automatic stroke assessment and rehabilitation management method according to claim 5, wherein The said Triplet Loss function consists of three parts: anchor A, positive sample P, and negative sample N. Among them, anchor A is a certain reference sample, positive sample P is a sample belonging to the same category as anchor A, and negative sample N is a sample of a different category from the anchor.
7. The multi-agent-based automatic stroke assessment and rehabilitation management method according to claim 6, wherein The said Triplet Loss function L is defined as L = max(d(A,P) - d(A,N) + margin, 0) where d(A,P) is the distance between the anchor and the positive sample, d(A,N) is the distance between the anchor and the negative sample, and margin is a preset margin.
8. The method for automatic stroke assessment and rehabilitation management based on multi-agent as claimed in claim 7, wherein The said Step 2 includes the following sub-steps: Step 21: Extract the temporal data of skeletal joint points through Openpose; Step 22: Calculate biomechanical indicators through motion features to complete the analysis of the motion trajectory; Step 23: Adopt feature splicing technology to fuse the RGB features and skeletal pose features of the video to complete modal fusion; Step 24: Complete the training of the rehabilitation assessment expert model based on the temporal classification model Transformer.
9. The multi-agent based automatic stroke assessment and rehabilitation management method according to claim 8, characterized in that, The said Step 4 includes the following sub-steps: Step 41: The patient uploads a recorded video. After passing through the patient video comparison expert model, the skeletal temporal data is obtained, and then the result is passed to the rehabilitation assessment expert model; Step 42: The rehabilitation evaluation expert model evaluates the patient's motor function based on the obtained results and the Fugl-Meyer assessment scale; Step 43: Transmit the motor function scoring result to the rehabilitation exercise recommendation expert model to generate a rehabilitation recommendation plan for the patient; Step 44: The patient performs data backtransmission; Step 45: Repeat Steps 41 to 44 to complete the collaborative connection between the three expert models and generate a safe and progressive rehabilitation plan.
10. An automatic stroke evaluation and rehabilitation management system based on multi-agent, comprising a mobile device and a server, characterized in that, It further includes multiple agents, The mobile device is used to obtain videos and pictures during the execution of rehabilitation action tasks and establish a connection with the server; The server side is responsible for communicating with the multiple agents, and uploading the patient data uploaded from the mobile device side to the multiple agents, The multiple agent model includes a first agent, a second agent, and a third agent, The first agent is trained as an expert model for patient video comparison; The second agent is trained as a rehabilitation evaluation expert model; The third agent is trained as a rehabilitation exercise recommendation expert model; The system uses the method described in Claim 1 for automatic stroke assessment and rehabilitation management based on multiple agents.
Citation Information
Patent Citations
Cerebral stroke patient rehabilitation training management system and method based on Internet platform
CN119400364A