A scoliosis posture sensing and feedback system, and a human posture training action evaluation method and application
By using a multi-camera setup and intelligent hardware system, combined with a digital management platform, personalized scoliosis rehabilitation training can be achieved. This solves the problems of insufficient personalization, inaccurate real-time monitoring, and poor posture assessment in existing systems, thereby improving training effectiveness and patient compliance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI ZHENGJI MEDICAL TECHNOLOGY CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing rehabilitation training systems lack personalization, real-time monitoring, and precise posture assessment, resulting in poor training effects, poor patient compliance, and difficulty in effectively correcting scoliosis in a home environment.
The hardware system, consisting of a multi-camera setup, display screen, voice interaction module, and artificial intelligence module, combined with a digital training management platform and cloud platform, enables personalized rehabilitation plans, real-time feedback, and precise posture assessment.
We provide highly personalized rehabilitation training programs to ensure training quality, improve patient compliance, reduce hospital visits, and enhance training effectiveness and efficiency.
Smart Images

Figure BDA0005164402090000042 
Figure BDA0005164402090000051 
Figure BDA0005164402090000052
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision and artificial intelligence technology, and relates to a scoliosis posture perception and feedback system and a method and application for human posture training movement evaluation. Background Technology
[0002] In the field of rehabilitation medicine, scoliosis is a common spinal deformity characterized by an abnormal lateral curvature of the spine, sometimes accompanied by spinal rotation. The main treatments include bracing, physical therapy, and rehabilitation training, aiming to restore normal posture through long-term training and correction. However, existing rehabilitation training systems face several technical challenges and limitations in practical application.
[0003] Insufficient personalization: Existing rehabilitation training methods typically employ standardized rehabilitation programs, which often fail to adequately consider the individual circumstances of each patient, such as age, severity of scoliosis, and physical condition. Therefore, it is impossible to dynamically adjust the training program according to the patient's specific situation, resulting in significant differences in training effectiveness.
[0004] Difficulties in supervising training: In traditional rehabilitation training, patients need to regularly visit hospitals or rehabilitation centers for guidance from doctors. When rehabilitation training is conducted at home, the lack of real-time supervision can lead to patients failing to perform movements correctly, resulting in poor rehabilitation outcomes. Furthermore, doctors cannot monitor patients' training progress at all times; assessment of rehabilitation progress relies primarily on patients' subjective descriptions or regular hospital checkups, which is inefficient.
[0005] Inadequate feedback mechanisms: Rehabilitation training requires long-term adherence from patients to achieve ideal results. However, due to the lack of immediate feedback and incentive mechanisms, many patients easily lose motivation during training, leading to poor training compliance. This not only delays the rehabilitation process but may also affect treatment outcomes.
[0006] Low accuracy in posture assessment: Currently widely used rehabilitation training systems mainly rely on physician observation or simple sensor devices for posture assessment, resulting in low accuracy. Traditional methods are primarily based on two-dimensional image data, making it difficult to accurately capture and analyze complex three-dimensional posture information, especially when patients are undergoing dynamic training, leading to poor accuracy and reliability of assessment results.
[0007] To address these issues, advanced technologies such as computer vision, deep learning, and multi-camera technology have been increasingly applied to the field of rehabilitation training in recent years. These technologies offer more precise posture recognition and assessment methods, potentially significantly improving the effectiveness of rehabilitation training. However, the application of these technologies still faces challenges, such as how to improve real-time performance while maintaining accuracy, and how to reduce system complexity to suit home environments. Therefore, there is an urgent need for a rehabilitation training system that can better adapt to patients' individual needs, provide real-time feedback, and perform accurate posture assessment. Summary of the Invention
[0008] To address the shortcomings of existing technologies, the purpose of this invention is to provide a scoliosis posture perception and feedback system and a method for evaluating human posture training movements.
[0009] This invention proposes a scoliosis posture sensing and feedback system, which includes a hardware component and a software component.
[0010] The hardware component includes a scoliosis posture sensing and feedback device, which includes a display screen, a multi-camera group, a voice interaction module, and an artificial intelligence module.
[0011] The software component includes a comprehensive digital training management platform with mini-programs and web interfaces, as well as a cloud platform that supports data storage and management, data processing and analysis, AI model training and inference, and ensures data security through encryption technology.
[0012] In one specific implementation, the scoliosis posture sensing and feedback device can be mounted on a base for easy placement on the ground; alternatively, it can be fixed to a suitable location on a wall by means of adhesive or hooks.
[0013] In the scoliosis posture sensing and feedback device of the present invention
[0014] The display screen is installed on the front of the scoliosis posture sensing and feedback device, occupying most of the front area, and is used for content display and image information interaction feedback.
[0015] The multi-camera group includes a color camera and a depth camera; the depth camera includes an infrared projector and an infrared camera, and is connected to an artificial intelligence module; the multi-camera group is used to collect the user's omnidirectional motion data.
[0016] The voice interaction module is located on the lower front of the scoliosis posture sensing and feedback device, and may also include a speaker, microphone, etc., to provide voice information interaction feedback.
[0017] The artificial intelligence module is installed in the device and is used to analyze and process images, voice information, etc., acquired or obtained through interaction, to perform human posture assessment.
[0018] The color camera has a horizontal field of view of 62.7°, a vertical field of view of 49°, and a diagonal field of view of 75.1°±5°; the infrared camera has a horizontal field of view of 58.4°, a vertical field of view of 45.5°, and a diagonal field of view of 70°±5°; and the infrared projector has a horizontal field of view of 67.26° and a vertical field of view of 51°±5°.
[0019] The software component of the digital training integrated management platform mainly includes a doctor-side mini-program, a patient-side mini-program, and a web page management terminal.
[0020] The doctor-side mini-program includes modules such as prescription management, patient management, personal center, and data statistics. The prescription management module allows doctors to dynamically adjust rehabilitation plans based on individual patient differences and treatment progress, and then send these plans to patients. The patient management module helps doctors access patients' training records and videos, and provides remote movement correction suggestions. The personal center integrates patient, team, and plan management functions, optimizing the allocation of doctors' work resources. The data statistics module utilizes time series analysis and machine learning algorithms to predict future patient traffic and prescription demand, assisting in medical resource planning.
[0021] The patient-side mini-program covers receiving and viewing rehabilitation prescriptions, improving understanding and adherence through visual interpretation. Its training management module provides historical training videos, scores, and doctor feedback, and incorporates gamification mechanisms, using a points system, progress display, social interaction, and virtual rewards to enhance patient motivation. The personal center supports management of rehabilitation information, follow-up appointment plans, and doctor communication, and provides 24 / 7 personalized support through intelligent customer service. The health education module provides scoliosis-related knowledge to help patients better understand the disease and treatment process.
[0022] The web page management interface includes a doctor management module, a patient management module, a system management module, and a content management module. The doctor management module uses machine learning to evaluate doctor performance, integrating patient feedback, treatment effectiveness, and workload to ensure fair evaluation. The patient management module integrates medical records and lifestyle habits to provide multi-dimensional patient profiles to support precision treatment. The system management module automatically adjusts the interface layout based on user behavior analysis to improve operational efficiency. The content management module updates the training video library and information, and uses artificial intelligence to evaluate content quality to ensure its accuracy and effectiveness; and / or,
[0023] The cloud platform stores patients' basic information, diagnostic records, and training video data, processes complex events, analyzes data streams, supports the training and deployment of AI models, and employs encryption technology to improve data security during data transmission and storage.
[0024] The present invention also provides a method for evaluating human posture training movements, the evaluation method comprising the following steps:
[0025] Step 1: Collect image data of the patient's training movements in real time using a multi-camera system, and preprocess the collected image data to make it meet the model input requirements;
[0026] Step 2: Extract RGB features and depth features from the preprocessed image data, and fuse the RGB features and depth features into a unified feature representation;
[0027] Step 3: Perform 2D pose estimation, generate heat maps and location affinity fields (PAFs) for key human body points, determine the 2D coordinates of key points; calculate the depth value of each key point, and combine the 2D pose information to convert the key points into 3D coordinates through geometric transformation to complete 3D pose reconstruction.
[0028] Step 4: Based on the 3D pose reconstructed in Step 3, conduct a comprehensive motion quality assessment based on motion similarity score, key angle evaluation, range of motion evaluation, motion speed evaluation, motion stability evaluation, posture symmetry evaluation, and duration evaluation.
[0029] In step one, the image data includes an RGB image containing pose information and a depth image containing depth information; the preprocessing includes image resizing, color space conversion, normalization, etc.; and / or,
[0030] In step two, a ResNet-50 deep convolutional neural network is used to extract features from the RGB image and the depth image respectively, obtaining two-dimensional visual information and depth information, and then fusing them into three-dimensional spatial feature information; and / or,
[0031] In step three, based on the multi-stage CNN architecture of OpenPose, pose features in the image are extracted layer by layer, intermediate results are output at each stage, and pose estimation is optimized step by step to obtain the two-dimensional coordinates of key points; in the depth estimation branch, a fully convolutional network is used to estimate the depth value of each key point from the fused features, representing the position of each body part of the user in three-dimensional space; and / or,
[0032] In step four, the comprehensive motion quality assessment is expressed by the following formula:
[0033] S total =w1S similarity +w2S ROM+w3S velocity +w4S stability +w5S symmetry +w6S duraton
[0034] Among them, S similarity S represents the action similarity score. ROM S represents the range of motion score. velocity S represents the speed of movement score. stability S represents the stability score. symmetry S represents the score for postural symmetry. duration The duration score represents the time spent; w1, w2, w3, w4, w5, and w6 represent the weights corresponding to different scores.
[0035] Furthermore, in step one, normalization of the RGB image and depth image stabilizes model training and accelerates convergence; the normalization of the RGB image is expressed by the following formula:
[0036] Among them, I normalized σ represents the normalized image pixel value, I represents the original image pixel value, μ represents the average value of the image channels, and σ represents the standard deviation of the image channels.
[0037] The normalization of the depth image is expressed by the following formula:
[0038] Among them, D normalized D represents the normalized depth value, and D represents the original depth value. min D represents the minimum depth value of all pixels in a depth image. max Represents the maximum depth value of all pixels in the depth image; and / or,
[0039] In step two, the extracted RGB image features F rgb and depth image features F depth They are expressed by the following formulas respectively:
[0040] F rgb =ResNet50(I rgb )
[0041] F depth =ResNet50(I depth )
[0042] The feature F obtained by fusing the RGB image features and the depth image features fused Represented as: F fused =Conv1x1([F rgb ,F depth ]);
[0043] Using 1x1 convolutional kernels for feature fusion aims to reduce the dimensionality of feature maps and integrate information from different sources while preserving spatial structure information; and / or,
[0044] In step three, the output of each stage in the multi-stage CNN architecture is represented by the following formula: S t =CNN t (F fused ,S t-1 );
[0045] Among them, S t This represents the output of stage t, containing heatmaps of key human body points and location affinity field (PAF) information, CNN. t This represents a convolutional neural network with layer t.
[0046] The depth estimation result output by the fully convolutional neural network is expressed as follows: D est =FCN(F fused );
[0047] Among them, D est The depth estimation result represents the depth value for each keypoint.
[0048] 3D pose reconstruction transforms the positional information of joints and keypoints in a 2D image into pose information in 3D space, providing precise spatial positional information. 3D pose reconstruction is expressed by the following formula:
[0049]
[0050] Where K represents the camera intrinsic parameter matrix, (u,v) represents the coordinates of the two-dimensional keypoints, and z represents the estimated depth; and / or,
[0051] In step four, the action similarity score calculates the similarity between the user's action and the standard action using a dynamic time warping algorithm; the key angle assessment evaluates joint angles including scoliosis angle, trunk tilt angle, and acromion tilt angle; the range of motion assessment calculates the difference between the range of motion of each joint and the standard range; the action speed assessment calculates the difference between the average speed of the assessed key points and the standard speed; the action stability assessment uses the standard deviation of acceleration to assess action stability; the posture symmetry assessment calculates the symmetry of the assessed left and right key points; and the duration assessment calculates the deviation between the actual execution time of the assessed action and the standard time.
[0052] In the human pose training motion evaluation process, temporal smoothing is used to improve the stability of pose estimation between consecutive frames; and / or,
[0053] By selecting appropriate optimization methods and training strategies that control the learning rate during the model's convergence process, model performance can be improved; and / or,
[0054] The loss function defines the objective of model optimization, guiding the model to learn accurate pose estimation and depth prediction; and / or,
[0055] Post-processing is used to further optimize and correct the model output, ensuring that the output conforms to human physical laws or standards; and / or,
[0056] Model optimization, including techniques such as quantization and pruning, can improve the model's inference efficiency and resource utilization.
[0057] Based on the aforementioned human posture training action evaluation results, this invention also provides a personalized training plan recommendation method. This method incorporates the latest deep learning technologies, including Transformer architecture, multimodal fusion, dynamic temporal attention, reinforcement learning, and uncertainty estimation, to obtain a personalized recommendation model. It can process various types of input data, capture long-term and short-term patterns in the patient's rehabilitation process, and generate personalized, interpretable rehabilitation training plan suggestions.
[0058] Step 1: Collect data including patient history, physical condition, and doctor's advice, and convert the above data into features;
[0059] Step II: Use a model based on Transformer architecture and multimodal fusion to process input data from different modalities, and combine it with a multi-head self-attention mechanism to process the data;
[0060] Step III: Fuse information from different modalities through a gating mechanism to generate a comprehensive hidden feature representation;
[0061] Step IV: Based on the model output, generate a personalized rehabilitation training plan for the patient using reinforcement learning methods, and capture important time points in the patient's rehabilitation process using a dynamic time attention mechanism.
[0062] And / or,
[0063] The model includes a feature embedding layer, a temporal encoder, a multi-head self-attention layer, a multimodal fusion layer, a feedforward network, and an output layer component;
[0064] The feature embedding layer transforms numerical, categorical, or time-series data into feature vector representations suitable for model processing.
[0065] e i =Embed i (x i )
[0066] Where, x i Embed represents different types of input. i Indicates different embedding methods;
[0067] The timing encoder converts the temporal information in the timing data into input that the model can understand through position encoding;
[0068]
[0069] Where pos is the position, i is the dimension, and d is the dimension. model It is the model dimension;
[0070] The multi-head self-attention layer captures the global dependencies between input features through a self-attention mechanism, thereby achieving information aggregation;
[0071]
[0072] MultiHead(Q,K,V)=Concat(head1,…,head h W o
[0073] in It is a learnable parameter matrix, Q, K, V: query, key, and value matrix; d k Indicates the column dimension in the key matrix;
[0074] The multimodal fusion layer uses a gating mechanism to fuse features from different data modalities to generate a unified comprehensive representation;
[0075] g i =σ(W g [h1;h2;…;h n ]+b g )
[0076]
[0077] Among them, h i It is a hidden representation of different modalities, g i It is a gating value used to control the information flow of each modality; its dimension is related to each h. i Same; W g It is a learnable weight matrix with dimension [d] out ,d total ], where d out It is the output dimension, d total b is the total dimension of the hidden representation of all connections; g It is a bias vector with dimension [d out ]; ⊙ represents element-wise multiplication; σ is the activation function that compresses the output to between 0 and 1; hfused This represents the final fused representation, which incorporates information from all modalities.
[0078] The feedforward network performs further nonlinear transformations on the fused features to generate the model output; FFN(x) = max(0, xW1+b1)W2+b2
[0079] Where W1, W2 are learnable weight matrices; b1, b2 are bias vectors.
[0080] The output layer generates the final recommendation result based on the results of the feedforward network; y = σ(W out h final +b out )
[0081] Among them, W out It is a learnable weight matrix; b out It is the bias vector; h final It is the feature representation of the last layer.
[0082] The human posture training movement evaluation method of the present invention may further incorporate an anomaly detection model to detect abnormal movements and assess the degree of danger; the anomaly detection includes the following steps:
[0083] Step 1: Extract the patient's training motion data from the 3D skeletal joint coordinate sequence and preprocess it to adapt to the model input requirements;
[0084] Step 2: Use a 3D convolutional neural network to extract spatial and temporal features from the input action sequence, capturing local and global information about the action;
[0085] Step 3: Model the extracted features using a Long Short-Term Memory (LSTM) network to capture the temporal dependencies in the action sequence;
[0086] Step 4: Use a fully connected layer to classify the output of the LSTM and calculate the probability that the current action is normal or abnormal.
[0087] Step 5: Calculate the risk level score of the abnormal movement through the regression layer to determine the potential risk of harm to the patient.
[0088] Step 6: Based on the probability and danger level score of the anomaly detection, the system determines whether the action is safe, abnormal but safe, or dangerous, and provides corresponding feedback.
[0089] And / or,
[0090] The anomaly detection model includes an input layer, a feature extraction layer, a sequence modeling layer, a classifier, a regression layer, and an output layer; wherein,
[0091] 1) Input layer: Receives 3D skeleton joint coordinate sequences;
[0092] 2) Feature extraction layer: 3D CNN feature extraction: F t =CNN(X) t-w:t )
[0093] Among them, F t X is the eigenvector at time step t; t-w:t is the input sequence from tw to t; w is the size of the time window; CNN(·) represents the 3D CNN operation;
[0094] 3) Sequence modeling layer: LSTM sequence modeling: h t ,c t =LSTM(F t ,h t-1 ,c t-1 )
[0095] Among them, h t It is the hidden state at time step t; c t It is the cell state at time step t; F t These are features extracted by CNN; LSTM(·) represents LSTM unit operation;
[0096] 4) Classifier: Anomaly classification: p(y t =Abnormal|X 1:t )=σ(W c h t +b c )
[0097] Where p(y t =Abnormal|X 1:t W represents the probability that an action is anomalous, given all inputs up to time step t. c It is the classifier weight matrix; h t It is the hidden state at time step t; b c It is the bias term; σ(·) is the sigmoid activation function;
[0098] 5) Regression layer: Risk assessment: r t =ReLU(W r h t +b r )
[0099] Where, r t It is a risk level score for time step t; W r It is the regression layer weight matrix; h t It is the hidden state at time step t; b r It is the bias term; ReLU(·) is the modified linear unit activation function;
[0100] 6) Output layer: Final decision:
[0101]
[0102] Where θ is the threshold for anomaly detection; δ is the threshold for the degree of danger.
[0103] The present invention also provides the application of the above-mentioned perception and feedback system, or the above-mentioned human posture training movement assessment method, in human training posture perception, personalized rehabilitation training, remote medical support, training anomaly detection and risk prevention and control, training data management, etc.
[0104] The beneficial effects of this invention include: The system possesses highly personalized functionality, enabling customized rehabilitation training plans based on the patient's specific condition. Through real-time monitoring with AI technology, the system can accurately analyze the patient's movements, ensuring the quality of each training session. Simultaneously, doctors can remotely view the training progress and provide feedback at any time, truly achieving real-time guidance. To improve patient compliance, the system incorporates a gamified training mechanism combined with instant feedback to stimulate patient engagement. Based on accurate evaluation using objective data, the system can analyze training effects in detail and effectively reduce potential risks during training through anomaly detection algorithms, ensuring safety and reliability. The system employs a data-driven approach to continuously optimize training plans and AI models, improving overall training effectiveness and efficiency. Furthermore, patients can conduct home training through the system, reducing the frequency of hospital visits, significantly saving time and costs, and making the rehabilitation process more convenient and efficient. Attached Figure Description
[0105] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0106] Figure 1 This is a schematic diagram of the smart hardware structure of the present invention.
[0107] Figure 2 This is a schematic diagram of the system architecture of the present invention.
[0108] Figure 3 This is a flowchart of the system of the present invention applied to scoliosis correction.
[0109] Figure 4 This is a schematic diagram of the interface of the staff (doctor) mini-program of this invention.
[0110] Figure 5 This invention is a schematic diagram of the user-side (patient-side) mini-program interface.
[0111] Figure 6 This is a schematic diagram of the web system interface of the present invention. Detailed Implementation
[0112] The present invention will be further described in detail below with reference to the specific embodiments and accompanying drawings. Except for the contents specifically mentioned below, the processes, conditions, and experimental methods for implementing the present invention are all common knowledge and general knowledge in the art, and the present invention does not have any particular limitations.
[0113] This invention provides a scoliosis posture perception and feedback system, comprising hardware and software components. The hardware component includes a scoliosis posture perception and feedback device, which includes a display screen, a multi-camera setup, a voice interaction module, and an artificial intelligence chip. The software component includes a comprehensive digital training management platform with mini-programs and a web interface, as well as a cloud platform supporting data storage and management, data processing and analysis, AI model training and inference, and ensuring data security through encryption technology. This invention also provides a method for evaluating human posture training movements, used to assess training movements and for recommending subsequent personalized training programs.
[0114] This invention relates to a scoliosis posture perception and feedback system and a corresponding human posture training movement assessment method, which can be used for subsequent home training. The posture perception and feedback system specifically involves intelligent hardware and software components. In specific implementation, the intelligent hardware component, referred to as the "scoliosis training magic mirror" (or simply "magic mirror" in this text), mainly consists of a display screen, a multi-camera system, a voice interaction module, and an artificial intelligence module. It enables the playback of standard videos, recording and uploading of user training movements, real-time judgment and feedback of human posture, and processing and display of interactive information. The software component mainly includes a comprehensive digital training management platform consisting of a web-based management interface, a WeChat mini-program for staff (doctors), and a user (patient) interface, as well as a cloud platform supporting data storage and management, data processing and analysis, AI model training and inference, and ensuring data security through encryption technology. The posture perception and feedback system can support functions such as diagnosis, prescription of training exercises, video review and feedback of training movements, and follow-up visit reminders.
[0115] For a specific user, once staff determines that they have scoliosis and require rehabilitation training, they use a mini-program to display a QR code on the staff's (doctor's) end. The user scans the code, fills in their personal information, and uploads necessary clinical data. After receiving and verifying the user's identity, the staff fills in the diagnosis and issues a training prescription, which is then sent directly to the user's mini-program. Once the user receives and confirms the training prescription, they can begin the corresponding training according to the prescription's requirements.
[0116] During training, the posture perception and feedback system displays details of the training prescription to the user through the "Magic Mirror," including the training content, duration, frequency, and key points. When the user clicks on a specific training item, the system first plays a standard instructional video for that item, then prompts the user to begin training. After detecting that the user's position is appropriate, it prompts the user to start. During training, the Magic Mirror uses a camera array and algorithm model to detect the user's human posture feature points in real time to determine whether their movements meet the standard requirements. For movements that are clearly substandard, it provides real-time voice prompts and correction requirements. The Magic Mirror automatically provides voice prompts at the start of training and at key points during the process (such as 50%, 75%, etc.). Upon completion, the system automatically provides an AI score for the completion of each training item. It should be noted that:
[0117] The algorithm model in this invention employs an improved OpenPose algorithm, which is mounted on the artificial intelligence module. In this invention, the magic mirror uses the lens assembly and the improved OpenPose algorithm model mounted on the artificial intelligence module to perform real-time detection and feedback of user actions. The specific usage process is as follows:
[0118] Data acquisition: The camera system captures images of the user's training movements in real time;
[0119] Preprocessing: The acquired images are preprocessed, including image resizing, color space conversion, normalization, etc., to adapt to the input requirements of the algorithm model;
[0120] The image size adjustment refers to scaling the acquired image so that the image input to the model has the same resolution and size, making it easier for the model to process it subsequently.
[0121] The color space conversion refers to converting the image in color spaces such as RGB, brightness, and chromaticity according to actual needs, so that the model can better capture information such as the edge and contrast of the object;
[0122] Normalization refers to scaling pixel values from an integer range of 0-255 to a floating-point range of 0-1, which can make the model more stable during training and speed up convergence.
[0123] Feature extraction: The improved OpenPose algorithm model used in this invention uses a multi-stage convolutional neural network (CNN) to extract features, and combines it with Part Affinity Fields (PAF) technology, and incorporates information from depth images to achieve the correlation between key points and ensure stability in complex backgrounds;
[0124] Key point recognition: Based on the extracted features, this invention uses deep learning and computer vision technology to detect the position and movement of 18 key points on the patient's body in real time (nose, right ear, left ear, right eye, left eye, neck, right shoulder, right elbow, right hand, left shoulder, left elbow, left hand, right hip, right knee, right foot, left hip, left knee, and left foot).
[0125] Motion evaluation: The pose estimated by 3D reconstruction is compared with a predefined standard motion template to calculate the deviation of key point positions and angle differences, and a reasonable motion evaluation algorithm model is designed for AI scoring;
[0126] Real-time feedback: Based on the action evaluation and analysis results and the corresponding scores, the system judges whether the user's actions meet the standards. If a significant deviation is found, it immediately provides correction suggestions through voice prompts.
[0127] Data recording: The detected posture data is recorded for subsequent training effect evaluation and personalized program adjustment.
[0128] The entire training process is recorded by the multi-camera system on the Magic Mirror and uploaded to the staff (doctor's) end for review and guidance. When staff receive a new training video, they will receive a system notification. After reviewing it, they can send their guidance in text form to the user (patient's) mini-program. Before each new training session, the user will receive feedback from staff to optimize the day's training. Based on prescription requirements, the system will notify staff and the user in text form 7 days before the follow-up appointment to ensure timely scheduling.
[0129] Specifically, in one embodiment, the scoliosis posture sensing and feedback system of the present invention mainly includes: a hardware part and a software part;
[0130] The hardware component is the smart hardware – the "Scoliosis Training Magic Mirror," with the following structure: Figure 1 As shown, it mainly includes:
[0131] High-definition display screen: used to display training videos, real-time feedback, and interactive information between staff and users; the high-definition display screen is installed on the front of the magic mirror, occupying most of the front area of the magic mirror;
[0132] Multi-camera group: Collects omnidirectional motion data of the patient, including a color camera and a depth camera (including an infrared projector and an infrared camera). The data obtained by the color camera and / or the depth camera will be processed by an artificial intelligence module.
[0133] The upper part of the scoliosis training magic mirror consists of an infrared projector, a color camera, and an infrared camera, from left to right.
[0134] Color camera FoV: H 62.7°, V 49°, D 75.1°±5°; Infrared camera FoV: H 58.4°, V 45.5°, D 70°±5°; Infrared projector FoV: H 67.26°, V 51°±5°;
[0135] FoV refers to the field of view, which describes the angular range of a camera observing a given scene. There are three main types: horizontal field of view (HFoV), vertical field of view (V FoV), and diagonal field of view (D FoV).
[0136] The artificial intelligence module includes a high-performance motherboard and an AI chip for handling complex computational tasks, installed inside the magic mirror. The AI chip supports deep image processing and edge computing. The system utilizes cutting-edge chips, such as Google's Edge TPU or Qualcomm's AI Engine, to achieve efficient local inference. The cloud employs a federated learning framework to achieve distributed updates and optimization of the model while protecting privacy.
[0137] Voice interaction module: including speakers, microphones, etc., providing voice guidance and feedback, installed at the bottom of the Magic Mirror.
[0138] The software component includes a comprehensive digital training management platform with mini-programs and web interfaces, as well as a cloud platform that supports data storage and management, data processing and analysis, AI model training and inference, and ensures data security through encryption technology.
[0139] In this invention, an improved OpenPose model is used to implement a method for evaluating human posture training movements. The evaluation method includes the following steps:
[0140] Step 1: Collect image data of the patient's training movements in real time using a multi-camera system, and preprocess the collected image data to make it meet the model input requirements;
[0141] Step 2: Extract RGB features and depth features from the preprocessed image data, and fuse the RGB features and depth features into a unified feature representation;
[0142] Step 3: Perform 2D pose estimation, generate heat maps and location affinity fields (PAFs) for key human body points, determine the 2D coordinates of key points; calculate the depth value of each key point, and combine the 2D pose information to convert the key points into 3D coordinates through geometric transformation to complete 3D pose reconstruction.
[0143] Step 4: Based on the 3D pose reconstructed in Step 3, conduct a comprehensive motion quality assessment based on motion similarity score, key angle evaluation, range of motion evaluation, motion speed evaluation, motion stability evaluation, posture symmetry evaluation, and duration evaluation.
[0144] The improved OpenPose model used in the evaluation method of this invention incorporates data from a depth camera and combines multi-stage CNN and fully convolutional networks to achieve high-precision 3D human pose estimation, outputting 18 human key points in real time. Furthermore, a temporal smoothing module is added to the model to ensure stability across frames, making it particularly suitable for rehabilitation training scenarios.
[0145] In this invention, the improved OpenPose model architecture includes: an input layer, a feature extraction network, a 2D pose estimation branch, a depth estimation branch, a 3D pose reconstruction module, and a temporal smoothing module;
[0146] The input layer includes: RGB image stream: 3 channels, resolution 640x480; depth image stream: 1 channel, resolution 640x480.
[0147] The feature extraction network uses ResNet-50 as the backbone network to process RGB images and depth images respectively, and performs feature fusion on the extracted RGB features and depth features.
[0148] The 2D pose estimation branch uses a multi-stage CNN structure to output heatmaps and part affinity fields (PAFs);
[0149] The depth estimation branch uses a fully convolutional network (FCN) structure to output the depth value of each keypoint;
[0150] The 3D pose reconstruction module combines the obtained 2D pose estimation results and depth estimation information to convert the 2D coordinates and depth information into 3D coordinates through geometric transformation.
[0151] The temporal smoothing module eliminates jitter through a Kalman filter, achieving smooth 3D pose estimation across frames.
[0152] In the human posture training action evaluation method of the present invention, in step one, the image data includes an RGB image containing posture information and a depth image containing depth information. The preprocessing includes image size adjustment, color space conversion, and normalization. The RGB image normalization is shown in the following formula:
[0153] Among them, I normalized σ represents the normalized image pixel value, I represents the original image pixel value, μ represents the average value of the image channels, and σ represents the standard deviation of the image channels.
[0154] Depth image normalization is shown in the following formula:
[0155] Among them, D normalized D represents the normalized depth value, and D represents the original depth value. min D represents the minimum depth value of all pixels in a depth image. max This represents the maximum depth value of all pixels in the depth image.
[0156] In step two, ResNet-50 deep convolutional neural network is used to extract features from RGB images and depth images respectively, to obtain two-dimensional visual information and depth information, and then fuse them into three-dimensional spatial feature information;
[0157] RGB feature F rgb and depth features F depth Extract as shown in the following formula:
[0158] F rgb =ResNet50(I rgb )
[0159] F depth =ResNet50(I depth )
[0160] The feature F obtained by fusing the RGB image features and the depth image features fused Represented as:
[0161] F fused =Conv1x1([F rgb ,F depth ])
[0162] Here, Conv1x1 indicates that a 1x1 convolution kernel is used for feature fusion, the purpose of which is to reduce the dimensionality of the feature map and integrate information from different sources while preserving spatial structure information;
[0163] In step three, based on the multi-stage CNN architecture of OpenPose, the pose features in the image are extracted layer by layer, intermediate results are output at each stage, and the pose estimation is optimized step by step to obtain the two-dimensional coordinates of the key points. In the depth estimation branch, the depth value of each key point is estimated from the fused features through a fully convolutional network, which represents the position of each body part of the user in three-dimensional space.
[0164] The 2D estimation stage employs a multi-stage CNN structure, with the output of each stage represented as follows: S t =CNN t (F fused ,S t-1 )
[0165] Where S t This is the output of stage t, containing heatmap information of key points on the human body and location affinity field (PAF) information, CNN. t This represents a convolutional neural network with layer t.
[0166] In the depth estimation stage, a fully convolutional network structure is adopted, as shown in the following equation: D est =FCN(F fused )
[0167] Among them, D est The depth estimation result represents the depth value for each keypoint.
[0168] 3D pose reconstruction is performed based on 2D coordinate and depth information. Given the camera intrinsic matrix K, 2D keypoints (u,v), and estimated depth z, the 3D pose reconstruction is expressed by the following formula:
[0169]
[0170] Where K represents the camera intrinsic parameter matrix, (u,v) represents the coordinates of the two-dimensional keypoint, and z represents the estimated depth;
[0171] In this invention, the stability across frames can be improved by adding temporal smoothing (Kalman filtering);
[0172] The state equation is as follows: x k =Fx k-1 +w k
[0173] The measurement equation is shown below: z k =Hx k +v k
[0174] Where F is the state transition matrix, H is the measurement matrix, and w k and v k These are process noise and measurement noise, respectively.
[0175] In the prediction step of Kalman filtering, the system is based on the state of the previous time step. And the state transition matrix F is used to predict the state at the current time step. The error covariance of the state is predicted, taking into account the uncertainty of the prediction due to process noise, as shown in the following equation:
[0176] Equations of state
[0177] Covariance prediction equation
[0178] in: It is the prior state estimate (predicted value) at time k. It is the posterior state estimate of the previous time step k-1. P is the prior error covariance matrix at time k. k-1 It is the posterior error covariance matrix of the previous time k-1, F is the state transition matrix, and Q is the process noise covariance matrix;
[0179] In the update step of the Kalman filter, the system uses the actual measured value z k The predicted value is corrected by calculating the Kalman gain K. k This is used to balance the weights of predicted and measured values. A larger Kalman gain indicates that the system trusts the measured values more, while a smaller gain indicates that the system trusts the prediction model more. Finally, the covariance matrix P... k It will also be adjusted based on the updated status to reflect the new level of uncertainty, as shown in the following formula:
[0180] Kalman gain calculation
[0181] State update equation
[0182] Covariance update equation
[0183] Where: K k Here, H is the Kalman gain, H is the observation matrix, R is the observation noise covariance matrix, and z is the Kalman gain. k It is the actual observed value at time k. It is the posterior state estimate (updated estimate) at time k, P k Let I be the posterior error covariance matrix at time k, and let I be the identity matrix.
[0184] In the 3D pose reconstruction process of step three, the loss functions for different steps are shown in the following equations:
[0185] 2D heatmap loss:
[0186] Among them, H i This is the heatmap of the i-th key point predicted by the model. It is the actual heatmap of the i-th key point, and N is the total number of key points;
[0187] PAF loss:
[0188] Among them, Pj It is the j-th PAF predicted by the model. is the true value of the j-th PAF, and M is the total number of PAFs;
[0189] Depth estimation loss:
[0190] Among them, D k It is the depth value of the k-th keypoint predicted by the model. It is the true depth value of the kth keypoint, where K is the total number of keypoints whose depth needs to be estimated.
[0191] 3D pose reconstruction loss:
[0192] Among them, X l These are the 3D coordinates of the l-th keypoint in the model reconstruction. is the true 3D coordinate of the l-th keypoint, and L is the total number of 3D keypoints;
[0193] Total loss: L total =αL heatmap +βL PAF +γL depth +δL 3D
[0194] Where α, β, γ, and δ are weighting coefficients;
[0195] In this invention, the model can be optimized by setting a training strategy, and the strategy parameters include the following:
[0196] Optimizer: Adam; Initial learning rate: 1e-4; Learning rate decay: 10% every 50 epochs; Batch size: 32; Number of training epochs: 300 epochs.
[0197] The present invention can further optimize and correct the output results of the model through post-processing, including the following:
[0198] Anatomical constraints:
[0199] For adjacent key points i, j, k, the angle θ is restricted to a reasonable range:
[0200]
[0201] θ min ≤θ≤θ max
[0202] Among them, X i ,X j ,X k Let i, j, k be the coordinates of the key points;
[0203] Joint angle calculation:
[0204] Where v1 and v2 are adjacent bone vectors;
[0205] In addition to the optimizations mentioned above, model optimization can also improve the accuracy of the model, including the following:
[0206] Quantization: Converting floating-point operations to fixed-point operations, for example: Q(r) = round(r / scale + zero_point);
[0207] Pruning: Remove weights whose importance is below a threshold t: W pruned =W⊙(|W|>t);
[0208] In step four, a comprehensive evaluation scheme was designed by combining medical expertise and AI technology, including scores for action similarity, range of motion, action speed, stability, postural symmetry, and duration; specifically,
[0209] 1) Action similarity score
[0210] The Dynamic Time Warping (DTW) algorithm is used to calculate the similarity between the patient's actions and the standard actions:
[0211]
[0212] Where X and Y represent the patient and the standard action sequence, respectively, d(x i ,y j () is the Euclidean distance between two poses;
[0213] Similarity score:
[0214] 2) Key Angle Assessment
[0215] Calculate the angles of important joints, such as:
[0216] Cobb angle:
[0217] Where m1 and m2 are the slopes of the two most inclined vertebrae on the spinal curve;
[0218] Trunk tilt angle:
[0219] x shoulder y shoulder The x and y coordinates represent the shoulders. hip y hip The x and y coordinates represent the buttocks;
[0220] acromion tilt angle:
[0221] This represents the x and y coordinates of the left shoulder. Represents the x and y coordinates of the right shoulder;
[0222] 3) Range of motion (ROM) assessment
[0223] Calculate the range of motion of each joint and compare it with the standard range: ROM joint =max(θ) joint )-min(θ joint )
[0224] score:
[0225] In range of motion assessment, the angle θ of joint movement is involved. joint This includes the spinal scoliosis angle, trunk tilt angle, and peak tilt angle in the key angle assessment in section 2);
[0226] 4) Action speed assessment
[0227] Calculate the average velocity at key points and compare it with the standard velocity:
[0228] Where p i This represents the position vector of a key point at a certain moment, where Δt is the time interval.
[0229] score:
[0230] v std This represents the standard speed, which is the ideal speed of movement set according to standard instructional videos or expert advice.
[0231] 5) Motion stability assessment
[0232] Use the standard deviation of acceleration to assess motion stability:
[0233] a i This represents the acceleration vector of a key point at a specific moment. It can be approximated by the second-order difference of the key point's position. It is the average value of acceleration; σ acc It is the standard deviation of acceleration, used to measure the stability of a movement; a smaller standard deviation indicates a more stable movement.
[0234] Stability score:
[0235] Here, k is an adjustable parameter used to control the sensitivity of the stability score. A larger k value makes the score more sensitive to changes in acceleration. Typical k values may be between 0.1 and 10, depending on the nature of the movement and the rigor of the evaluation. For example:
[0236] For movements that require very high stability (such as balance training), a larger k value, such as 5 or 10, may be used.
[0237] For actions that allow for some degree of variation, a smaller k value, such as 0.5 or 1, might be used.
[0238] 6) Postural symmetry assessment
[0239] Calculate the symmetry of the key points on the left and right sides:
[0240] Where, d left,i d represents the position of the left keypoint in the i-th keypoint pair; rght,i This indicates the position of the right keypoint in the i-th keypoint pair;
[0241] 7) Duration Assessment
[0242] For each rehabilitation exercise, there is usually a recommended standard execution time, and the actual execution time T is... actual With standard time T std The deviation.
[0243]
[0244] Among them, T std In collaboration with medical experts, the execution time for each rehabilitation exercise is set based on factors such as the patient's age and the severity of their condition;
[0245] The final action quality score is obtained by combining the scores from all the above.
[0246] S total =w1S similarity +w2S ROM +w3S velocity +w4S stability +w5S symmetry +w6S duration
[0247] Here, w1, w2, w3, w4, w5, and w6 are the weights of each score, and their sum equals 1. These weights can be adjusted according to specific rehabilitation training needs and the doctor's professional judgment.
[0248] The overall score is normalized to a percentage scale, ranging from 0 to 100, for easier understanding and comparison:
[0249] 90-100 points: Excellent, fully meets or exceeds expectations; 80-89 points: Good, basically meets expectations, with minor room for improvement; 70-79 points: Average, meets basic requirements, needs significant improvement; 60-69 points: Pass, barely completes, needs substantial improvement; Below 60 points: Fail, does not meet basic requirements, needs retraining.
[0250] Based on the aforementioned human posture training movement evaluation results, this invention also provides a personalized training plan recommendation method. This method combines the latest deep learning technologies, including Transformer architecture, multimodal fusion, dynamic temporal attention, reinforcement learning, and uncertainty estimation, to obtain a personalized recommendation model. This personalized recommendation model can process various types of input data, capture long-term and short-term patterns in the patient's rehabilitation process, and generate personalized, interpretable rehabilitation training plan suggestions.
[0251] This model can better adapt to the needs of different users, and through continuous learning and uncertainty estimation, it can continuously improve the quality of its recommendations over time. This advanced approach will greatly improve the effectiveness and efficiency of training for scoliosis users.
[0252] The specific scheme of the personalized recommendation method for the training plan is as follows:
[0253] Step 1: Collect data including patient training history, physical condition, and doctor's advice, and convert the data into feature vectors;
[0254] The data to be collected includes the following:
[0255] Patient history data: including past training records, scores, and progress;
[0256] Physical condition: such as age, gender, height, weight, degree of scoliosis (Cobb angle);
[0257] Doctor's advice includes recommended training types, intensity, and frequency;
[0258] The collected data is transformed into features usable by the model: X = [x1, x2, ..., x n ]
[0259] Where: X is the feature vector, x i These represent different characteristics, such as: x1: age; x2: Cobb angle; x3: overall score of the last training session; x4: average training duration over the past week; etc.
[0260] Step II: Use a model based on Transformer architecture and multimodal fusion to process input data from different modalities, and combine it with a multi-head self-attention mechanism to process the data;
[0261] The personalized recommendation model in this invention uses a Transformer-based Multimodal Fusion Network model structure. This model is based on the Transformer architecture and combines multimodal fusion technology. It mainly includes the following components: feature embedding layer, temporal encoder, multi-head self-attention layer, multimodal fusion layer, feedforward network, and output layer.
[0262] 1) Feature embedding layer
[0263] For each type of input x i (e.g., numerical features, categorical features, time series data, etc.) can be embedded using different methods:
[0264] e i =Embed i (x i )
[0265] Embed i It could be a linear layer, word embedding, or other appropriate embedding method;
[0266] 2) Timing encoder
[0267] For time-series data, use position encoding:
[0268]
[0269] Where pos is the position, i is the dimension, and d is the dimension. model It is the model dimension.
[0270] 3) Multi-head self-attention layer
[0271]
[0272] MultiHead(Q,K,V)=Concat(head1,…,head h W O
[0273] head i =Attention(QW i Q ,KW i K VW i V ), W i Q W i K W i V W OIt is a learnable parameter matrix, Q, K, V: query, key, and value matrices; dk represents the column dimension in the key matrix;
[0274] 4) Multimodal fusion layer
[0275] Using gating mechanisms to fuse information from different modalities:
[0276] g i =σ(W g [h1;h2;…;h n ]+b g )
[0277]
[0278] Among them, h i : These are hidden representations of different modalities. For example, h1 might be a representation of the patient's numerical characteristics (such as age, height, weight, etc.), h2 might be a representation of the patient's time-series data (such as past training records), h3 might be a representation of the doctor's text suggestions, and so on.
[0279] g i : is the gating value, used to control the information flow of each modality; its dimension is related to each h. i same;
[0280] W g : is a learnable weight matrix. Its dimension is [d out ,d total ], where d out It is the output dimension, d total It is the total dimension of the hidden representation of all connections;
[0281] b g This is a bias vector with dimension [d] out ];
[0282] ⊙: Indicates element-wise multiplication;
[0283] σ: This is an activation function, usually the sigmoid function is chosen, which compresses the output to between 0 and 1;
[0284] h fused This is the final fused representation, which combines information from all modalities;
[0285] 5) Feedforward network
[0286] FFN(x) = max(0, xW1+b1)W2+b2
[0287] Where W1, W2 are learnable weight matrices; b1, b2 are bias vectors.
[0288] 6) Output layer
[0289] y=σ(W out h final +b out )
[0290] Among them, W out : is a learnable weight matrix; b out : is the bias vector; h final This is the feature representation of the last layer;
[0291] This invention uses a multi-task learning framework, combining multiple loss functions to optimize the recommendation results of a personalized recommendation system;
[0292]
[0293] in:
[0294] It is a binary cross-entropy loss, used for selecting training items; It is the mean squared error loss, used for predicting training intensity and duration; It is the ranking loss, used to optimize the order of training items; λ1, λ2, λ3, and λ4 are regularization terms used to prevent overfitting; λ1, λ2, λ3, and λ4 are weights that balance different losses.
[0295] In the personalized recommendation process of this invention, a dynamic time attention mechanism is also introduced to better capture important time points in the patient's recovery process:
[0296] α t =softmax(v T tanh(W h h t +W x x t +b))
[0297]
[0298] Where: t: represents the time step, which in the context of rehabilitation training may represent different training dates or training sessions;
[0299] h t This is the hidden state at time step t. It is typically generated by a recurrent neural network such as an RNN or LSTM and contains information from previous time steps.
[0300] x t This is the input for time step t. In the scenario described in this application, this may include information such as the training data for the day and the user's status.
[0301] W h W t : is a learnable weight matrix;
[0302] v: is a vector parameter used to map the transformed state and input to a scalar;
[0303] b: is a bias vector;
[0304] α t : is the attention weight at time step t, representing the degree of attention the model pays to that time step;
[0305] c: is the context vector, which is a weighted sum of all hidden states, with the weights determined by the attention mechanism;
[0306] Based on the model output, reinforcement learning methods are used to generate the final rehabilitation training plan:
[0307]
[0308] Where: s t This is the current state, a t It's an exercise (recommended training program), r t It is a reward (such as the patient's progress), and γ is a discount factor;
[0309] This invention also uses incremental learning methods to periodically update the personalized recommendation model:
[0310]
[0311] Where, θ t : Represents the model's parameters at time step t, which includes all trainable weights and biases in the model;
[0312] θ t+1 : Represents the updated model parameters, i.e., the parameters at time step t;
[0313] α: Learning rate, which controls the step size of each update; it is a hyperparameter and needs to be adjusted according to the specific situation.
[0314] This represents the gradient operator with respect to the parameter θ;
[0315] Loss function; in practical applications, this may be a composite loss function, including multiple sub-objectives, such as recommendation accuracy and security;
[0316] This is newly collected data; it may include recent patient training records, feedback, etc.
[0317] The memory replay buffer contains important data that has been previously learned, used to prevent catastrophic forgetting.
[0318] The personalized recommendation model in this invention is interpretable, employing Integrated SHAP to explain the model's decisions:
[0319]
[0320] Where φ i is the SHAP value of feature i, and f is the model function;
[0321] This invention also uses Monte Carlo Dropout to estimate the uncertainty of the model:
[0322]
[0323] Where y t It is the prediction of the t-th forward propagation. It is the average forecast. It is the variance of each prediction;
[0324] The present invention also provides an anomaly detection method for identifying potentially risky actions and ensuring safety.
[0325] To ensure training safety, the system integrates algorithms based on deep anomaly detection networks. These algorithms, by learning deep representations of normal actions, are able to more accurately identify potentially dangerous actions.
[0326] The "dangerous actions" in the anomaly detection algorithm module mainly refer to actions that may cause harm to the patient's body. This includes, but is not limited to: actions that are beyond the patient's current ability, incorrect postures or techniques that may lead to acute injury, worsen scoliosis, or cause other complications; at the same time, it will also detect actions that are incorrect compared to the teaching demonstration as a reference for improving the patient's technique, but these are not necessarily classified as "dangerous" unless they may cause injury.
[0327] The detection model architecture used in the anomaly detection method of this invention is as follows:
[0328] The detection model, by combining 3D CNN and LSTM, can effectively capture the spatial-temporal features in motion sequences, thereby accurately detecting abnormal movements and assessing their degree of danger. It can not only identify abnormalities that deviate from standard movements but also assess the potential harm these abnormalities may cause, providing safer guidance for rehabilitation training.
[0329] The structure of the above detection model includes:
[0330] Input layer: Receives action sequence data, typically a 3D skeletal joint coordinate sequence; Feature extraction layer: Uses a 3D convolutional neural network (3D CNN) to extract spatial-temporal features; Sequence modeling layer: Uses a long short-term memory network (LSTM) to model temporal dependencies; Classifier: Uses fully connected layers to classify actions as normal or abnormal; Regression layer: Evaluates the degree of danger of the action; Output layer: Combines the classification and regression results to output the anomaly detection result and the degree of danger assessment.
[0331] 1) Input layer: Receives 3D skeleton joint coordinate sequences;
[0332] 2) Feature extraction layer: 3D CNN feature extraction: F t =CNN(X) t-w:t )
[0333] Among them, F t : is the eigenvector at time step t; X t-w:t : is the input sequence from tw to t; w: is the size of the time window; CNN(·): indicates a 3D CNN operation;
[0334] 3) Sequence modeling layer: LSTM sequence modeling: h t ,c t =LSTM(F t ,h t-1 ,c t-1 )
[0335] Where: h t : is the hidden state at time step t; c t : is the cell state at time step t; F t : represents the features extracted by CNN; LSTM(·) represents the LSTM unit operation;
[0336] 4) Classifier: Anomaly classification: p(y t =Abnormal|X 1:t )=σ(W c h t +b c )
[0337] Where: p(y t =Abnormal|X 1:t ): represents the probability that an action is abnormal given all inputs up to time step t; W c : is the classifier weight matrix; b c : is the bias term; σ(·) is the sigmoid activation function;
[0338] 5) Regression layer: Risk assessment: r t =ReLU(W r h t +br )
[0339] Where: r t It is a risk level score for time step t; W r It is the regression layer weight matrix; b r It is the bias term; ReLU(·) is the modified linear unit activation function;
[0340] 6) Output layer: Final decision:
[0341]
[0342] Where: θ: is the threshold for anomaly detection; δ: is the threshold for the degree of danger;
[0343] In one specific embodiment, the present invention also provides a digital therapy software system based on WeChat mini-programs and web interfaces for use between doctors and patients, comprising:
[0344] 1) Doctor's side (WeChat mini-program), such as Figure 4 As shown, it includes the following structure and content:
[0345] Prescription Management: This module provides patients with functions such as issuing diagnoses, selecting attending physicians, providing prescription plans, rehabilitation plans, contraindications, precautions, and follow-up visit schedules, and sending these to patients. It also features a dynamic template library that automatically adjusts prescription content based on individual patient differences and treatment progress, enabling precise formulation of rehabilitation plans.
[0346] Patient Management: This module allows users to query training records of patients who have not received guidance or feedback, view the patient's training video for the day and provide guidance. It can also automatically analyze training videos to provide suggestions for correcting patient movements and improve the accuracy of remote guidance.
[0347] Personal Center: Includes patient management, my team management, treatment plan management, qualification certification, online customer service, etc. The team management module is designed with an intelligent team collaboration scheduling solution, which automatically allocates patients based on doctors' expertise and workload, optimizing the allocation of medical team resources.
[0348] Data statistics: This module provides information such as the number of patients, prescriptions, and follow-up appointments for this week and this month. It uses time series analysis and machine learning algorithms to predict patient flow and prescription demand in the future, thus assisting in the planning of medical resources.
[0349] 2) Patient-side (WeChat mini-program), such as Figure 5 As shown, it includes the following structure and content:
[0350] Prescription Information: Receives and confirms rehabilitation prescriptions and views progress management information. With the help of the visual prescription interpretation function, professional medical terms are transformed into easy-to-understand graphic explanations, improving patients' understanding and execution of prescriptions.
[0351] Training Management: View historical training videos, scores, and doctor feedback for training programs. It's worth noting that this function incorporates a gamified training mode online, dynamically adjusting difficulty and reward mechanisms based on patient training data to enhance training motivation. The specific gamification method in this invention is as follows:
[0352] Points system: Patients may earn points for each training session they complete. Different levels are set based on the points, such as "Rehabilitation Novice", "Persistence Expert", "Rehabilitation Master", etc. Each registration will have a corresponding achievement badge. Points can be calculated based on the quality, completion and persistence of the training.
[0353] Progress visualization: Use progress bars or other graphical methods to display the patient's rehabilitation progress; and include daily, weekly or monthly training challenges and completion rates, such as "Completed training for 7 consecutive days" or "Training score reached 90 points this week".
[0354] Social interaction: It may allow patients to interact with other patients within the system and share their rehabilitation experiences; set up leaderboards to display training duration, persistence, etc., to stimulate healthy competition;
[0355] Virtual rewards: Unlock virtual items after completing specific goals, such as personalized avatars and special themes;
[0356] Real-time feedback system: Use vivid visual and audio feedback, use animation effects to intuitively show the difference between the patient's movements and the standard movements, such as playing cheers when the patient completes a movement; provide encouraging voice prompts during training, and display the accuracy and score of the movements in real time;
[0357] Personal Center: This module manages information on recovered patients, online customer service, follow-up appointment plans, and online doctor communication. Specifically, it automatically recommends the best follow-up appointment time based on the patient's recovery progress and the doctor's schedule, and implements an intelligent customer service system based on natural language processing to provide patients with 24 / 7 instant response and personalized suggestions.
[0358] Health education: Acquire knowledge about scoliosis and improve understanding of the disease and treatment. This part will customize educational content based on the patient's cognitive level and rehabilitation stage, develop a personalized learning path, and optimize the knowledge absorption effect.
[0359] 3) Management interface (web page), such as Figure 6 As shown:
[0360] Physician Management: This module includes a physician list and service information. It introduces a machine learning-based physician performance evaluation system that comprehensively considers patient feedback, treatment effectiveness, and workload to achieve a fair and reasonable evaluation mechanism.
[0361] Patient Management: This module includes patient lists, rehabilitation prescriptions, and training details. It develops a multi-dimensional patient profiling system that integrates medical records, lifestyle habits, and social factors to provide comprehensive data support for precision treatment.
[0362] System management includes user, department, and job responsibility management. It features an adaptive interface based on user behavior analysis, automatically adjusting the function layout according to the usage habits of different roles to improve operational efficiency.
[0363] Content Management: This section updates the standard training video library, including science popularization and informational videos, as well as statistical reports on doctors and patients, and evaluates system effectiveness. It implements an AI-powered content quality assessment algorithm to automatically review and optimize training videos and science popularization articles, ensuring the accuracy and effectiveness of the content.
[0364] The relevant data storage, processing, and model training in this invention can be performed on a cloud platform:
[0365] Data storage:
[0366] MySQL: Stores structured patient basic information, account data, etc.; MongoDB: Stores semi-structured text data (such as doctors' diagnostic records), etc.; HDFS: Stores large-scale raw training data and log files; S3 Database: Stores and distributes media files (such as training videos); Redis: Used for caching and session management.
[0367] Data processing:
[0368] Batch processing: Apache Spark: Large-scale data processing and analysis; Apache Flink: Complex event processing and data stream analysis; Stream processing: Apache Kafka: Real-time data stream collection and distribution; Apache Storm: Real-time data stream processing;
[0369] Model training:
[0370] Model training: TensorFlow on Kubernetes: distributed model training; Model service: TensorFlowServing: high-performance model inference service;
[0371] Load balancing:
[0372] Load balancing: Utilizing NGINX, a high-performance web server and reverse proxy; Service discovery: Nacos service registration and discovery;
[0373] Auto-scaling: Kubernetes HPA / VPA: Load-based auto-scaling;
[0374] Application services
[0375] Microservice framework: Spring Boot is used;
[0376] This invention also incorporates a security module to enhance security, including:
[0377] Data encryption:
[0378] Transmission Encryption: End-to-end encryption using the TLS 1.3 protocol ensures secure data transmission. For example, when a patient's training data is transmitted from the smart hardware "Scoliosis Training Mirror" to the cloud platform, all data is encrypted using the TLS 1.3 protocol. This means that even if someone intercepts the data in transit, they will not be able to read its content, thus protecting the patient's privacy and sensitive information.
[0379] Storage Encryption: Static data is encrypted using the AES-256-GCM algorithm. For example, when patient training videos and scoring data are stored in a cloud platform database (such as an S3 database), the system uses the AES-256-GCM algorithm to encrypt this data. Even if a hacker manages to gain access to the database, they cannot read the actual content without the decryption key, further protecting the security of patient data.
[0380] Homomorphic encryption: Introducing partial homomorphic encryption allows for limited computational operations on encrypted data, further protecting sensitive data. For example, when a system needs to perform statistical analysis on training data from multiple patients (e.g., calculating average training time or improvement), homomorphic encryption can be used. This allows the system to perform calculations and obtain statistical results without decrypting specific patient data. This protects patient privacy while obtaining valuable data analysis results.
[0381] Identity verification:
[0382] Multi-factor authentication: Combining biometrics (such as fingerprint or facial recognition) with one-time passwords (OTPs) to achieve strong identity authentication. Examples of usage scenarios and methods are as follows:
[0383] When a doctor logs into the system: When a doctor logs into the doctor's terminal for the first time via the WeChat mini program, the system sends a one-time password to the doctor's registered mobile phone number. The doctor needs to enter this password to complete the first login.
[0384] When a patient uses the Magic Mirror for the first time: To ensure that only authorized patients use the device, multi-factor authentication is required upon first use. Method: The patient may need to use facial recognition (via the Magic Mirror camera) and then enter a one-time verification code obtained from a mini-program.
[0385] Role-based access control (RBAC): provides fine-grained control over user access permissions to system functions and data. Examples of its use cases and methods are as follows:
[0386] When medical team members access patient data: Different roles of medical personnel (such as attending physicians, rehabilitation therapists, nurses, etc.) have different access permissions to patient data. Method: The system automatically assigns permissions based on the logged-in user's role. For example, an attending physician can view and modify all patient data, while a rehabilitation therapist may only be able to view training-related data.
[0387] When administrators use the web-based management system: different levels of administrators have different access permissions to system functions. For example, system administrators can set and modify permissions for all users, while content administrators may only be able to update the training video library and popular science content.
[0388] OAuth 2.0 and OpenID Connect: Used for secure user authorization and authentication. Examples of usage scenarios and methods are as follows:
[0389] When patients register and log in via WeChat Mini Program: WeChat's OAuth 2.0 authentication mechanism is used to achieve fast and secure identity verification. Method: The patient clicks the "Log in with WeChat" button. The system uses the OAuth 2.0 protocol to request user authorization from WeChat and obtain necessary user information (such as openid) for identity verification.
[0390] When integrating the system with other healthcare platforms: If integration with a hospital's HIS system or other third-party healthcare service platforms is required, the method is to use the OpenID Connect protocol, allowing users to log in to the system using their existing hospital accounts without re-registration, while ensuring the secure transmission of identity information.
[0391] Privacy protection:
[0392] Data anonymization: Sensitive information is dynamically anonymized; Data minimization principle: Only necessary personal information is collected and processed; Strict data access logs and auditing mechanisms: All data access operations are recorded for easy traceability; Compliance with relevant national medical data protection regulations.
[0393] Figure 2 The overall architecture of the scoliosis posture perception and feedback system is shown, which is divided into five main layers. The specific functions of each layer are as follows:
[0394] 1) Application Layer: Contains multiple user interfaces, specifically:
[0395] Doctor's Mini Program: Doctors can use this mini program to manage patients, view prescriptions, and schedule follow-up appointments.
[0396] Patient-side mini-program: Patients can use this program to receive rehabilitation training prescriptions, view health education content, and track training progress.
[0397] Web backend: Administrators manage system content through the web interface, while data and content for doctors and patients are managed and updated through the web backend;
[0398] Smart Magic Mirror: This is the hardware device of the system, used for real-time feedback and motion detection of patients during rehabilitation training, displaying training videos, and capturing patients' posture data through a camera array.
[0399] 2) Interface layer:
[0400] API Gateway: Handles requests from the application layer and serves as the system's interface management center, forwarding requests for different protocols such as HTTP, REST, and JSON to the business layer through the API Gateway;
[0401] Nginx: As a reverse proxy for the server, it handles the distribution and load balancing of user requests to ensure the efficient operation of the system.
[0402] 3) Business layer: includes services from multiple endpoints:
[0403] Doctor-side services: Provides doctors with functions for managing patients and rehabilitation plans, including modules such as prescription management, patient management, personal center, and data statistics;
[0404] Patient-side services include modules for patient prescription information, training management, and health education, helping patients manage their rehabilitation process;
[0405] Web-based services: Used by doctors and administrators to manage the system backend, including modules such as doctor management, system management, and content management;
[0406] Magic Mirror Terminal Service: Primarily responsible for image and voice processing, including image modules, voice modules, assessment modules, etc., used for real-time monitoring and feedback of patients' training movements;
[0407] 4) Application Service Layer: Supports the core functions of the system and provides the following services:
[0408] Application services: Provide support for standard application functions;
[0409] Model services: including the training and deployment of AI models, and responsible for the optimization of motion recognition and rehabilitation training programs;
[0410] Data encryption: Ensures the security of data transmitted within the system;
[0411] Data caching: Caching data improves system efficiency;
[0412] 5) Basic Service Layer: The system's underlying services, providing data storage and processing capabilities.
[0413] Storage: Responsible for storing patients' training data, images, videos, and other information;
[0414] Data processing: This includes data preprocessing and analysis, supporting data analysis and feedback during the rehabilitation process;
[0415] AI model training: responsible for training machine learning and deep learning models, and optimizing posture analysis models for rehabilitation training;
[0416] 6) Operating environment: The environment on which the system depends for operation, including cloud hosts, Linux, and Docker container technology, used to achieve efficient deployment, expansion, and management of the system;
[0417] 7) Other functions:
[0418] Access control: Controlling the permissions of doctors, patients, and administrators in the system to ensure data security;
[0419] Log recording: Records system operation logs to facilitate subsequent traceability and auditing.
[0420] Figure 3 The workflow of the scoliosis rehabilitation system is described in detail, consisting of multiple steps, each linked sequentially, as follows:
[0421] 1) Generation of doctor's diagnosis QR code:
[0422] The doctor makes an initial diagnosis of the patient, determines their rehabilitation needs, and generates a unique QR code on the doctor's mini-program.
[0423] 2) Patients scan the QR code to register and upload their information:
[0424] Patients can access the patient-side mini-program by scanning a QR code with their mobile phones, fill in their personal information, and upload necessary clinical data (such as X-rays, examination reports, etc.) so that doctors can make further diagnoses.
[0425] 3) The doctor confirms the patient's identity and issues a rehabilitation prescription:
[0426] After verifying the patient's identity and submitted documents, the doctor confirms the diagnosis on their end and issues a personalized rehabilitation training prescription. The prescription includes a specific training plan, frequency, and training programs.
[0427] 4) The patient receives the prescription and begins training:
[0428] After receiving the prescription from the doctor, the patient uses the smart mirror device for rehabilitation training. The mirror plays standardized training videos and reminds the patient to prepare for the training.
[0429] 5) The Magic Mirror plays training videos and performs real-time motion detection:
[0430] The Magic Mirror captures the patient's training movements in real time through a multi-camera setup (including a color camera and a depth camera), and analyzes the patient's movements through an image processing module and AI algorithms to ensure that the movements meet the training requirements.
[0431] 6) AI Analysis and Real-time Feedback:
[0432] The AI module analyzes the patient's posture and movements, and provides voice reminders or visual feedback for non-standard movements to ensure that the patient can adjust their posture in time and avoid incorrect training.
[0433] 7) Uploading training videos, motion evaluation, review and feedback:
[0434] After each training session, the Magic Mirror automatically uploads the training video and related data to the cloud for doctors to view. The system scores the training movements, including the accuracy of posture, range of motion, and speed of movement.
[0435] Doctors review patients' training videos through a doctor-side mini-program and provide feedback based on the patients' performance. Doctors can also use AI to analyze the data and adjust the training plan, providing more suitable rehabilitation suggestions for the patients.
[0436] 8) The system dynamically adjusts the rehabilitation plan and follow-up appointment schedule:
[0437] Based on the patient's training data and the doctor's feedback, the system will dynamically adjust the training plan to ensure the gradual optimization of rehabilitation training and provide personalized rehabilitation suggestions. The system will also generate rehabilitation progress reports based on the patient's training data and remind the doctor and patient when the follow-up appointment date is approaching to ensure the continuity of rehabilitation and the tracking of results.
[0438] Example 1: Scoliosis Correction Training Assessment
[0439] ●Patient Information: An 11-year-old female diagnosed with mild scoliosis (Cobb angle of 20°), and the doctor prescribed a 24-week home rehabilitation training program.
[0440] ●Training content: 30 minutes of training daily (details of the plan will not be disclosed), including: xxx stretching (10 minutes), xx posture (10 minutes), xx movement (5 minutes), xx balance (5 minutes).
[0441] ●Evaluation Process
[0442] ○ Key Angle Assessment: The system measures the patient's Cobb angle, trunk tilt angle, and acromion tilt angle. For example, at week 20, the Cobb angle decreased from an initial 20° to 17°.
[0443] ○Movement Similarity Score: Using the DTW algorithm, the similarity between the patient's movements and the standard movements was compared: the similarity of the xx stretching movement improved from 0.75 in week 1 to 0.92 in week 18.
[0444] ○ Range of motion (ROM) assessment: Measure the scoliosis and rotation ROM of the spine: The scoliosis ROM increased from the initial 25° to 32°.
[0445] ○ Movement speed assessment: Monitor the speed at which the patient completes the movement: The average speed of the xx posture improved from the initial 0.8 times the standard speed to 1.1 times.
[0446] ○ Motion stability assessment: Calculate the standard deviation of acceleration at key points: The stability score for the xx motion improved from 0.6 to 0.85.
[0447] ○ Postural symmetry assessment: Compare the symmetry of key points on the left and right sides: Shoulder symmetry improved from 0.8 to 0.9.
[0448] ○ Duration Assessment: Record the actual execution time of each action: xx balance time increased from the initial 2 minutes to 4.5 minutes.
[0449] ●Overall score: The system calculates the overall score based on the weight of each indicator, which improved from 70 points in week 1 to 88 points in week 24, showing significant progress.
[0450] ●Doctor's Feedback: Based on the system's evaluation results, the doctor adjusted the training plan in week 8, increasing the time spent in the [specific posture] posture and providing specific suggestions for improving the [specific stretching posture].
[0451] Example 2: Scoliosis Correction Training Assessment
[0452] ●Patient Information: A 14-year-old male diagnosed with mild idiopathic scoliosis (thoracic Cobb angle 18°, lumbar Cobb angle 12°). The doctor has developed a 20-week early intervention rehabilitation plan.
[0453] ●Training content: 45 minutes of training per day (details of the plan will not be disclosed), including: A. Posture correction training 1 (15 minutes), B. Core muscle strengthening training 2 (10 minutes), C. Spinal flexibility training 3 (10 minutes), D. Balance training 4 (10 minutes).
[0454] ●Evaluation Process
[0455] ○ Key Angle Assessment: The system measured the patient's Cobb angle, trunk tilt angle, and acromion tilt angle weekly. At week 15, the thoracic Cobb angle decreased from 18° to 16°, and the lumbar Cobb angle decreased from 12° to 10°.
[0456] ○Movement similarity score: Patients' movements were compared with standard movements using the improved DTW algorithm: A. The similarity of trained movements improved from 0.72 in week 1 to 0.83 in week 12.
[0457] ○ Range of motion (ROM) assessment: Measure the ROM of scoliosis, rotation, and flexion of the spine: The ROM of scoliosis increased from the initial 22° to 28°.
[0458] ○ Movement speed assessment: Monitor the patient's speed in completing Training B: The average speed of xx improved from the initial 0.85 times the standard speed to 1.1 times.
[0459] ○Movement stability assessment: Calculate the standard deviation of acceleration at key points during training D: the stability score for xx improved from 0.65 to 0.88.
[0460] ○ Postural symmetry assessment: Compare the symmetry of the left and right sides of the body when standing: The shoulder height difference decreased from the initial 1.5 cm to 0.5 cm.
[0461] ○ Duration Assessment: Record the time the patient was able to maintain each training movement: C. The duration of training movement xx increased from 30 seconds to 90 seconds.
[0462] ○ Quality of life assessment: The patient's quality of life was assessed using the SRS-22 questionnaire: the total score improved from the initial 3.8 points to 4.5 points, especially with a significant improvement in the self-image dimension.
[0463] ●Overall Score: The system calculates the overall score based on the weight of each indicator. The score gradually increases from 75 points in week 1 to 89 points in week 20, showing significant improvement.
[0464] ● Doctor's Feedback: Based on the system's assessment results, the doctor fine-tuned the training plan in week 4, increasing the time spent on A training specific to the patient's scoliosis pattern. In week 8, the doctor introduced more challenging E training programs based on the patient's progress.
[0465] ●Parental Feedback: The patient's parents reported that their child's standing and sitting postures had significantly improved, and the child was more confident in their body image. They also noticed that the child was paying more attention to maintaining correct posture in daily life.
[0466] ●Patient Feedback: The patient stated that the "Scoliosis Training Magic Mirror" made training fun and easy to stick to through real-time feedback and gamification. He particularly liked the system's progress tracking function, which motivated him to continuously strive to improve his performance.
[0467] ●Long-term plan: The doctor plans to continue assessments every 3 months after the training ends, until the patient's bones are fully developed. Simultaneously, the patient is encouraged to integrate the learned posture corrections and core training into daily life to prevent further progression of scoliosis.
[0468] These two embodiments demonstrate how the training effectiveness assessment method of this application can be applied to the early intervention process for adolescents with mild scoliosis. Through comprehensive assessment indicators, the system can accurately track changes in the patient's physical condition and assess improvements in quality of life. This data not only helps doctors adjust rehabilitation plans in a timely manner but also enhances the confidence of patients and their parents, improves treatment adherence, and ultimately achieves the goals of preventing scoliosis progression and improving quality of life.
[0469] The scope of protection of this invention is not limited to the above embodiments. Any variations and advantages that can be conceived by those skilled in the art without departing from the spirit and scope of this invention are included in this invention and are protected by the appended claims.
Claims
1. A scoliosis posture sensing and feedback system, characterized in that, The system includes hardware and software components; The hardware component includes a scoliosis posture sensing and feedback device, which includes a display screen, a multi-camera group, a voice interaction module, and an artificial intelligence module. The software component includes a digital training integrated management platform with mini-programs and web interfaces, as well as a cloud platform that supports data storage and management, data processing and analysis, AI model training and inference, and ensures data security through encryption technology. The display screen is installed on the front of the scoliosis posture sensing and feedback device and is used for content display and image information interaction feedback. The multi-camera group includes a color camera and a depth camera; the depth camera includes an infrared projector and an infrared camera, and is connected to an artificial intelligence module; the multi-camera group is used to collect the user's omnidirectional motion data. The voice interaction module is located on the front of the scoliosis posture sensing and feedback device, and includes a speaker and a microphone, for providing voice information interaction feedback. The artificial intelligence module can analyze and process image and voice information acquired through collection or interaction to perform human posture assessment.
2. The sensing and feedback system as described in claim 1, characterized in that, The color camera has a horizontal field of view of 62.7°, a vertical field of view of 49°, and a diagonal field of view of 75.1°±5°; the infrared camera has a horizontal field of view of 58.4°, a vertical field of view of 45.5°, and a diagonal field of view of 70°±5°; and the infrared projector has a horizontal field of view of 67.26° and a vertical field of view of 51°±5°.
3. The sensing and feedback system as described in claim 1, characterized in that, The digital training integrated management platform includes a doctor-side mini-program, a patient-side mini-program, and a web page management interface. The doctor-side mini-program includes a prescription management module, a patient management module, a personal center, and a data statistics module. The prescription management module allows doctors to dynamically adjust rehabilitation plans based on individual patient differences and treatment progress and send them to patients. The patient management module helps doctors query patients' training records and videos and provides remote movement correction suggestions. The personal center integrates patient, team, and plan management functions to optimize the allocation of doctors' work resources. The data statistics module uses time series analysis and machine learning algorithms to predict future patient traffic and prescription demand, assisting in medical resource planning. The patient-side mini-program covers receiving and viewing rehabilitation prescriptions, improving understanding and execution of prescriptions through visual interpretation; the training management module provides historical training videos, scores, and doctor feedback, and introduces gamification mechanisms, using a points system, progress display, social interaction, and virtual rewards to enhance patients' training motivation; the personal center supports the management of rehabilitation information, follow-up visit plans, and doctor communication, and provides 24 / 7 personalized support through intelligent customer service; The health education module provides information about scoliosis to help patients better understand the disease and treatment process; The web page management terminal includes a doctor management module, a patient management module, a system management module, and a content management module; The doctor management module uses machine learning to evaluate doctor performance, taking into account patient feedback, treatment effectiveness, and workload to ensure fair evaluation; the patient management module integrates medical records and lifestyle habits to provide multi-dimensional patient profiles to support precision treatment. The system management module automatically adjusts the interface layout based on user behavior analysis to improve operational efficiency; the content management module updates the training video library and information, and evaluates content quality through artificial intelligence to ensure its accuracy and effectiveness. And / or, The cloud platform stores patients' basic information, diagnostic records, and training video data, processes complex events, analyzes data streams, supports the training and deployment of AI models, and employs encryption technology to improve data security during data transmission and storage.
4. A method for evaluating human posture training movements, characterized in that, The evaluation method includes the following steps: Step 1: Collect image data of the patient's training movements in real time using a multi-camera system, and preprocess the collected image data to make it meet the model input requirements; Step 2: Extract RGB features and depth features from the preprocessed image data, and fuse the RGB features and depth features into a unified feature representation; Step 3: Perform 2D pose estimation, generate heat maps and location affinity fields (PAFs) for key human body points, determine the 2D coordinates of key points; calculate the depth value of each key point, and combine the 2D pose information to convert the key points into 3D coordinates through geometric transformation to complete 3D pose reconstruction. Step 4: Based on the 3D pose reconstructed in Step 3, conduct a comprehensive motion quality assessment based on motion similarity score, key angle evaluation, range of motion evaluation, motion speed evaluation, motion stability evaluation, posture symmetry evaluation, and duration evaluation.
5. The evaluation method as described in claim 4, characterized in that, In step one, the image data includes an RGB image containing pose information and a depth image containing depth information, and the preprocessing includes image resizing, color space conversion, and normalization; And / or, In step two, ResNet-50 deep convolutional neural network is used to extract features from RGB images and depth images respectively, to obtain two-dimensional visual information and depth information, and then fuse them into three-dimensional spatial feature information; And / or, In step three, based on the multi-stage CNN architecture of OpenPose, the pose features in the image are extracted layer by layer, intermediate results are output at each stage, and the pose estimation is optimized step by step to obtain the two-dimensional coordinates of the key points. In the depth estimation branch, the depth value of each key point is estimated from the fused features through a fully convolutional network, which represents the position of each body part of the user in three-dimensional space. And / or, In step four, the comprehensive motion quality assessment is expressed by the following formula: S total =w1S similarity +w2S ROM +w3S velocity +w4S stability +w5S symmetry +w6S duration , Among them, S similarity S represents the action similarity score. ROM S represents the range of motion score. velocity S represents the speed of movement score. stability S represents the stability score. symmetry S represents the score for postural symmetry. duration The duration score represents the time spent; w1, w2, w3, w4, w5, and w6 represent the weights corresponding to different scores.
6. The evaluation method as described in claim 5, characterized in that, In step one, normalization is performed on the RGB image and depth image to stabilize model training and accelerate convergence; the normalization of the RGB image is expressed by the following formula: Among them, I normalized σ represents the normalized image pixel value, I represents the original image pixel value, μ represents the average value of the image channels, and σ represents the standard deviation of the image channels. The normalization of the depth image is expressed by the following formula: Among them, D normalized D represents the normalized depth value, and D represents the original depth value. min D represents the minimum depth value of all pixels in a depth image. max This represents the maximum depth value of all pixels in the depth image. And / or, In step two, the extracted RGB image features F rgb and depth image features F depth They are expressed by the following formulas respectively: F rgb =ResNet50(I rgb ), F depth =ResNet50(I depth ), The feature F obtained by fusing the RGB image features and the depth image features fused Represented as: F fused =Conv1x1([F rgb ,F depth ]); And / or, In step three, the output of each stage in the multi-stage CNN architecture is represented by the following formula: S t =CNN t (F fused ,S t-1 ), Among them, S t This represents the output of stage t, containing heatmaps of key human body points and location affinity field (PAF) information, CNN. t This represents a convolutional neural network with layer t. The depth estimation result output by the fully convolutional neural network is expressed as follows: D est =FCN(F fused ), Among them, D est The depth estimation result represents the depth value for each keypoint. 3D pose reconstruction transforms the positional information of joints and keypoints in a 2D image into pose information in 3D space, providing precise spatial positional information. 3D pose reconstruction is expressed by the following formula: Where K represents the camera intrinsic parameter matrix, (u,v) represents the coordinates of the two-dimensional keypoint, and z represents the estimated depth; And / or, In step four, the action similarity score calculates the similarity between the user's action and the standard action using a dynamic time warping algorithm; the key angle assessment evaluates joint angles including scoliosis angle, trunk tilt angle, and acromion tilt angle; the range of motion assessment calculates the difference between the range of motion of each joint and the standard range; the action speed assessment calculates the difference between the average speed of the assessed key points and the standard speed; the action stability assessment uses the standard deviation of acceleration to assess action stability; the posture symmetry assessment calculates the symmetry of the assessed left and right key points; and the duration assessment calculates the deviation between the actual execution time of the assessed action and the standard time.
7. The evaluation method as described in claim 4, characterized in that, In the human pose training motion evaluation process, temporal smoothing is used to improve the stability of pose estimation between consecutive frames; and / or, By selecting appropriate optimization methods and training strategies that control the learning rate during the model's convergence process, model performance can be improved; and / or, The loss function defines the objective of model optimization, guiding the model to learn accurate pose estimation and depth prediction; and / or, Post-processing is used to further optimize and correct the model output, ensuring that the output conforms to human physical laws or standards; and / or, Model optimization, including techniques such as quantization and pruning, can improve the model's inference efficiency and resource utilization.
8. The evaluation method as described in claim 4, characterized in that, The method also includes a step of recommending personalized training plans based on the evaluation results, including: Step 1: Collect data including patient training history, physical condition, and doctor's advice, and convert the data into features; Step II: Use a model based on Transformer architecture and multimodal fusion to process input data from different modalities, and combine it with a multi-head self-attention mechanism to process the data; Step III: Fuse information from different modalities through a gating mechanism to generate a comprehensive hidden feature representation; Step IV: Based on the model output, generate a personalized rehabilitation training plan for the patient using reinforcement learning methods, and capture important time points in the patient's rehabilitation process using a dynamic time attention mechanism. And / or, The model includes a feature embedding layer, a temporal encoder, a multi-head self-attention layer, a multimodal fusion layer, a feedforward network, and an output layer component; The feature embedding layer transforms numerical, categorical, or time-series data into feature vector representations suitable for model processing. e i =Embed i (x i ), Where, x i Embed represents different types of input. i Indicates different embedding methods; The timing encoder converts the temporal information in the timing data into input that the model can understand through position encoding; Where pos is the position, i is the dimension, and d is the dimension. model It is the model dimension; The multi-head self-attention layer captures the global dependencies between input features through a self-attention mechanism, thereby achieving information aggregation; MultiHead(Q,K,V)=Concat(head1,...,head h )W O , in It is a learnable parameter matrix, Q, K, V: query, key, and value matrix; d k Indicates the column dimension in the key matrix; The multimodal fusion layer uses a gating mechanism to fuse features from different data modalities to generate a unified comprehensive representation; g i =σ(W g [h1;h2;...;h n ]+b g ), Among them, h i It is a hidden representation of different modalities, g i It is a gating value used to control the information flow of each modality; its dimension is related to each h. i Same; W g It is a learnable weight matrix with dimension [d] out d total ], where d out It is the output dimension, d total b is the total dimension of the hidden representation of all connections; g It is a bias vector with dimension [d out ]; ⊙ represents element-wise multiplication; σ is the activation function that compresses the output to between 0 and 1; h fused This represents the final fused representation, which incorporates information from all modalities. The feedforward network performs further nonlinear transformations on the fused features to generate the model output; FFN(x)=max(0,xW1+b1)W2+b2, Where W1, W2: are learnable weight matrices; b1, b2: are bias vectors; The output layer generates the final recommendation result based on the results of the feedforward network. y=σ(W out h final +b out ), Among them, W out It is a learnable weight matrix; b out It is the bias vector; h final It is the feature representation of the last layer.
9. The evaluation method as described in claim 4, characterized in that, The evaluation method incorporates an anomaly detection model to detect abnormal actions and assess the degree of danger; the anomaly detection includes the following steps: Step 1: Extract the patient's training motion data from the 3D skeletal joint coordinate sequence and preprocess it to adapt to the model input requirements; Step 2: Use a 3D convolutional neural network to extract spatial and temporal features from the input action sequence, capturing local and global information about the action; Step 3: Model the extracted features using a Long Short-Term Memory (LSTM) network to capture the temporal dependencies in the action sequence; Step 4: Use a fully connected layer to classify the output of the LSTM and calculate the probability that the current action is normal or abnormal. Step 5: Calculate the risk level score of the abnormal movement through the regression layer to determine the potential risk of harm to the patient. Step 6: Based on the probability and danger level score of the anomaly detection, the system determines whether the action is safe, abnormal but safe, or dangerous, and provides corresponding feedback. And / or, The anomaly detection model includes an input layer, a feature extraction layer, a sequence modeling layer, a classifier, a regression layer, and an output layer; wherein, 1) Input layer: Receives 3D skeleton joint coordinate sequences; 2) Feature extraction layer: 3D CNN feature extraction: F t =CNN(X t-w:t ), Among them, F t X is the eigenvector at time step t; t-w:t is the input sequence from tw to t; w is the size of the time window; CNN(·) represents the 3D CNN operation; 3) Sequence modeling layer: LSTM sequence modeling: h t ,c t =LSTM(F t ,h t-1 ,c t-1 ), Among them, h t It is the hidden state at time step t; c t It is the cell state at time step t; F t These are features extracted by CNN; LSTM(·) represents LSTM unit operation; 4) Classifier: Anomaly Classification: p(y t =Abnormal|X 1:t )=σ(W c h t +b c ), Where p(y) t =Abnormal|X 1:t W represents the probability that an action is anomalous, given all inputs up to time step t. c It is the classifier weight matrix; h t It is the hidden state at time step t; b c It is the bias term; σ(·) is the sigmoid activation function; 5) Regression layer: Risk assessment: r t =ReLU(W r h t +b r ), Where, r t It is a risk level score for time step t; W r It is the regression layer weight matrix; h t It is the hidden state at time step t; b r It is the bias term; ReLU(·) is the modified linear unit activation function; 6) Output layer: Final decision: Where θ is the threshold for anomaly detection; δ is the threshold for the degree of danger.
10. The application of the perception and feedback system as described in any one of claims 1-3, or the human posture training movement assessment method as described in any one of claims 4-9, in human posture perception during training, personalized rehabilitation training, remote medical support, training anomaly detection and risk prevention and control, and training data management.