Real-time attitude detection and feedback system and method based on computer vision and deep learning
The real-time posture detection and feedback system based on computer vision and deep learning technology solves the problem of traditional rehabilitation training relying on manual guidance, realizes real-time feedback and data recording, and improves the efficiency and quality of rehabilitation training.
Patent Information
- Application Number
- CN202510599464.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-11
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional rehabilitation training methods rely on manual guidance, which is costly and has uneven professional levels. They lack timely and accurate feedback and make it difficult to record and track training data, limiting the quantitative evaluation and optimization of rehabilitation progress.
A real-time posture detection and feedback system based on computer vision and deep learning is used to provide real-time feedback through dual-video comparison. Combined with Mediapipe skeleton tracking and BGR color coding, it can realize the detection and dynamic feedback of elbow flexion and extension angles and trunk offset, and generate a structured rehabilitation assessment report.
It improves the efficiency and success rate of rehabilitation training, lowers the technical threshold, provides scientific and personalized rehabilitation guidance, and enhances user participation and training quality.
Smart Images

Figure CN120766337A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a real-time posture detection and feedback system and method based on computer vision and deep learning, belonging to the technical field of computer vision and deep learning. Background Art
[0002] There is a significant imbalance in the distribution of rehabilitation resources worldwide, and some countries face a severe shortage of rehabilitation therapists. According to statistics from the World Health Organization (WHO), the global density of rehabilitation professionals is less than 10 people per 100,000 population, and in some low-income countries, it is even less than 1 person per 100,000 population. This lack of resources has resulted in a large number of people in need of rehabilitation being unable to obtain professional support. At the same time, as the global aging trend intensifies, it is estimated that by 2050, the number of people over 65 years old will exceed 1.6 billion, which will further exacerbate the contradiction between supply and demand of rehabilitation resources. Therefore, the development of an intelligent, lightweight, and low-cost rehabilitation training system has important research value and practical significance.
[0003] Traditional rehabilitation training methods often rely on manual guidance, which is not only costly but also has varying levels of expertise. This often leads to poor training results for the average person due to a lack of timely and accurate feedback. Furthermore, traditional methods make it difficult to record and track training data, further limiting the ability to quantitatively assess and optimize rehabilitation progress.
[0004] With the rapid development of computer vision and deep learning technologies, applications based on human posture detection have gained widespread attention in areas such as rehabilitation training, sports coaching, and virtual reality. For example, the emergence of open-source libraries such as MediaPipe has made human keypoint detection more efficient and accurate. Integrating these tools, researchers have developed a variety of posture detection applications for sports performance optimization and rehabilitation training assistance. However, existing systems generally suffer from limited functionality and inadequate user experience. Therefore, integrating advanced AI technologies with user needs has become a pressing issue.
[0005] Rehabot's dual-video comparison function allows users to simultaneously monitor demonstration videos and their own movements. It analyzes the user's posture in real time and provides immediate feedback, significantly improving training efficiency. Furthermore, the system supports recording and playback of the entire training process and simultaneously generates detailed training reports containing 12 types of statistical data, providing users with scientific and personalized rehabilitation guidance. This innovative design not only lowers the technical threshold for rehabilitation training but also provides users with more intuitive and comprehensive data support.
[0006] It's worth noting that the design concept of the Rehabot system isn't isolated; it's built on a wealth of related research. For example, according to research by Kishore and Bindu, machine learning techniques can be effectively used to estimate yoga poses, and the performance comparison of different pose models was presented, which ultimately led to the use of the Mediapipe framework in this system. Similarly, Klaus et al. explored the effectiveness of real-time feedback (RTF) in rehabilitation training, demonstrating that such digital feedback systems can significantly improve user exercise execution quality and treatment compliance. These research findings provide theoretical support and technical reference for the development of Rehabot. From a practical application perspective, the Rehabot system is significant in several ways. First, in the field of medical rehabilitation, the system can analyze user training performance, assisting doctors and therapists in developing more precise rehabilitation plans and helping patients recover faster. Second, in the context of sports injury rehabilitation, the Rehabot system can provide real-time movement guidance for upper limb rehabilitation, preventing injuries caused by improper posture from worsening. Furthermore, for the elderly, the system can serve as a low-cost, efficient health management tool to help them maintain physical function.
[0007] The Rehabo system can also provide support for efficient and accurate problem solving in existing research areas. The Angle() and X_Offset() methods in the PoseDetector class enable the system to accurately calculate joint angles and body offsets. This method of combining computer vision with high-precision posture detection directly improves the limitations of traditional rehabilitation training, which relies on manual observation. For example, Smith noted that automated posture detection based on computer vision reduces errors by approximately 30% compared to manual evaluation.
[0008] In summary, the development of the Rehabot system provides a valuable reference for the development of this field. By combining advanced artificial intelligence technology with user-friendly interactive design, it can play an important role in multiple fields such as medical treatment, sports injury rehabilitation, and health management. Summary of the Invention
[0009] The present invention provides a real-time posture detection and feedback system and method based on computer vision and deep learning. The system adopts a humanized real-time feedback mechanism and provides real-time feedback through color coding (green, orange and red). The vivid changes in color help users to more intuitively recognize the current errors, and can help users quickly correct wrong actions, thereby reducing the error rate in rehabilitation training and improving the quality of training. At the same time, the green correct action feedback can also increase the user's confidence in rehabilitation training. Studies have shown that real-time feedback can increase the success rate of rehabilitation training by 25%. The Rehabot system included in the present invention also uses the process_video() method to support the processing of dual video streams. Combined with the threading.Thread method of the Thread open source library, multi-threaded operation of dual videos and GUI interface is realized. It allows users or researchers to compare demonstration videos with user real-time images at the same time, and combined with a rich GUI interface, it will not cause process congestion.
[0010] The technical solution adopted by the present invention to solve its technical problems is: a real-time posture detection and feedback system based on computer vision and deep learning. The system realizes basic guidance and quantitative evaluation of rehabilitation training through the synchronous comparison demonstration of dual video streams and the user's real-time movements, combined with human posture estimation and human body mechanics analysis algorithms. The core functions of the system cover three modules, specifically including: (1) a real-time action guidance system using a dual-window mirror presentation mechanism, which can synchronously play standard rehabilitation videos and user camera images; (2) a human body mechanics parameter evaluation system based on Mediapipe skeleton tracking, which supports the detection and dynamic feedback of key indicators such as elbow flexion and extension angles and trunk symmetry (quantification of the spatial offset of the shoulder-hip horizontal axis), and is named the "AOJ" calculation model; (3) an intelligent data management system that integrates data report generation, full training process recording and progress visualization, can provide instant action correction with graphics through BGR color coding, and generate a structured comprehensive report containing training time trajectory, error statistics and rehabilitation suggestions, providing a traceable quantitative basis for clinical rehabilitation evaluation.
[0011] Furthermore, the present invention includes a dual video comparison module. The dual video comparison module serves as the core functional unit of the system, and its interface layout is as follows: Figure 5 This module uses a dual-window synchronous mirroring presentation mechanism: the left window (light blue border) plays a pre-selected upper limb rehabilitation training demonstration video (Demonstration Video), and the right window (pink border) displays the user's real-time motion capture image (User Screen). The color difference encoding strategy significantly improves visual differentiation. The system integrates the following human motion parameter visualization functions:
[0012] ① Real-time frame rate (FPS): reflects the efficiency of dual-channel video stream processing;
[0013] ② Elbow joint angle (Left / Right): Calculate the bilateral elbow flexion and extension angles based on the Mediapipe skeleton key point detection algorithm;
[0014] ③ Trunk offset (Offset 13-23 / 14-24): Quantifies the degree of spatial offset of the shoulder-hip horizontal axis.
[0015] Furthermore, the system of the present invention adopts a multi-threaded parallel processing architecture and realizes asynchronous rendering of dual video streams (Thread() framework) through an independent thread resource allocation strategy. This design effectively reduces interference between processes and ensures that the 720P video stream maintains a processing rate of ≥14FPS. The superimposed skeleton tracking lines in the figure are derived from the OpenCV-Mediapipe collaborative algorithm, and its color coding mechanism (dynamic color scale in User Screen) intuitively reflects the degree of deviation between the user's action and the standard action range, providing a quantitative basis for clinical rehabilitation evaluation.
[0016] Furthermore, a progress bar is provided below the video window of the system of the present invention to dynamically display the training progress. Rehabot uses the OpenCV constant .get(cv2.CAP_PROP_FRAME_COUNT) to obtain the total number of frames of the training demonstration video, and obtains the real-time progress value through calculation, which is finally reflected in the progress bar control.
[0017] The present invention also provides a method for implementing a real-time posture detection and feedback system based on computer vision and deep learning, specifically including: based on the video display area, Rehabot provides a rich visual experience. In the system of the present invention, the Mediapipe Pose module is used to implement human posture detection. The module extracts human key points (landmarks) through a deep learning model, which usually has 33 key points on the whole body (such as elbows, knees and ankles, etc.). In response to the needs of upper limb rehabilitation assessment, Rehabot uses 8 left-right symmetrical key points on both sides of the acromion (Landmark 11 / 12), elbow (Landmark 13 / 14), wrist (Landmark 15 / 16) and hip (Landmark 23 / 24) to construct a motion evaluation system. ( Figure 7 After the key point coordinates are normalized, the pixel space mapping is realized by formula (1), and the coordinate information of the three key points on the unilateral arm is recorded as (x i ,y i ), stored in the self.lmlist key point list.
[0018]
[0019] where, The normalized coordinates of the key points, (w, h) are the width and height of the image, respectively. This ensures the consistency of the coordinates under different input resolutions.
[0020] Based on the principle of vector geometry, a model for calculating the flexion angle of the elbow joint is constructed Figure 8 ), and the above coordinates are converted into new vectors, defined as:
[0021] v1 = (x1 - x2, y1 - y2) # (2)
[0022] v2 = (x3 - x2, y3 - y2)
[0023] Using trigonometric functions, the cosine similarity between the two vectors is calculated:
[0024]
[0025] To facilitate calculation, the angle value is converted to radian value:
[0026]
[0027] According to the relevant literature of sports and rehabilitation medicine, the maximum extension angle of the human elbow joint is close to 180°, and the allowable angle error range is ±10°. However, in actual movement, due to soft tissue limitations or individual differences, we set the correct action range to 160°-180°. This range not only ensures the standardization of the action (i.e., the arm is straight), but also avoids joint hyperextension injury.
[0028] The above angle is calculated by Rehabot through the calculation of the angle between the shoulder, elbow and wrist, which is used to evaluate whether the arm is straight during the user's training process. In order to further improve the evaluation standard of action standardization, we introduce the calculation method of upper limb X-axis offset offset, which is calculated by (5), as shown in Figure 9 .
[0029] Offest=|x1-x2|#(5) Trunk symmetry is also one of the key indicators for evaluating posture correctness. This offset is used to determine the horizontal axis spatial distance between the shoulder and hip, thereby evaluating whether the arm is straight forward when raised. The horizontal offset (d) between the shoulder and hip on the same side should be less than 3 to 5 cm to ensure the stability of the movement. Considering that the user needs to maintain a highly symmetrical posture during rehabilitation and exercise, while ensuring that the arm is raised straight forward, Rehabot combined the relationship between the camera's horizontal field of view (W) and the camera's horizontal field of view (FOV) mentioned in the official OpenCV documentation, and obtained a typical FOV value of 60° to 90° for consumer-grade cameras. Assuming that the user is 1 meter away from the camera and the maximum horizontal offset degree d=5 cm is taken to ensure that the user has greater fault tolerance, the formula (6) is:
[0030]
[0031] The horizontal field of view (W) can be deduced to be between 60 cm and 100 cm. In the Rehabot system, we take an average value of W = 80 cm.
[0032] Combined with ITU-T H.264, an international standard for video coding widely used in video communication and streaming media transmission, 720p (1280*720) is defined as one of the high-definition (HD) resolutions, which is suitable for most consumer devices (such as laptops, smartphones, webcams, etc.). Therefore, Rehabot selects the horizontal resolution as 720p, and the horizontal resolution (R x ) is 1280 pixels, and it can be derived from formula (7):
[0033]
[0034] The maximum pixel value (P) corresponding to the horizontal axis offset (Offset) between the shoulder and hip is 80 pixels. Therefore, in the Rehabot system, we define the Offset parameter value as less than 80, which is the correct forward movement.
[0035] Furthermore, the Rehabot system of the present invention automatically generates a rehabilitation assessment report that meets basic standards through structured data collection and analysis ( Figure 10Its core functions include: (1) Multi-dimensional data integration: recording 13 core indicators such as training start and end time, demonstration video path, and movement error statistics (elbow joint angle deviation, trunk deviation number); (2) Intelligent scoring mechanism: Based on the weighted algorithm of angle error (±0.2 points / time) and deviation error (±0.1 points / time), it generates a 0-10 point rehabilitation score (four levels of evaluation: excellent / good / fair / needs improvement); (3) Personalized suggestion engine: Automatically generates correction suggestions based on the error type (such as "adjust the left elbow flexion and extension angle"), and supports CSV format export to meet the archiving needs of medical records.
[0036] Furthermore, the system of the present invention adopts a multi-threaded asynchronous processing architecture to achieve holographic recording of the training process: (1) Synchronous recording of dual video streams, by creating an independent video writer (cv2.VideoWriter) to synchronously save the demonstration video stream and the user action stream (including skeletal tracking visualization); (2) Dynamic quality control: Based on the H.264 encoding standard, the video is output in 720P format, and the frame rate dynamic adjustment technology is used to ensure the recording smoothness of ≥25fps. (3) Intelligent storage management: Supports custom storage paths in the system "Playback Settings" module and automatically creates a timestamp directory (%Y_%M_%D%H:%M:%S format).
[0037] The assessment report and video recording are mapped through timestamps, which supports the synchronous viewing of training impact and corresponding stored quantitative assessment data through the playback system (Playback Library) at a later time, providing a multi-module traceability basis for the rehabilitation process.
[0038] Furthermore, the Rehabot system of the present invention uses BRG color coding (green indicates correct, orange indicates adjustment is required, and red indicates error) to classify training movement errors and generate prompt statements for different levels, which will not only be integrated into the report but also be fed back in real time in the corresponding module of the system ( Figure 11 Based on the AOJ calculation model, it also provides guidance on correcting different errors. This feedback mechanism is presented in both graphic and text form in the system's text browser, ensuring that users can intuitively and comprehensively understand the performance of their actions.
[0039] The experimental data source of the system of the present invention adopts the WLURehabilitation Posture dataset released by Wilfrid Laurier University (available at www.kaggle.com), which contains 3 types of standardized rehabilitation training action samples. This study selected its upper limb training subset (Arm Raise Correct) as the standard training paradigm, and excluded sample videos with substandard imaging quality (blurriness > threshold) and human posture occlusion (limb overlap > 30%) through the quality assessment process. The 10 groups of screened valid video examples have been archived to Rehabot's GitHub repository to build a standardized training reference database. The example demonstration screen is as follows Figure 12 In actual operation, you can follow the order of the following points. The following is a detailed introduction:
[0040] The Rehabot system supports two ways to import demonstration video files, such as Figure 13 As shown:
[0041] ①Please Select Video: Import a single MP4 video file.
[0042] ②Please Select Folder: Import a video folder and automatically rotate the videos in it.
[0043] In order to further meet personalized needs, Rehabot also provides live training video recording function, such as Figure 14 As shown. The system records the entire GUI interface and saves the screen recording file in MP4 format to the user path. In the "PlaybackSettings" module:
[0044] ①Enable Playback: Enable the screen recording function.
[0045] ②Playback: Screen recording timer.
[0046] ③Please...Save Path: Click the icon on the right to customize the save path of the screen recording file.
[0047] ④Playback Library: Open the screen recording folder to view the saved video files.
[0048] like Figure 9 As shown, for users who need long-term training, Rehabot provides a flexible data export function that supports exporting training data into CSV format report files for subsequent analysis and summary. Complete the relevant settings in the "Report Settings" section in advance:
[0049] ①Enable Report: Enable data recording and reporting function.
[0050] ②Please Select Report Save Path: Click the right icon to select the report file save path.
[0051] ③Save Report: Save the report, the example content is as shown in Figure 15 .
[0052] When the above settings are completed, go to the "Detection Control" module shown in Figure 16 . It contains four controls. Start: Start training. Stop: Stop training. Reset: Reset the settings options. Help: Help.
[0053] In today's professional rehabilitation resources are scarce, Rehabot with its flexible use of scene, lower cost and clear and simple user experience, for simple rehabilitation training provides a new solution. Combined with computer vision technology, Rehabot not only presents the user with basic but far-reaching rehabilitation training guidance, but also injects the power of technology into the field of rehabilitation, so that everyone can enjoy high-quality rehabilitation services.
[0054] Beneficial effects:
[0055] 1、The Rehabot system of the present application provides a new research direction for the field of rehabilitation training and sports science by combining computer vision (OpenCV and MediaPipe posture detection, etc.) and real-time feedback mechanism. The Rehabot system has the function of personalized rehabilitation training optimization. The system can monitor the elbow posture angle range (Angle_condition) and the lateral offset range of the shoulder and hip bone (Offset_condition) of the user in real time, and dynamically adjust the training intensity and plan according to the physical condition of different users. In the field of multi-modal data fusion, the Rehabot system uses real-time video and real-time data analysis module to provide ideas for future research on the application of multi-modal data in the field of rehabilitation training.
[0056] 2、The Rehabot system of the present application simplifies the operation process of rehabilitation training through an intuitive and easy-to-use user interface and functions. Users only need to equip with basic computer hardware environment, and then through simple button operation, they can perform basic rehabilitation training and exercise body function. Such ease of use greatly reduces the technical threshold, so that non-professionals can easily use the system. In addition, the save_report() method can generate a more detailed training report.
[0057] 3. The Rehabot system of the present invention can provide patients, doctors, and rehabilitation therapists with a certain scientific basis to help them develop more accurate rehabilitation plans. These changes not only improve the training effect, but also enhance user participation and experience.
[0058] 4. The present invention adopts a humanized real-time feedback mechanism, which provides real-time feedback through color coding (green, orange and red). The vivid changes in color help users to more intuitively recognize current errors, and can help users quickly correct wrong actions, thereby reducing the error rate in rehabilitation training and improving the quality of training. At the same time, the green feedback on correct actions can also increase the user's confidence in rehabilitation training. Studies have shown that real-time feedback can increase the success rate of rehabilitation training by 25%. The Rehabot system included in the present invention also uses the process_video() method to support the processing of dual video streams. Combined with the threading.Thread method of the Thread open source library, multi-threaded operation of dual videos and GUI interface is realized. It allows users or researchers to compare demonstration videos with user real-time images at the same time, and combined with the rich GUI interface, it will not cause process congestion. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Schematic diagram of the front end of the Rehabot system of the present invention.
[0060] Figure 2 Schematic diagram of the Rehabot system of the present invention.
[0061] Figure 3 This is a functional diagram of the system of the present invention.
[0062] Figure 4 Flow chart of the method of the present invention.
[0063] Figure 5 Schematic diagram of dual video comparison of the system of the present invention.
[0064] Figure 6 Schematic diagram of a progress bar during operation of the system of the present invention.
[0065] Figure 7 Schematic diagram of the Mediapipe posture points used in the Rehabot system of the present invention.
[0066] Figure 8 Schematic diagram of the bilateral elbow flexion and extension angles (Angle) of the system of the present invention.
[0067] Figure 9 Schematic diagram of the shoulder-hip horizontal axis spatial offset (Offset) of the system of the present invention.
[0068] Figure 10The Report system of the present invention automatically generates a preview image of a rehabilitation assessment report that complies with basic specifications.
[0069] Figure 11 The system of the present invention ensures that the user can intuitively and comprehensively understand the three error classification prompt diagrams of the performance of his own actions.
[0070] Figure 12 An example diagram illustrating the construction of a standardized training reference database for the present invention.
[0071] Figure 13 Schematic diagram of two exemplary video file importing methods supported by the Rehabot system of the present invention.
[0072] Figure 14 This is a schematic diagram of the Playback Settings module of the system of the present invention.
[0073] Figure 15 Schematic diagram of the Report Setting module of the system of the present invention.
[0074] Figure 16 This is a schematic diagram of the Detection Control module of the system of the present invention. DETAILED DESCRIPTION
[0075] The present invention will be described in further detail below with reference to the accompanying drawings.
[0076] like Figures 1-3 As shown in the figure, the Rehabot system of the present invention is a real-time posture detection and feedback system based on computer vision and deep learning technology, which is designed to provide intelligent support for upper limb rehabilitation training, sports guidance and health management. The system of the present invention uses the high-precision human posture key point detection algorithm provided by MediaPipe, combined with the user-friendly graphical interface (GUI) built with PyQt5, to achieve real-time analysis and feedback of user movements. Through dual video stream input, multi-threaded processing and screen recording functions, Rehabot provides a comprehensive interactive experience, suitable for various scenarios such as rehabilitation training, fitness guidance and scientific research analysis.
[0077] like Figure 1 As shown, the present invention takes "modularity" as the core design concept of Rehabot, dividing the entire system into functionally distinct components, each with its own specific tasks. At the same time, the barriers between components are broken down, and through clear interfaces, efficient and accurate communication and control between components are achieved. This is reflected in the front-end interface of Rehabot, which provides users with concise but precise guidance through graphic boxes and text prompts, allowing users of all ages to interact in a friendly manner.
[0078] like Figure 2, the Rehabot system of the present invention adopts a dual-channel multi-threaded processing mechanism during operation, supports dual video stream input, and designs two independent processing logics according to the functional requirements of different video streams. Among them, the processing of demonstration videos is relatively simplified, aiming to optimize the user experience and increase the standardization of training, and provide a reference; while the user's real-time picture becomes the core focus of system processing and analysis. Through the independently developed threshold algorithm, the system can evaluate the user's movements in real time, and provide scientific judgment results and targeted correction suggestions. During use, the system uses the user's picture as the main analysis object, and combines the demonstration video as an auxiliary reference to provide users with a more reliable training basis, thereby effectively enhancing the user's confidence and sense of participation. The Rehabot system of the present invention is an intelligent rehabilitation training aid based on human posture key point detection technology.
[0079] like Figure 3 As shown, the system of the present invention realizes basic guidance and quantitative evaluation of rehabilitation training through the synchronous comparison demonstration of dual video streams and the real-time actions of the user, combined with human posture estimation and human body mechanics analysis algorithms. The core functions of the system cover three major modules: (1) A real-time action guidance system using a dual-window mirror presentation mechanism, which can synchronously play standard rehabilitation videos and user camera images; (2) A human body mechanics parameter evaluation system based on Mediapipe bone tracking, which supports the detection and dynamic feedback of key indicators such as elbow flexion and extension angles, trunk symmetry (quantification of spatial offset of the shoulder-hip horizontal axis), and is named the "AOJ" calculation model; (3) It also has an intelligent data management system that integrates data report generation, full training process recording and progress visualization. It can provide instant action correction with graphics through BGR color coding, and generate a structured comprehensive report containing training time trajectory, error statistics and rehabilitation suggestions, providing a traceable quantitative basis for clinical rehabilitation evaluation. The specific process is as follows: Figure 4 shown.
[0080] The dual video comparison module is the core functional unit of the system, and its interface layout is as follows: Figure 5 This module uses a dual-window synchronous mirroring presentation mechanism: the left window (light blue border) plays a pre-selected upper limb rehabilitation training demonstration video (Demonstration Video), and the right window (pink border) displays the user's real-time motion capture image (UserScreen). The color difference encoding strategy significantly improves visual differentiation. The system integrates the following human motion parameter visualization functions:
[0081] ① Real-time frame rate (FPS): reflects the efficiency of dual-channel video stream processing;
[0082] ② Elbow joint angle (Left / Right): Calculate the bilateral elbow flexion and extension angles based on the Mediapipe skeleton key point detection algorithm;
[0083] ③Torsos Offset 13-23 / 14-24: Quantify the spatial offset of the shoulder-hip horizontal axis.
[0084] The system adopts a multi-thread parallel processing architecture, and realizes asynchronous rendering of double video streams through independent thread resource allocation strategy (Thread() framework). This design effectively reduces the interference between processes and ensures that the processing rate of 720P video stream is ≥14FPS. The superimposed skeleton tracking lines in the figure are derived from the OpenCV-Mediapipe cooperative algorithm, and the color coding mechanism (dynamic color scale in User Screen) intuitively reflects the deviation degree of user action from the standard action range, providing quantitative basis for clinical rehabilitation evaluation.
[0085] A progress bar is provided below the video window, such as Figure 6 , which is used to dynamically display the training progress. Rehabot uses OpenCV constant.get(cv2.CAP_PROP_FRAME_COUNT) to obtain the total frame number of the training demonstration video, and calculates the real-time progress value, which is finally reflected in the progress bar control.
[0086] On the basis of the video display area, Rehabot provides rich visual experience. In this system, the Mediapipe Pose module is used to realize human posture detection, which extracts human key points (landmarks) through a deep learning model, usually with 33 key points (such as elbows, knees and ankles, etc.) throughout the body. For upper limb rehabilitation evaluation needs, Rehabot adopts 8 left-right symmetrical key points including shoulder peak (Landmark 11 / 12), elbow (Landmark 13 / 14), wrist (Landmark 15 / 16) and hip (Landmark 23 / 24) to construct an action evaluation system, Figure 7 After normalization of the key point coordinates, pixel space mapping is realized through formula (1), and the coordinate information of 3 key points on the single arm is recorded as (x i , y i ), stored in the self.lmlist key point list.
[0087]
[0088] wherein, represents the normalized coordinates of the key points, and (w, h) are the width and height of the image respectively, which ensures the consistency of the coordinates under different resolution inputs.
[0089] Based on the principle of vector geometry, an elbow flexion angle calculation model is constructed Figure 8), using the above coordinate points to transform into a new vector, defined as:
[0090] v1=(x1-x2, y1-y2)#(2)
[0091] v2=(x3-x2, y3-y2)
[0092] Using trigonometric functions, calculate the cosine similarity between two vectors:
[0093]
[0094] To facilitate calculation, convert the angle value to radian value:
[0095]
[0096] According to relevant literature on sports and rehabilitation medicine, the maximum extension angle of the human elbow joint is close to 180°, and the allowable angle error range is ±10°
[11] . However, in actual exercise, due to soft tissue limitations or individual differences, we set the correct movement range to 160° to 180°. This range not only ensures the standardization of the movement (i.e., straightening the arm), but also avoids joint hyperextension injuries.
[0097] The above angles are calculated by Rehabot by measuring the angles between the shoulder, elbow, and wrist to evaluate whether the user's arms are straight during training. In order to further improve the evaluation criteria for movement standardization, we introduced a method for calculating the upper limb X-axis offset, which is calculated by (5), as follows: Figure 9 As shown:
[0098] Offest=|x1-x2|#(5) Trunk symmetry is also one of the key indicators for evaluating posture correctness. The offset is used to determine the horizontal axis spatial distance between the shoulder and hip, thereby evaluating whether the arm is straight forward when raised. The horizontal offset (d) between the shoulder and hip on the same side should be less than 3 to 5 cm to ensure the stability of the movement. Considering that the user needs to maintain a highly symmetrical posture during rehabilitation and exercise, while ensuring that the arm is raised straight forward, Rehabot combined the relationship between the camera's horizontal field of view (W) and the camera's horizontal field of view (FOV) mentioned in the official OpenCV documentation
[12] , and obtained a typical FOV value of 60° to 90° for consumer-grade cameras. Assuming that the user is 1 meter away from the camera and the maximum horizontal offset degree d=5 cm is taken to ensure that the user has greater fault tolerance, the formula (6) is:
[0099]
[0100] The horizontal field of view (W) can be deduced to be between 60 cm and 100 cm. In the Rehabot system, we take an average value of W = 80 cm.
[0101] Combined with ITU-T H.264, an international standard for video coding widely used in video communication and streaming media transmission, 720p (1280*720) is defined as one of the high-definition (HD) resolutions, which is suitable for most consumer devices (such as laptops, smartphones, webcams, etc.). Therefore, Rehabot selects the horizontal resolution as 720p, and the horizontal resolution (R x ) is 1280 pixels, and it can be derived from formula (7):
[0102]
[0103] The maximum pixel value (P) corresponding to the horizontal axis offset (Offset) between the shoulder and hip is 80 pixels. Therefore, in the Rehabot system, we define the Offset parameter value as less than 80, which is the correct forward movement.
[0104] The Rehabot system of the present invention automatically generates a rehabilitation assessment report that meets basic standards through structured data collection and analysis ( Figure 10 Its core functions include: (1) Multi-dimensional data integration: recording 13 core indicators such as training start and end time, demonstration video path, and movement error statistics (elbow joint angle deviation, trunk deviation number); (2) Intelligent scoring mechanism: Based on the weighted algorithm of angle error (±0.2 points / time) and deviation error (±0.1 points / time), it generates a 0-10 point rehabilitation score (four levels of evaluation: excellent / good / fair / needs improvement); (3) Personalized suggestion engine: Automatically generates correction suggestions based on the error type (such as "adjust the left elbow flexion and extension angle"), and supports CSV format export to meet the archiving needs of medical records.
[0105] The system of the present invention uses a multi-threaded asynchronous processing architecture to achieve holographic recording of the training process: (1) Synchronous recording of dual video streams, by creating an independent video writer (cv2.VideoWriter) to synchronously save the demonstration video stream and the user action stream (including skeletal tracking visualization); (2) Dynamic quality control: Based on the H.264 encoding standard, the video is output in 720P format, and the frame rate dynamic adjustment technology is used to ensure the recording smoothness of ≥25fps. (3) Intelligent storage management: Supports custom storage paths in the system "Playback Settings" module and automatically creates a timestamp directory (%Y_%M_%D%H:%M:%S format).
[0106] The evaluation report and the video recording are mapped by time stamp, supporting the later synchronization viewing of the training influence and the corresponding stored quantitative evaluation data through the playback system (Playback Library), and providing multi-module traceability basis for the rehabilitation process.
[0107] The Rehabot system grades the training action errors by BRG color coding (green for correct, orange for adjustment required, and red for error), generates prompt sentences for different grades, which are not only integrated into the report, but also real-time feedback in the corresponding modules of the system Figure 11 ). At the same time, based on the AOJ calculation model, correction guidance is provided for different errors. This feedback mechanism is presented in the form of text and pictures in the text browser of the system, ensuring that users can intuitively and comprehensively understand their own action performance, as shown in Figure 11 .
[0108] The operation guide of the system is detailed in the README.txt document in the software package. The experimental data source uses the WLU Rehabilitation Posture dataset published by Wilfrid Laurier University (available at www.kaggle.com), which contains 3 types of standardized rehabilitation training action samples. This study selects the upper limb training subset (ArmRaise Correct) as the standard training paradigm, and excludes samples with poor imaging quality (blur>threshold) and human posture occlusion (limb overlap>30%) through quality evaluation process. The 10 sets of effective video examples selected after screening have been archived in the GitHub repository of Rehabot, building a standardized training reference database. The example demonstration screen is shown in Figure 12 . In actual operation, the operation can be performed in the following order, and the following is a detailed introduction:
[0109] The Rehabot system supports two ways of importing demonstration video files, as shown in Figure 13 .
[0110] ①Please Select Video: Import a single MP4 video file.
[0111] ②Please Select Folder: Import a video folder, automatically cycle through the videos in it.
[0112] In order to further meet the individual needs, Rehabot also provides a training live video recording function, as shown in Figure 14 . The system records the complete GUI interface and saves the recording file in MP4 format to the user path. In the "PlaybackSettings" module:
[0113] ①Enable Playback: Enable the screen recording function.
[0114] ②Playback: Screen recording timer.
[0115] ③Please...Save Path: Click the icon on the right to customize the save path of the screen recording file.
[0116] ④Playback Library: Open the screen recording folder to view the saved video files.
[0117] like Figure 9 As shown, for users who need long-term training, Rehabot provides a flexible data export function that supports exporting training data into CSV format report files for subsequent analysis and summary. Pre-complete the relevant settings in the "Report Settings" section:
[0118] ①Enable Report: Enable data recording and reporting functions.
[0119] ②Please Select Report Save Path: Click the icon on the right and select the report file save path.
[0120] ③Save Report: Save the report, the sample content is as follows Figure 15 shown.
[0121] When the above settings are completed, come to Figure 16 The Detection Control module is shown. It contains four controls. Start: Starts training. Stop: Stops training. Reset: Resets the settings. Help: Help.
[0122] In today's world of scarce professional rehabilitation resources, Rehabot offers a new solution for simple rehabilitation training with its flexible use cases, lower costs, and clear, concise user experience. Incorporating computer vision technology, Rehabot not only provides users with basic yet profound rehabilitation training guidance, but also injects the power of technology into the rehabilitation field, ensuring that everyone can enjoy high-quality rehabilitation services.
[0123] The Rehabot system of the present invention provides a new research direction for the fields of rehabilitation training and sports science by combining computer vision (OpenCV and MediaPipe posture detection, etc.) and real-time feedback mechanism. The Rehabot system has personalized rehabilitation training optimization. The system can monitor the user's elbow posture angle range (Angle_condition) and the lateral offset range (Offset_condition) of the shoulder and hip in real time. In order to dynamically adjust the training intensity and plan according to the physical conditions of different users, scholar Liu has also proposed a certain scheme design in similar research. The Rehabot system also provides a reference algorithm and data range for this. In the field of multimodal data fusion, the Rehabot system uses real-time video and real-time data analysis modules to provide a concept for future research on the application of multimodal data in the field of rehabilitation training. The Rehabot system of the present invention simplifies the rehabilitation training operation process through an intuitive and easy-to-use user interface and functions. Users only need to be equipped with a basic computer hardware environment and then perform basic rehabilitation training and exercise body functions through simple button operations. This kind of ease of use greatly reduces the technical threshold, making it easy for non-professionals to use the system. In addition, the save_report() method can generate more detailed training reports. The Rehabot system of the present invention can provide patients, doctors, and rehabilitation therapists with a certain scientific basis to help them formulate more accurate rehabilitation plans. These changes not only improve the training effect, but also enhance user participation and experience.
[0124] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, the division of the above-mentioned functional units and modules is only used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system described in this application is divided into different functional units or modules to complete all or part of the functions described above.
Claims
1. A real-time posture detection and feedback system based on computer vision and deep learning, characterized in that: The system uses dual video streams for simultaneous comparison demonstration and user real-time movements, combined with human posture estimation and human body mechanics analysis algorithms, to achieve basic guidance and quantitative evaluation of rehabilitation training. The core functions of the system cover three major modules, including: (1) a real-time action guidance system using a dual-window mirror presentation mechanism, which can simultaneously play standard rehabilitation videos and user camera images; (2) The human body mechanical parameter evaluation system based on Mediapipe skeleton tracking supports the detection and dynamic feedback of key indicators such as elbow flexion and extension angle, trunk symmetry (quantification of spatial offset of the shoulder-hip horizontal axis), and is named the "AOJ" calculation model; (3) It also has an intelligent data management system that integrates data report generation, full training process recording and progress visualization. It can provide instant action correction through BGR color coding and generate a structured comprehensive report containing training time trajectory, error statistics and rehabilitation suggestions, providing a traceable quantitative basis for clinical rehabilitation evaluation.
2. A real-time posture detection and feedback system based on computer vision and deep learning according to claim 1, characterized in that: The system includes a dual video comparison module; The dual-video comparison module serves as the core functional unit of the system. Its interface layout adopts a dual-window synchronous mirroring presentation mechanism: the left window (light blue border) plays a pre-selected upper limb rehabilitation training demonstration video (Demonstration Video), and the right window (pink border) displays the user's real-time motion capture image (UserScreen). The color difference encoding strategy significantly improves visual differentiation. The system integrates the following human motion parameter visualization functions: ① Real-time frame rate (FPS): reflects the efficiency of dual-channel video stream processing; ② Elbow joint angle (Left / Right): Calculate the bilateral elbow flexion and extension angles based on the Mediapipe skeleton key point detection algorithm; ③ Trunk offset (Offset 13-23 / 14-24): Quantifies the degree of spatial offset of the shoulder-hip horizontal axis.
3. A real-time posture detection and feedback system based on computer vision and deep learning according to claim 1, characterized in that: The system adopts a multi-threaded parallel processing architecture and realizes asynchronous rendering of dual video streams (Thread() framework) through an independent thread resource allocation strategy.
4. A real-time posture detection and feedback system based on computer vision and deep learning according to claim 1, characterized in that: A progress bar is provided below the system video window to dynamically display the training progress. Rehabot uses the OpenCV constant .get(cv2.CAP_PROP_FRAME_COUNT) to obtain the total number of frames of the training demonstration video, and obtains the real-time progress value through calculation, which is finally reflected in the progress bar control.
5. The real-time posture detection and feedback system based on computer vision and deep learning according to claim 1, characterized in that: The system automatically generates a rehabilitation assessment report that meets basic standards through structured data collection and analysis, including: (1) multi-dimensional data integration: recording 13 core indicators such as training start and end time, demonstration video path, and movement error statistics (elbow joint angle deviation, trunk offset number); (2) intelligent scoring mechanism: based on a weighted algorithm of angle error (±0.2 points / time) and offset error (±0.1 points / time), a 0-10 point rehabilitation score is generated (four levels of assessment: excellent / good / fair / needs improvement); (3) personalized suggestion engine: automatically generates correction suggestions ("adjust the left elbow flexion and extension angle") according to the error type, and supports CSV format export to meet the archiving needs of medical records.
6. The real-time posture detection and feedback system based on computer vision and deep learning according to claim 1, characterized in that: The system adopts a multi-threaded asynchronous processing architecture to achieve holographic recording of the training process: (1) synchronous recording of dual video streams, by creating an independent video writer (cv2.VideoWriter) to synchronously save the demonstration video stream and the user action stream (including skeleton tracking visualization); (2) dynamic quality control: based on the H.264 encoding standard, the video is output in 720P format, and the frame rate dynamic adjustment technology is used to ensure the basic recording smoothness of ≥25fps; (3) intelligent storage management: support for customizing the storage path in the system "Playback Settings" module, automatically create a timestamp directory (%Y_%M_%D%H:%M:%S format), and establish a mapping relationship between the evaluation report and the video recording through the timestamp. It supports the later synchronous viewing of the training impact and the corresponding stored quantitative evaluation data through the playback system (Playback Library), providing a multi-module traceability basis for the rehabilitation process.
7. The real-time posture detection and feedback system based on computer vision and deep learning according to claim 1, characterized in that: The system uses BRG color coding (green indicates correct, orange indicates adjustment is required, and red indicates error) to classify training movement errors and generate prompt statements for different levels. These prompt statements will not only be integrated into the report, but will also be fed back in real time in the corresponding module of the system. Based on the AOJ calculation model, correction guidance will be provided for different errors. This feedback mechanism is presented in the system's text browser in both graphic and text forms to ensure that users can intuitively and comprehensively understand the performance of their own movements.
8. A method for implementing a real-time posture detection and feedback system based on computer vision and deep learning, characterized in that include: The Mediapipe Pose module is used to implement human posture detection. The Mediapipe Pose module extracts human key points (landmarks) through a deep learning model. It usually has 33 key points on the whole body. To meet the needs of upper limb rehabilitation assessment, Rehabot uses 8 bilaterally symmetrical key points on both sides of the acromion (Landmark 11 / 12), elbow (Landmark 13 / 14), wrist (Landmark 15 / 16) and hip (Landmark 23 / 24) to build a motion assessment system. After the coordinates of the key points are normalized, the pixel space mapping is realized by formula (1). The coordinate information of the three key points on the unilateral arm is recorded as (x i ,y i ), stored in the self.lmlist key point list: in, Represents the normalized coordinates of the key points, (w,h) are the width and height of the image respectively, which ensures the coordinate consistency under different resolution inputs; Based on the principle of vector geometric relationship, a calculation model for the elbow flexion and extension angle is constructed, and the above coordinate points are converted into new vectors, which are defined as: v1=(x1-x2, y1-y2)#(2) v2=(x3-x2, y3-y2) Using trigonometric functions, calculate the cosine similarity between two vectors: To facilitate calculation, convert the angle value to radian value:
9. The method for implementing a real-time posture detection and feedback system based on computer vision and deep learning according to claim 8, characterized in that: The method sets the correct motion range to 160° to 180°; The above angles are calculated by the Rehabot system between the shoulder, elbow, and wrist to evaluate whether the user's arms are straight during training. In order to further improve the evaluation criteria for movement standardization, the calculation method of the upper limb X-axis offset is introduced, which is calculated by (5) as follows: Offest=|x1-x2|#(5) Trunk symmetry is also one of the key indicators for evaluating posture correctness. The offset is used to determine the horizontal axis spatial distance between the shoulder and hip, thereby evaluating whether the arm is raised straight forward. The horizontal offset (d) between the shoulder and hip on the same side should be less than 3 to 5 cm to ensure the stability of the movement. Considering that the user needs to maintain a highly symmetrical posture during rehabilitation and exercise, while ensuring that the arm is raised straight forward, Rehabot combined with the official OpenCV documentation about the relationship between the camera's horizontal field of view (W) and the camera's horizontal field of view (FOV), and obtained a typical FOV value of 60° to 90° for consumer-grade cameras. Assuming that the user is 1 meter away from the camera, and taking the maximum horizontal offset degree d = 5 cm to ensure that the user has greater fault tolerance, its formula (6) is: It can be deduced that the horizontal field of view (W) is 60 cm to 100 cm, and in the Rehabot system, the average value is W=80 cm.
10. The method for implementing a real-time posture detection and feedback system based on computer vision and deep learning according to claim 8, characterized in that: Combined with the ITU-TH.264 video coding international standard widely used in video communication and streaming media transmission, 720p (1280*720) is defined as one of the high-definition (HD) resolutions, which is suitable for most consumer-grade devices. Therefore, Rehabot selects the horizontal resolution as 720p. The horizontal resolution (R x ) is 1280 pixels, and it can be derived from formula (7): The maximum pixel value (P) corresponding to the horizontal axis offset (Offset) between the shoulder and hip is 80 pixels. In the Rehabot system, when the Offset parameter value is less than 80, it is defined as a correct forward extension movement.