Vision-based motion capture system for rehabilitation training

By using a video-based motion capture method, and employing a monocular camera and deep neural networks to assess rehabilitation movements, this approach addresses the issues of high cost and complexity associated with traditional sensor systems, enabling low-cost and efficient rehabilitation training assessment and guidance.

CN116634936BActive Publication Date: 2025-10-31TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180085985.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-04
Filing Date
2021-12-16
Publication Date
2025-10-31
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

Existing sensor-based motion capture systems are costly, complex, and lack scalability in rehabilitation training, resulting in poor rehabilitation outcomes or even physical injury for patients without professional guidance.

Method used

A video-based motion capture method is adopted, which uses a monocular camera and deep neural network to identify key points of the human body, evaluates rehabilitation movements through video data, and generates movement features and scores, reducing reliance on professionals.

Benefits of technology

It enables effective rehabilitation training assessments to be conducted at home or outside of clinics, reduces equipment costs, improves the availability and safety of rehabilitation training, and reduces reliance on professionals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116634936B_ABST
    Figure CN116634936B_ABST
Patent Text Reader

Abstract

This invention includes a video-based motion capture method, apparatus, and storage medium. The method includes: acquiring video data including at least one body part of a person; selecting key points of at least one body part based on a predetermined rehabilitation category; extracting motion features of at least one body part from the video data; scoring the motion features based on the predetermined rehabilitation category; and generating a display illustrating the motion features and the scores of the motion features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to technical solutions for a vision-based motion capture system (VMCS) for rehabilitation training, and in particular to a video-based motion capture method, apparatus, and storage medium. Background Technology

[0002] Physical therapists can provide physical rehabilitation for patients, including the elderly, who may have motor dysfunction-related conditions and / or injuries. However, such rehabilitation training requires the personal supervision of a professional physical therapist, and ongoing guidance from a rehabilitation specialist is also necessary as the patient performs rehabilitation exercises. Thus, training and exercise require diligent monitoring and attention from these professionals and specialists, demanding significant time and effort from the rehabilitation process. Furthermore, even without professionals and specialists, patients training on their own (e.g., at home rather than in a clinic) are expected to lack the professional rehabilitation experience and guidance provided by a physical therapist, potentially leading to poor rehabilitation outcomes or even further physical injury.

[0003] Attempts to automate assessment systems using sensor-based methods (e.g., wearable cameras, infrared cameras) are complex, expensive, and lack scalability. For example, such sensor-based motion capture systems (SMCS) are technically inadequate because they require specific hardware, specialized programming for processing data, and complex requirements for capturing human movement. The high cost of such software and hardware limits their practical application in rehabilitation training. Summary of the Invention

[0004] The present invention includes a video-based motion capture method, apparatus, and storage medium.

[0005] The video-based motion capture method includes: acquiring video data including at least one body part of a person; selecting key points of the at least one body part in the video data based on a predetermined rehabilitation category; extracting motion features of the at least one body part from the video data; determining a score for the motion features based on the predetermined rehabilitation category; and generating a display illustrating the motion features and the scores of the motion features.

[0006] The video-based motion capture device includes: an acquisition module for acquiring video data including at least one body part of a person; a selection module for selecting key points of at least one body part based on a predetermined rehabilitation category; an extraction module for extracting motion features of at least one body part from the video data; a scoring module for scoring the motion features (i.e. determining their scores) based on the predetermined rehabilitation category; and a generation module for generating a display of the illustrated motion features and the scores of the motion features.

[0007] According to an exemplary embodiment, scoring the motion characteristics of at least one body part based on video data includes: determining at least one of key point distance and angle, wherein the key point distance includes Euclidean distance between multiple key points of at least one body part of a person, and the angle includes angle between parts of at least one body part of a person.

[0008] According to an exemplary embodiment, the device further includes a scaling module for scaling at least one body part of a person to a predetermined size based on the person's height, wherein the scoring of movement characteristics based on a predetermined rehabilitation category includes: scoring after scaling at least one body part of the person.

[0009] According to an exemplary embodiment, the device further includes an application module for applying one or more Gaussian filters to motion features.

[0010] According to an exemplary embodiment, the generation and display of the illustrated motion features and the score of the motion features includes: drawing the motion features in the video.

[0011] According to an exemplary embodiment, the selection of key points of at least one body part based on a predetermined rehabilitation category includes: predicting the predetermined rehabilitation category using a deep neural network (DNN), the DNN being configured to predict N possible regions representing possible locations of the key points relative to at least one body part, where N is an integer.

[0012] According to an exemplary embodiment, the step of predicting a predetermined rehabilitation category using a DNN includes: comparing video data including at least one body part of a person with a plurality of anchor poses, each of the plurality of anchor poses including poses of a plurality of predetermined rehabilitation categories, wherein the plurality of predetermined rehabilitation categories include the predetermined rehabilitation category.

[0013] According to an exemplary embodiment, the step of predicting a predetermined rehabilitation category using a DNN includes: sorting N*K pose regions, where K is an integer indicating the number of predetermined key points on the human body.

[0014] According to an exemplary embodiment, the generation of the illustrated motion feature and the display of the score of the motion feature includes: generating a display such that at least one of the anchoring postures is illustrated as covering at least one body part of a person.

[0015] According to an exemplary embodiment, video data of at least one body part of a person includes a red-green-blue (RGB) image of at least one body part of a person obtained by a monocular camera.

[0016] A non-volatile computer-readable storage medium that stores a program that causes a computer to execute a process, the process including the above-described video-based motion capture method. Attached Figure Description

[0017] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0018] Figure 1 This is a simplified illustration of a schematic diagram according to an embodiment.

[0019] Figure 2 This is a simplified illustration of a schematic diagram according to an embodiment.

[0020] Figure 3 This is a simplified illustration of a diagram according to an embodiment.

[0021] Figure 4 This is a simplified illustration of a flowchart according to an embodiment.

[0022] Figure 5 This is a simplified illustration of a diagram according to an embodiment.

[0023] Figure 6 This is a simplified illustration of a diagram according to an embodiment.

[0024] Figure 7 This is a simplified illustration of a diagram according to an embodiment.

[0025] Figure 8 This is a simplified illustration of a diagram according to an embodiment.

[0026] Figure 9 This is a simplified illustration of a diagram according to an embodiment.

[0027] Figure 10 This is a simplified illustration of a diagram according to an embodiment.

[0028] Figure 11 This is a simplified illustration of a flowchart according to an embodiment. Detailed Implementation

[0029] The features discussed below can be used individually or in any combination in any order. Furthermore, embodiments can be implemented using processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-volatile computer-readable medium.

[0030] Figure 1A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is illustrated. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, the first terminal 103 may encode video data locally for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the encoded video data from the other terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission is common in media service applications, etc.

[0031] Figure 1 The illustration shows a second pair of terminals 101 and 104, configured to support bidirectional transmission of encoded video, for example, during video conferencing. For bidirectional data transmission, each terminal 101 and 104 can encode video data captured at a local location for transmission to the other terminal via network 105. Each terminal 101 and 104 can also receive encoded video data sent by the other terminal, decode the encoded data, and display the recovered video data on a local display device.

[0032] exist Figure 1 In this disclosure, terminals 101, 102, 103, and 104 may be illustrated as servers, personal computers, and smartphones, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 105 refers to any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data among terminals 101, 102, 103, and 104. Communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, the architecture and topology of network 105 may be of little importance to the operation of this disclosure, unless otherwise stated below.

[0033] Figure 2 The illustration shows the arrangement of a video encoder and decoder in a streaming environment as an example application of the disclosed subject matter. The disclosed subject matter is equivalent to other video-enabled applications, including, for example, video conferencing, digital TV, and storing compressed video on digital media such as CDs, DVDs, and Memory Sticks.

[0034] The streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, that creates, for example, an uncompressed video sample stream 213. The sample stream 213 may be emphasized as having a high data volume compared to the encoded video stream and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. The encoded video stream 204 may be stored on a streaming server 205 for future use; the encoded video stream 204 may be emphasized as having a lower data volume compared to the sample stream. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video stream 204. Client 212 may include video decoder 211, which decodes an incoming copy of the encoded video stream 208 and creates an outgoing video sample stream 210 that can be displayed on display 209 or other presentation device (not depicted). In some streaming systems, video streams 204, 206, and 208 may be encoded according to certain video encoding / compression standards. Examples of these standards have been pointed out above and are further described herein.

[0035] Figure 3 Example Figure 300 illustrates an image 301 of video captured by at least a monocular camera (such as a smartphone described herein), where such image 301 is a person 311. The exemplary embodiments described herein are configured to identify various joints of a model 312 of the person 311 shown in layer image 302, and as shown in image 303, the model 312 (which can be considered a keypoint estimate) can be superimposed onto the person 311. Image 303 may also include commentary 313, such as suggested physical therapy modalities described herein. Image 303 may also be an image output to the aforementioned smartphone for providing commentary and analysis to a patient. Specific illustrations may be modified in various ways and are not limited to the specific example shown herein.

[0036] Figure 4An exemplary flowchart 400 is illustrated, in which data can be received at S401. According to an exemplary embodiment, such a flowchart 400 can represent any of different scenarios, such as scenarios in a rehabilitation room / clinic and scenarios in a patient's home / outside such a clinic. In a rehabilitation room scenario, the physical therapist may have already instructed the patient to perform rehabilitation movements, after which camera-recorded data of the patient's rehabilitation movements can be obtained, for example, using the camera of their mobile phone. These recorded videos (i.e., the data in S401) can then be uploaded to a server (such as a cloud server and / or a doctor's office computer) for assessing the correctness of the patient's movements, such as S402-S411 described below. Following the assessment process, our system sends the assessment results, along with comments such as 313 described above, to the physical therapist, thereby assisting the physical therapist in performing analyses and diagnoses of the patient's recovery status that would not be possible in other situations. According to an embodiment in a home scenario, the patient can use their mobile phone's camera to record their rehabilitation movements, where, after assessing the correctness of their rehabilitation movements, such a cloud server can directly provide the patient with performance scores and rehabilitation suggestions.

[0037] For example, in such a motion curve generation pipeline, there exists an analysis and visualization of a series of 3D body keypoints as input. For example, at S402, there is 3D human keypoint estimation, such that the human keypoint detection module of such a cloud server is designed to extract keypoints from input RGB images (such as...) Figure 5 The example image 301 and / or image 501 in Figure 500 captures the 3D position of the body joints. Figure 3 Image 302 in the image is an estimated human pose represented by a set of 3D keypoints 312. Additionally, at S402, there is keypoint selection, where specific keypoints can be selected based on the type of rehabilitation movement, such as, for example, in... Figure 5 In image 502, in selection 523, for hand rehabilitation movements, only key points on the hand are selected.

[0038] Such 3D human poses can be estimated using deep neural networks, for example, but not limited to... Figure 6 The exemplary network architecture 600 is shown in the diagram. (As shown in the diagram...) Figure 6As shown, the network architecture 600 includes at least two modules: a region candidate module 611 and a combination of modules 630 and 640 as pose fine-tuning modules. The region candidate module 611 can be used to predict possible regions for each body keypoint. An exemplary embodiment utilizes a pose candidate network to predict N possible regions to represent the possible location of each body keypoint. Assuming a human pose consists of K body keypoints, such an exemplary region candidate network will output N*K pose regions. Further, such pose fine-tuning modules (i.e., modules 630 (classification) and 640 (regression)) are used to rank the pose candidates using classifiers and regressors, and output the final pose estimation result as described herein.

[0039] For example, at S403, motion feature extraction can be implemented, thereby enabling the extraction of one or more motion features based on rehabilitation movement categories. The embodiment provides two motion feature extractors, including a keypoint distance extractor and a keypoint angle extractor. For the keypoint distance extractor, the Euclidean distance between two specific keypoints is calculated. For the angle extractor, the angle between two limbs is calculated, such as the angle between the upper and lower parts of an arm, separated by the elbow. This calculation is performed relative to a neural network and an input image such as that from a smartphone.

[0040] At S404, it can be determined whether motion feature normalization should be preset. This determination may also include determining whether the video shooting conditions are different from one or more preset conditions, making motion feature analysis using motion feature normalization necessary. If so, at S405, normalization can be performed to normalize the patient's proportions and position, scaling the patient's body to a fixed size based on the patient's height.

[0041] At S406, it can be determined whether to pre-determine whether to perform any feature smoothing and denoising. For example, due to possible incorrect pose estimation, some estimated 3D keypoints in the video may be inaccurate, which will result in unsmooth estimated motion features. Thus, if such smoothing and denoising are to be achieved, at S407, the feature smoothing and denoising process can employ a Gaussian filter to smooth the motion features.

[0042] At S407, motion curve visualization is provided, thereby visualizing the patient's analyzed motion through a graph of all motion features in the input video. At S408, performance score estimation is provided, such as customizing the performance score estimator as a deep neural network (DNN) classifier. If it is determined at S410 that such a performance score estimator needs to be trained, it can be trained at S411 using the estimated motion curve. An exemplary embodiment includes a score range of {0, 1, 2, 3, 4} for the performance score estimator.

[0043] In this way, technical improvements have been made to address the technical problems in physical therapy, allowing the embodiments described herein to avoid unrealistic limitations on physical therapy guidance. For example, the goal of the motion curve generation process described herein can be to visualize the patient's motion trajectory in the input video, and the estimated motion curve improves the practicality for any physical therapist in monitoring the patient's recovery status. Furthermore, the motion curve can be used to train a performance rating estimator; given a patient's motion video, the embodiments of this application can advantageously estimate 3D human pose (3D body keypoints) from each frame of the input video, and according to... Figure 4 and Figure 6 The pipeline includes the process of generating motion curves.

[0044] According to an exemplary embodiment, Figure 5 Example 500 is illustrated, showing an image 501 from a camera (such as a smartphone recording video), image 501 including people 511 and 512. According to an exemplary embodiment, at S402, selections 521, 522, and 523 can be made, where selection 521 is selecting the entire person 511, selection 522 is selecting the entire person 512, and as explained above, selection 523, for example, is selecting such a portion of person 511, since it may only specifically analyze some parts of the person.

[0045] See Figure 7 Figure 700 illustrates exemplary poses 701, 702, and 703, but may also include anchoring poses and other poses. Such anchoring poses can be analyzed based on any provided image and the choices made therefrom. For example, as... Figure 8 The feature 800 in the diagram can be analyzed at points 801, 802, and 803 based on poses 701, 702, and 703, respectively. Figure 5 Option 521. Similarly, such as Figure 9 The feature 900 in the diagram can be analyzed at points 901, 902, and 903 based on poses 701, 702, and 703, respectively. Figure 5 Option 522. Additionally, such as... Figure 10 Feature 1000 in the diagram can be analyzed at points 1001, 1002, and 1003 based on poses 701, 702, and 703, respectively. Figure 5 Option 523.

[0046] observe Figure 6 The network architecture in 600, Figure 5 Image 501 can be input data to a simplified network box 621, which can provide features about a combination of region candidate module 611 (localization) and modules 630 (classification) and 640 (regression) as pose fine-tuning modules. For example, as Figure 6 As shown, the orientation 700 is fed into the selection of image 502, respectively, as compared with... Figure 8 , Figure 9 and Figure 10 The features corresponding to these features are 800, 900, and 1000.

[0047] Accordingly, RGB images can be used to determine estimates of 3D human keypoints with estimated 3D human joints. The core functionality of an exemplary embodiment in the VMCS includes three main modules: human keypoint estimation, motion curve generation, and performance score estimation. Accordingly, the exemplary embodiment advantageously utilizes a single camera on a mobile device to capture 3D human motion, greatly alleviating the inconvenience of existing SMCSs, reducing equipment costs, and improving the practicality of physical therapy for healthcare professionals by applying the VMCS to various rehabilitation movements. For example, the goal of the motion curve generation process is to visualize the patient's movement trajectory in the input video, and the estimated motion curves can help physical therapists monitor the patient's rehabilitation status. Furthermore, in addition, the motion curves can be used to train a performance score estimator, such that, given a patient's motion video, the exemplary embodiment advantageously estimates 3D human pose (3D body keypoints) from each frame of the input video, such as based on... Figure 4 and Figure 6 Described.

[0048] Additionally, regarding module 640, features 800, 900, and 1000 can be analyzed to generate data such as graphs 631, 632, and 633 from module 630. For example, classification graph 631 shows that, regarding feature 800, pose 803 is... Figure 5 The pose of person 511 is the most accurate. Similarly, classification chart 632 shows that, regarding feature 900, pose 901 is the most accurate for... Figure 5 The pose of person 512 is the most accurate. Furthermore, classification chart 633 shows that, regarding feature 1000, none of the anchor poses 700 accurately reflects the pose of selection 523.

[0049] The embodiments demonstrate significant technological improvements in vision-based motion capture systems for rehabilitation training, enabling the system to rely solely on a single camera in any mobile device. This greatly improves usability compared to traditional sensor-based motion capture systems, making it more practical for diagnosing and rehabilitating patients with motor dysfunction disorders.

[0050] Thus, according to an exemplary embodiment, there exists a vision-based motion capture system for rehabilitation training that relies solely on a single camera in any mobile device. Compared to conventional sensor-based motion capture systems, this vision-based motion capture system greatly improves usability. Thus, for the embodiments of this application, patients with motor dysfunction can be diagnosed and rehabilitated more practically through a VMCS with a monocular camera for rehabilitation training.

[0051] As described in this application, there may be one or more hardware processors and computer components, such as buffers, arithmetic logic units, and memory instructions, configured to determine or store a predetermined incremental value (difference) between values ​​described in this application according to exemplary embodiments.

[0052] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media, or implemented by one or more specially configured hardware processors. For example, Figure 11 A computer system 1100 is shown that is suitable for implementing certain embodiments of the disclosed subject matter.

[0053] The computer software can be encoded using any suitable machine code or computer language, which can be assembled, compiled, linked, or similarly processed to create code including instructions that can be executed directly or by a computer's central processing unit (CPU), graphics processing unit (GPU), etc., through interpretation, microcode execution, etc.

[0054] The instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0055] Figure 11 The components shown for computer system 1100 are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing embodiments of this application. Nor should the configuration of the components be construed as having any dependency or requirement on any one or combination of components shown in the exemplary embodiments of computer system 1100.

[0056] Computer system 1100 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users via, for example, tactile input (e.g., key presses, swipes, data glove movements), audio input (e.g., voice, taps), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0057] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard 1101, mouse 1102, trackpad 1103, touch screen 1110, joystick 1105, microphone 1106, scanner 1108, and camera 1107.

[0058] Computer system 1100 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback from touchscreen 1110 or joystick 1105, but tactile feedback devices that do not act as input devices may also exist), audio output devices (e.g., speaker 1109, headphones (not depicted)), visual output devices (e.g., screen 1110, including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which are capable of outputting two-dimensional or greater than three-dimensional visual output in a manner such as stereoscopic flat image output; virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).

[0059] The computer system 1100 may also include human-accessible storage devices and associated media of the storage devices, such as optical media, including CD / DVD ROM / RW 1120 with media such as CD / DVD 1111, thumb drives 1122, removable hard disk drives or solid-state drives 1123, legacy magnetic media such as magnetic tapes and floppy disks (not depicted), dedicated devices based on ROM / Application-Specific Integrated Circuits (ASICs) / Programmable Logic Devices (PLDs), such as security protection devices (not depicted), and so on.

[0060] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the currently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0061] Computer system 1100 may also include an interface 1199 to one or more communication networks 1198. Network 1198 may be, for example, wireless, wired, or optical. Network 1198 may also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of network 1198 include, for example, local area networks such as Ethernet and wireless LANs; cellular networks including Global System for Mobile Communications (GSM), 3G, 4G, 5G, and LTE; wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including Controller Area Network Bus (CANBus). Some networks 1198 typically require external network interface adapters attached to certain general-purpose data ports or peripheral buses (1150 and 1151) (e.g., the Universal Serial Bus (USB) port of computer system 1100); other networks are typically integrated into the core of computer system 1100 by attaching to system buses as described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smartphone computer system via a cellular network interface). By using any of these networks 1198, computer system 1100 can communicate with other entities. Such communication can be one-way receiving (e.g., broadcasting TV), one-way transmitting (e.g., a CANBus connected to a CANBus device), or bidirectional, such as connecting to other computer systems using a local area digital network or a wide area digital network. Certain protocols and protocol stacks can be used on each of those networks and network interfaces described above.

[0062] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be attached to the core 1140 of the computer system 1100.

[0063] Core 1140 may include one or more central processing units (CPUs) 1141, graphics processing units (GPUs) 1142, graphics adapters 1117, dedicated programmable processing units 1143 in the form of field-programmable gate areas (FPGAs), hardware accelerators 1144 for certain tasks, and so on. These devices, along with read-only memory (ROM) 1145, random access memory 1146, and internal high-capacity storage devices 1147 such as internal non-user-accessible hard disk drives and solid-state drives (SSDs), can be connected via system bus 1148. In some computer systems, system bus 1148 may be accessed via one or more physical connectors to allow expansion through additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via peripheral bus 1151 to the core's system bus 1148. Architectures for the peripheral bus include peripheral device interconnect (PCI), USB, and so on.

[0064] CPU 1141, GPU 1142, FPGA 1143, and accelerator 1144 can execute certain instructions, which, when combined, constitute the aforementioned computer code. The computer code can be stored in ROM 1145 or RAM 1146. Transient data can also be stored in RAM 1146, while permanent data can be stored, for example, in an internal mass storage device 1147. Fast storage and retrieval of any memory device can be achieved using a cache memory, which can be closely associated with one or more CPUs 1141, GPU 1142, mass storage device 1147, ROM 1145, RAM 1146, etc.

[0065] The computer-readable medium may contain computer code for performing various computer-implemented operations. The medium and computer code may be designed and constructed specifically for the purposes of this application, or may belong to a class well-known and available to those skilled in the art of computer software.

[0066] For example, but not as a limitation, a computer system having architecture 1100, and particularly core 1140, can provide functionality resulting from the execution of software embodied in one or more tangible computer-readable media by a processor (including CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with certain non-transitory storage devices of the user-accessible mass storage described above, as well as core 1140 (e.g., internal mass storage 1147 or ROM 1145). Software implementing various embodiments of this application can be stored in such devices and executed by core 1140. Depending on specific needs, the computer-readable media may include one or more memory devices or chips. The software can cause core 1140, and specifically the processors therein (including CPU, GPU, FPGA, etc.), to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 1146 and modifying such data structures according to processes defined by the software. Alternatively or as an alternative, the computer system may provide functionality generated by logic hardwired or otherwise embodied in circuitry (e.g., accelerator 1144), which may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may cover logic, and vice versa. Where appropriate, references to computer-readable media may cover circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This application covers any suitable combination of hardware and software.

[0067] Although this application describes several exemplary embodiments, various modifications, arrangements, and alternative equivalents are possible within the scope of this application. Therefore, it should be understood that those skilled in the art can design various systems and methods, though not expressly shown or described herein, that embody the principles of this application, within the spirit and scope of the application.

Claims

1. A video-based motion capture method, characterized in that, The method includes: Obtain video data including at least one part of a person's body; In the video data, key points of at least one body part are selected based on a predetermined rehabilitation category; Extract motion features of at least one body part from the video data; Based on the predetermined rehabilitation category, a score for the motor characteristic is determined; and Generate and display the motion features and scores shown in the figure; The step of selecting key points of at least one body part based on a predetermined rehabilitation category includes: predicting the predetermined rehabilitation category using a deep neural network (DNN), wherein the DNN is configured to predict N possible regions representing possible locations of the key points relative to the at least one body part, where N is a positive integer. The step of predicting the predetermined rehabilitation category using a DNN includes comparing the video data, which includes at least one body part of a person, with multiple anchored poses. Each of the plurality of anchoring postures includes a plurality of postures of a predetermined rehabilitation category, wherein the plurality of predetermined rehabilitation categories include the predetermined rehabilitation category.

2. The method according to claim 1, characterized in that, Determining a score for the motion features of at least one body part from the video data includes: determining at least one of keypoint distances and angles, wherein... The keypoint distance includes the Euclidean distance between multiple keypoints of the at least one body part of the person. The angle includes the angle between portions of at least one body part of the person.

3. The method according to claim 2, characterized in that, Further includes: Based on the person's height, at least one body part of the person is scaled to a predetermined size. The determination of the score for the movement characteristic based on the predetermined rehabilitation category includes: performing the score after scaling the at least one body part of the person.

4. The method according to claim 3, characterized in that, Further includes: One or more Gaussian filters are applied to the motion feature.

5. The method according to claim 1, characterized in that, The generation of the display of the motion features and the scores of the motion features includes: drawing the motion features in the video.

6. The method according to claim 1, characterized in that, The step of predicting the predetermined rehabilitation category using a DNN includes: sorting N*K pose regions. Where K is an integer indicating the number of predetermined key points on the human body.

7. The method according to claim 1, characterized in that, The generation of the display illustrating the motion features and the scores of the motion features includes: generating the display such that at least one of the anchoring postures is illustrated as covering the at least one body part of the person.

8. The method according to claim 1, characterized in that, Video data of the at least one body part of the person includes red-green-blue (RGB) images of the at least one body part of the person obtained by a monocular camera.

9. A video-based motion capture device, characterized in that, The device includes: The acquisition module is used to acquire video data including at least one body part of a person. The selection module is used to select key points of the at least one body part based on a predetermined rehabilitation category; An extraction module is used to extract motion features of at least one body part from the video data; A scoring module is used to determine a score for the motor characteristic based on the predetermined rehabilitation category; and A generation module is used to generate a display of the motion features illustrated and the scores of the motion features; The step of selecting key points of at least one body part based on a predetermined rehabilitation category includes: predicting the predetermined rehabilitation category using a deep neural network (DNN), wherein the DNN is configured to predict N possible regions representing possible locations of the key points relative to the at least one body part, where N is an integer. The step of predicting the predetermined rehabilitation category using a DNN includes comparing the video data, which includes at least one body part of a person, with multiple anchored poses. Each of the plurality of anchoring postures includes a plurality of postures of a predetermined rehabilitation category, wherein the plurality of predetermined rehabilitation categories include the predetermined rehabilitation category.

10. The apparatus according to claim 9, characterized in that, Determining a score for the motion features of at least one body part from the video data includes: determining at least one of keypoint distances and angles, wherein... The keypoint distance includes the Euclidean distance between multiple keypoints of the at least one body part of the person. The angle includes the angle between portions of at least one body part of the person.

11. The apparatus according to claim 10, characterized in that, The device further includes a scaling module for scaling at least one body part of the person to a predetermined size based on the person's height. The determination of the score for the movement characteristic based on the predetermined rehabilitation category includes: performing the score after scaling the at least one body part of the person.

12. The apparatus according to claim 11, characterized in that, The device further includes an application module for applying one or more Gaussian filters to the motion feature.

13. The apparatus according to claim 9, characterized in that, The generation of the display of the motion features and the scores of the motion features includes: drawing the motion features in the video.

14. The apparatus according to claim 9, characterized in that, The step of predicting the predetermined rehabilitation category using a DNN includes: sorting N*K pose regions. Where K is an integer indicating the number of predetermined key points on the human body.

15. The apparatus according to claim 9, characterized in that, The generation of the display illustrating the motion features and the scores of the motion features includes: generating the display such that at least one of the anchoring postures is illustrated as covering the at least one body part of the person.

16. A non-volatile computer-readable storage medium storing a program, the program causing a computer to execute a process, characterized in that, The process includes the method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Simulation of physiological functions for monitoring and evaluation of bodily strength and flexibility

    US20200085348A1