Motion standardization evaluation method and system based on human motion prediction algorithm
Through the method based on the human motion prediction algorithm, the human posture estimation algorithm and prediction model can achieve standardized motion evaluation without wearing equipment, which solves the problems of high cost and low accuracy of wearable sensors, improves the convenience and accuracy of motion capture and evaluation, and promotes the intelligence of rehabilitation training and diagnosis.
Patent Information
- Application Number
- CN202510276317.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, wearable sensor devices are cost-effective to use, are inconvenient to wear, and are difficult to accurately capture and identify individual unique motion characteristics, resulting in the inability to achieve standardized motion judgments, which affects training efficiency and safety.
Using a method based on human motion prediction algorithm, the human posture estimation algorithm and human motion prediction model are used to normalize the motion evaluation using historical and real moving image sequences. There is no need for wearable devices. The prediction model is used to generate multiple future motion postures and calculate the average joint error to judge the motion normativeness.
It realizes accurate capture and personalized evaluation of movements, improves the convenience and popularity of use, reduces costs, can monitor and guide movement norms in real time, and improves the rehabilitation training effect and diagnosis accuracy.
Smart Images

Figure CN120279593A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of human motion prediction, and particularly to a method and system for evaluating the standardization of actions based on a human motion prediction algorithm. Background Art
[0002] In daily physical exercises, work, and study activities, inappropriate or non-standard human actions often lead to problems such as low training efficiency, muscle strains, and cervical spondylosis. Especially in some sports training programs and tests, incorrect movement patterns are extremely likely to cause physical injuries, and at the same time, it will also result in unqualified results, poor training effects, muscle damage, and cervical spondylosis. In order to detect the standardization of human actions, most implementation solutions mainly use wearable devices to collect data or are based on the field of human action detection.
[0003] Users can wear complex sensor devices to capture motion data. However, there are significant limitations in this process. The most prominent problem is that although these technologies can collect a certain amount of motion information, they rarely can make intelligent judgments and provide immediate feedback on whether these actions are standard. For the vast majority of user groups, wearable sensor devices not only have high costs, increasing the usage cost, but also are extremely inconvenient to wear in daily use, which greatly limits the popularization and application scope of these technologies. More critically, due to the lack of effective algorithms and mechanisms, even if motion data is collected, the standardized judgment cannot be directly achieved, and users often cannot accurately understand the deficiencies and improvement directions in their actions, which undoubtedly weakens the practical significance of technical assistance in training and improving effects.
[0004] In the field of human action detection, in most cases, the detection and evaluation system based on human actions needs to compare with the action data of other people stored in the database. This method does improve the generalization ability of the system to a certain extent because it can refer to a wide range of data sets for action recognition and evaluation. However, when we apply this comparison method to a single individual, problems gradually emerge. Each person's movement pattern is unique, and there may be significant individual differences in terms of movement amplitude, speed, or rhythm. Therefore, simply relying on the comparison with the standard or average data in the database often makes it difficult to accurately capture and identify the unique action characteristics of an individual. In this case, the system is prone to false detection or missed detection. The limitations and deficiencies of this method based on database comparison are particularly obvious when dealing with action data with distinct individual characteristics. Therefore, it is particularly important to explore a more convenient, economical, and intelligent judgment function new technical solution. Summary of the Invention
[0005] The present application provides a method for evaluating action standardization based on a human motion prediction algorithm to solve the problems in the prior art that wearable sensor devices are costly to use, inconvenient to wear, and direct implementation of standardized evaluation is impossible, and it is difficult for the human action detection and evaluation system to accurately capture and identify the unique action features of individuals.
[0006] Correspondingly, the present application also provides an action standardization evaluation system based on a human motion prediction algorithm, an electronic device, and a computer-readable storage medium for ensuring the implementation and application of the above method.
[0007] To solve the above technical problems, the present application discloses a method for evaluating action standardization based on a human motion prediction algorithm, and the method includes:
[0008] Inputting the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose;
[0009] Inputting the historical patient motion pose into a pre-trained human motion prediction model to obtain multiple predicted future patient motion poses;
[0010] Calculating the average joint error for each frame between the predicted future patient motion pose and the real future patient motion pose, and determining whether the patient's motion conforms to the standard according to the error.
[0011] The present application also discloses an action standardization evaluation system based on a human motion prediction algorithm, and the system includes:
[0012] A human pose detection module for inputting the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose;
[0013] A human pose prediction module for inputting the historical patient motion pose into a pre-trained human motion prediction model to obtain multiple predicted future patient motion poses;
[0014] A standard evaluation module for calculating the average joint error for each frame between the predicted future patient motion pose and the real future patient motion pose, and determining whether the patient's motion conforms to the standard according to the error.
[0015] The present application also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements one or more of the methods in the present application when executing the program.
[0016] The present application also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods described in one or more of the present application are implemented.
[0017] In the present application, the historical patient motion image sequence and the corresponding real future patient motion image sequence are respectively input into a preset human pose estimation algorithm to obtain the historical patient motion pose and the real future patient motion pose. Without wearing devices, the convenience and popularity of use are greatly improved, and the cost is reduced. Then, the historical patient motion pose is input into a pre-trained human motion prediction model to obtain various predicted future patient motion poses. Finally, the average joint error is calculated for each frame between the predicted future patient motion pose and the real future patient motion pose, and it is determined whether the patient's motion complies with the specification according to the error. By predicting the motion pose of each patient through the human motion prediction model and comparing it with the real future patient motion pose, accurate capture and personalized evaluation of the patient's actions are realized. The present application can capture and analyze the action information of patients in real time, and can realize the standardized evaluation and guidance of human actions in various medical scenarios such as rehabilitation training conditions.
[0018] Additional aspects and advantages of the present application will be given in the following description section, which will become apparent from the following description or be understood through the practice of the present application. Description of the Drawings
[0019] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0020] Figure 1 is a flowchart of the action standardization evaluation method based on the human motion prediction algorithm provided by the embodiment of the present application;
[0021] Figure 2 is a modified voxelPose model diagram provided by the embodiment of the present application;
[0022] Figure 3 is the basic structure diagram of the human motion prediction model provided by the embodiment of the present application;
[0023] Figure 4 is the basic structure diagram of the encoder provided by the embodiment of the present application;
[0024] Figure 5 is the cross-attention layer structure diagram provided by the embodiment of the present application;
[0025] Figure 6 is the basic structure diagram of the decoder provided by the embodiment of the present application;
[0026] Figure 7The overall flowchart of the action standardization evaluation method provided by the embodiments of this application;
[0027] Figure 8 The structural schematic diagram of the action standardization evaluation system based on the human motion prediction algorithm provided by the embodiments of this application;
[0028] Figure 9 The structural schematic diagram of the electronic device provided by the embodiments of this application. Detailed implementation manners
[0029] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described by referring to the drawings below are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.
[0030] Those skilled in the art of this technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0031] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0032] The solution provided by the embodiments of the present application can be executed by any electronic device. For example, it can be a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this. Regarding the technical problems existing in the prior art, the action normalization evaluation method and system based on the human motion prediction algorithm provided by the present application aim to solve at least one of the technical problems in the prior art.
[0033] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in conjunction with the accompanying drawings.
[0034] The embodiments of the present application provide a possible implementation manner, such as Figure 1 shown, a flowchart of an action normalization evaluation method based on a human motion prediction algorithm is provided. This solution can be executed by any electronic device. Optionally, it can be executed on the server side or the terminal device.
[0035] As Figure 1 shown, the method may include the following steps:
[0036] Step 101, input the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose;
[0037] Step 102, input the historical patient motion pose into a pre-trained human motion prediction model to obtain multiple predicted future patient motion poses;
[0038] Step 103, calculate the average joint error for each frame between the predicted future patient motion pose and the real future patient motion pose, and judge whether the patient's motion conforms to the specification according to the error.
[0039] In the embodiments of the present application, based on the image sequence of the original patient's motion, the future motion is predicted through the human pose estimation algorithm and the human motion prediction model used in this method, and then it is judged whether there is a problem with the patient's motion.
[0040] In the medical field, some minor differences in a patient's movement may indicate subtle changes in their physical condition or potential health problems. These differences, whether it is a slight stagger in gait, incoordination in movements, or a subtle decrease in strength, can all be early warning signals sent by the body. For example, unsteady gait may imply some abnormality in the nervous system or muscular system; incoordination in movements may be related to a disorder of balance or the initial manifestation of a certain movement disorder; and a gradual decrease in strength may be associated with various causes such as muscle degeneration and neuromuscular junction problems. Therefore, doctors and caregivers need to closely monitor these subtle changes in the patient's movement in order to detect and take appropriate treatment measures early, thereby effectively managing and improving the patient's health condition.
[0041] In the embodiments of the present application, a preset human pose estimation algorithm and a human movement prediction model are used to evaluate these movements in real time in a standardized manner. By comparing the differences between the patient's actual movements and the standard rehabilitation movements, the non-standard parts in the patient's movements can be quickly identified, thereby providing targeted guidance and suggestions. This not only helps to improve the rehabilitation effect but also effectively avoids secondary injuries caused by improper movements during the patient's rehabilitation process. Similarly, when patients perform movements during rehabilitation training, they need to strictly follow the established professional norms and standards. These norms not only ensure the safety and effectiveness of medical services but also promote the smooth progress of the patient's rehabilitation process. Therefore, the method proposed in the embodiments of the present application can monitor and evaluate every operation method of medical staff and every rehabilitation movement of patients in real time, which has important and far-reaching practical application value. This method can not only help medical staff timely discover and correct improper operations, improve professional skills and service quality, but also provide personalized rehabilitation guidance for patients to ensure that their rehabilitation movements are correct, accelerate the rehabilitation process, and improve the rehabilitation effect.
[0042] Based on the method in the embodiments of the present application, it is helpful for disease diagnosis and evaluation. By capturing and analyzing the movement information of patients, doctors can evaluate their movement quality and functional status. Especially for movement dysfunctions caused by neurological diseases and musculoskeletal diseases, it provides a more accurate diagnostic basis for movement judgment. In the aspect of ward monitoring, by judging the movements of patients, the computer can give real-time responses to abnormal behaviors and notify doctors in a timely manner, which can save every minute and second for treating patients. At the same time, these data can also be used for formulating subsequent treatment and rehabilitation plans. Meanwhile, in the aspect of rehabilitation training and treatment, the movement judgment technology can formulate personalized rehabilitation plans according to the specific movement performance of patients and give real-time feedback and guidance during the training process, thereby improving the accuracy and effectiveness of the training. In addition, the application of the human movement standardization evaluation technology also reduces the medical labor cost, improves the diagnostic accuracy, and promotes the intelligent development of medical services. By accurately capturing and analyzing the movements of patients and combining artificial intelligence and deep learning technologies, the medical field can provide more efficient and accurate medical services, accelerate the rehabilitation process of patients, and achieve an overall improvement in medical quality.
[0043] In the embodiments of the present application, the historical patient movement image sequence and the corresponding real future patient movement image sequence are respectively input into a preset human pose estimation algorithm to obtain the historical patient movement pose and the real future patient movement pose. Without wearing devices, it greatly improves the convenience and popularity of use and reduces the cost. Then the historical patient movement pose is input into a pre-trained human movement prediction model to obtain multiple predicted future patient movement poses. Finally, the average joint error is calculated for each frame between the predicted future patient movement pose and the real future patient movement pose, and it is judged whether the patient's movement conforms to the standard according to the error. By predicting the movement pose of each patient through the human movement prediction model and comparing it with the real future patient movement pose, accurate capture and personalized evaluation of the patient's movement are realized. The embodiments of the present application can capture and analyze the movement information of patients in real time, and can realize the standardization evaluation and guidance of human movements in various medical scenarios such as rehabilitation training conditions.
[0044] In an optional embodiment, before the historical patient movement image sequence and the corresponding real future patient movement image sequence are respectively input into a preset human pose estimation algorithm to obtain the historical patient movement pose and the real future patient movement pose, the method further includes:
[0045] Obtain the RGB video sequence of the patient's movement;
[0046] Sample and segment the RGB video sequence to obtain the historical patient movement image sequence and the corresponding real future patient movement image sequence.
[0047] In the embodiments of the present application, an RGB video sequence of a patient's movement is read from the database of the terminal device. Suppose the video sequence contains M frames. First, temporal dimension sparse average sampling is performed on the RGB video to obtain an image sequence <I1, I2, I3,..., I T >, where T represents the number of video frames obtained by sampling (assuming that the action contains 100 frames, and after coefficient average sampling, T = 50, that is, 50 frames are used to represent the action per second). The video is divided into a historical patient movement image sequence Pose his and the corresponding real future patient movement image sequence Pose future . Usually, a 2.5-second video is divided into a 2-second historical patient movement image sequence and a 0.5-second real future patient movement image sequence.
[0048] In an alternative embodiment, the historical patient movement image sequence and the corresponding real future patient movement image sequence are respectively input into a preset human pose estimation algorithm to obtain the historical patient movement pose and the real future patient movement pose, including:
[0049] Taking the historical patient movement image sequence and the corresponding real future patient movement image sequence as inputs respectively, using a two-dimensional image backbone network to predict the two-dimensional key point heat maps of the image frames, and projecting the two-dimensional key point heat maps into a discrete first three-dimensional space to construct a first 3D feature volume; the key points include the center point;
[0050] Inputting the 3D feature volume into the first three-dimensional convolutional neural network to locate the center point positions of all human bodies;
[0051] Centered on the center point position, projecting the two-dimensional key point heat maps into a second three-dimensional space to construct a second 3D feature volume;
[0052] Inputting the second 3D feature volume into the second three-dimensional convolutional neural network to predict the three-dimensional key point positions of the image frames;
[0053] Determining the historical patient movement pose based on the three-dimensional key point positions of the image frames in the historical patient movement image sequence, and determining the real future patient movement pose based on the three-dimensional key point positions of the image frames in the real future patient movement image sequence.
[0054] In the embodiments of the present application, the obtained historical patient motion image sequence his and the corresponding future real future patient motion image sequence future are respectively input into the human pose estimation algorithm. Optionally, the human pose estimation algorithm in the embodiments of the present application may adopt voxelPose. Since it is a single video stream in the embodiments of the present application, certain modifications are made to the voxelPose model, such as Figure 3 shown: In the first step, a common two-dimensional image backbone network (2D backbone) is used to predict the two-dimensional key point heatmap of the picture, and then the two-dimensional key point heatmap is projected into a discrete first 3D space to construct the first 3D feature volume. VoxelPose first detects the human body center point and then estimates the key point positions of each person. In the second step, the first 3D feature volume is input into the first three-dimensional convolutional neural network (3D CNN) to locate the center point positions of all human bodies in the scene. In the third step, for each human body center point located in the second step, taking this position as the center, the two-dimensional key point heatmap is projected into a second 3D space with a smaller space and finer grid to construct a more refined second 3D feature volume, which is input into the second 3D CNN to predict the 3D key point positions. The entire network includes the 2D backbone and 3D CNN parts and can be trained end-to-end. In the embodiments of the present application, the trained model is directly used for pose estimation to obtain the historical patient motion pose Pose his and the real future patient motion pose Pose future .
[0055] In an alternative embodiment, as Figure 3 shown, the human motion prediction model adopts a variational autoencoder, including an encoder and a decoder;
[0056] The historical patient motion pose is input into the pre-trained human motion prediction model to obtain multiple predicted future patient motion poses, including:
[0057] The historical patient motion pose is input into the encoder to generate multiple different latent variables;
[0058] The multiple different latent variables are input into the decoder for prediction to obtain the corresponding multiple different predicted future patient motion poses.
[0059] Human motion generation has two modes: deterministic and non-deterministic. Human intention is very complex and serves as an internal stimulus, prompting humans to act in different ways. Therefore, there is a certain degree of uncertainty in human motion. However, previous models often focus on deterministic generation, which does not conform to the laws of human motion. Adding random Gaussian noise to the input can increase the randomness of human motion generation, but simply adding noise may make the prediction discontinuous. Therefore, the embodiments of this application use variational autoencoders to generate more diverse results and achieve diversified prediction of the future, which can produce excellent effects in uncertain predictions.
[0060] In an optional embodiment, the historical motion postures of patients are input into the encoder to generate multiple different latent variables, including:
[0061] For each time frame feature in the historical motion postures of patients, multi-scale features are extracted on the spatial scale;
[0062] A graph convolutional neural network is used to extract features from the multi-scale features to obtain multi-scale graph convolutional features;
[0063] The multi-scale graph convolutional features are subjected to discrete cosine transform to obtain multi-scale discrete transform features;
[0064] A cross-attention mechanism is used to fuse the multi-scale discrete transform features to obtain discrete fusion features;
[0065] The discrete fusion features are subjected to inverse discrete cosine transform, and the time scale is compressed through a convolutional neural network to obtain the latent variable of each time frame.
[0066] In the design of the encoder of the model, in order to better capture the required features, the embodiments of this application adopt a multi-scale method. Optionally, for the features of each time frame, third-order scale features are extracted on the spatial scale, such as Figure 4 shown, using black dots to represent the joint points we use. Three levels use different numbers of joint points to model the complex motion features of the human body. Starting from joint points, bones, and the limbs and torso of the human body respectively, multi-scale features are obtained: where M represents the batch, T represents the input time length, N s represents the number of input joint points, d represents the feature dimension of each joint point, and different levels of human body representations are obtained.
[0067] After that, parameterized graph convolution (GCN) is added to extract features. The formula for graph convolution is X s1 =A s1 x s1 W s1 where A s1 and Ws1 They are two parameterized adjacency matrices and weight matrices. Since in human motion patterns, the relationships between many bones are not explicitly given, there are some implicit motion relationships in human motion, and we can obtain such features through the learning of neural networks. After that, the corresponding multi-scale graph convolution features are obtained: X s2 、X s3 , where D is the new feature dimension.
[0068] Considering that frequency domain analysis has played a good role in previous tasks, the embodiment of this application uses discrete cosine transform (DCT) to extract features in each window, discarding the high-frequency noise part to make the action features smoother. This process is expressed as X s2 、X s3 Similarly, after performing discrete cosine transform on each block, multi-scale discrete transform features are obtained: X p2_DCT 、X p3_DCT , where f represents the frequency size selected for use.
[0069] In the fusion of information at different scales, the embodiment of this application adopts different methods: for spatial scale information, a structure based on cross-attention mechanism (such as Figure 5 shown) is used for feature fusion. The discrete transform features of the first scale are used as Q and K respectively to perform cross-attention with the discrete transform features of the second scale and the discrete transform features of the third scale. The formula is:
[0070]
[0071] Then, the result in the spatio-temporal domain is obtained through inverse discrete cosine transform (IDCT), and the time scale is compressed through a convolutional neural network (CNN) to obtain latent variables.
[0072] In an optional embodiment, multiple different latent variables are input into the decoder for prediction to obtain corresponding multiple different predicted future patient motion postures, including:
[0073] After fusing the latent variables of different time scales, they are input into the graph convolution layer to obtain graph convolution features;
[0074] The graph convolution features are predicted frame by frame through a gated recurrent unit to obtain the predicted future patient motion posture of each time frame;
[0075] Among them, the predicted future patient motion posture of the previous time frame is connected in a residual manner with the current gated recurrent unit after passing through the graph convolutional layer to predict the predicted future patient motion posture of the current time frame.
[0076] In the design of the decoder of the model, considering that in the time series prediction task, the Recurrent Neural Network (RNN) has always shown good results. The embodiment of the present application uses a gated recurrent unit (GRU) with graph convolution for prediction decoding, as Figure 6 shown. When fusing different time scale information, the input data first passes through a graph convolutional layer, that is, h = AZW, After the first graph convolutional layer, h(t) = Ah(t)W. Then, prediction is performed frame by frame through the GRU unit, and the residual connection method is also used during prediction to improve the accuracy. For each GRU unit, this overall process is expressed as:
[0077] h(t) = Ah(t)W
[0078] r t = σ(r in (Input(t)) + r hid (h))
[0079] u t = σ(u in (Input(t)) + u hid (h))
[0080] c t = tanh(c in (Input(t)) + c hid (h))
[0081] h(t + 1) = u t ⊙ Input(t) + (1 - u t ) ⊙ c t
[0082] Input(t + 1) = MLP(h(t + 1)) + Input(t)
[0083] Among them, r in , r hid , u in , u hid , c in , c hid , and MLP respectively represent different linear layers, σ(·) and tanh(·) represent the sigmoid and tanh activation functions. The input data will obtain a new predicted frame through each GRU unit and pass the hidden variable backward.
[0084] The human motion prediction model in the embodiments of this application adopts a long-term 3D human motion prediction algorithm that can ensure diverse generation. This model fuses the attention mechanism with convolution, and performs discrete cosine transform on the time series to better extract human features. It combines the gated recurrent unit (GRU) in the recurrent neural network (RNN) and graph convolution, hoping to overcome problems such as error accumulation in RNN modeling and insufficient receptive fields of CNN. In the traditional resampling method, the introduction of random Gaussian noise is reduced, and samples are directly generated using the input, and generated through VAE to generate more diverse results, realizing diversified prediction of the future.
[0085] In an optional embodiment, the average joint error is calculated for each frame between the predicted future patient motion pose and the real future patient motion pose, and whether the patient's motion conforms to the specification is judged according to the error. The specific implementation process is as follows:
[0086] Input the historical patient motion pose Pose his into the human motion prediction model mentioned above, and K possible future motion patterns can be obtained. The K predicted future patient motion poses Pose pre are compared with the real patient motion pose Pose future for each frame to calculate the average joint error. That is, for the N joint points (i.e., key points) of each patient obtained in VoxelPose, the average error is calculated in terms of time and joint points. Assuming the time length of the future pose is T, the formula for calculating the error err is:
[0087]
[0088] Based on the magnitude of the obtained error err, it is judged whether the patient's motion conforms to the specification. The smaller the err, the higher the degree of standardization. The larger the err, the more likely the posture of the human motion is abnormal, and there may still be a sub-healthy state.
[0089] The above method in the embodiments of this application can also evaluate the rehabilitation training actions of patients after treatment, such as limb stretching, balance exercises, breathing control, etc. Input the corresponding motion actions into the algorithm model of the embodiments of this application, and judge whether the rehabilitation training actions of the patients follow the established professional norms and standards in the medical field according to the output results, and achieve the effects of improving the professional skill level of medical staff, reducing the occurrence of medical errors and complications, effectively guiding patients to correctly perform rehabilitation training, accelerating their recovery process, improving the overall medical quality and patient satisfaction.
[0090] In an optional embodiment, before inputting the historical patient motion postures into a pre-trained human motion prediction model to obtain various predicted future patient motion postures, the method further includes:
[0091] In the first stage, the human motion prediction model is trained using KL divergence and mean squared loss;
[0092] In the second stage, the human motion prediction model is trained using mean squared loss;
[0093] Among them, the first stage is: for the input data, the mean and variance are generated by the encoder, the latent variables are generated by resampling, input into the decoder, and the prediction result is generated by the decoder;
[0094] The second stage is: for the input data, multiple latent variables are generated by the new encoder and input into the decoder trained in the first stage, and the prediction result is generated by the decoder.
[0095] For the method proposed in the embodiments of the present application, the entire training process can be summarized as follows:
[0096] The first stage: The input is a vector of dimension (B, H, D). The corresponding mean and variance are generated by the Encoder, that is, the number of latent variables is 2, and then the prediction result vector y of dimension (B, F, D) is generated by resampling. In this stage, the traditional KL divergence and mean squared loss are used, that is:
[0097] loss = KL(N(μ, σ)||N(0, 1)) + MSE(y, GT)
[0098] The second stage: The input is a vector of dimension (B, H, D). Using the Decoder trained in the first stage and a new Encoder, K latent variables are generated by the Encoder and input into the above-trained Decoder to obtain K predicted results. Mean squared loss is used for training:
[0099]
[0100] Among them, B represents the batch size, H represents the historical sequence length, K represents the total number of output samples, K represents the number of samples, D represents the feature dimension, t and D represent the intermediate feature dimensions. F represents the future sequence length, and GT represents the ground truth. Encoder1 and Encoder2 are both based on the above structure, and the only difference is the size of their output dimensions. KL represents KL divergence, and MSE represents mean squared loss.
[0101] Based on the above method, in the embodiments of the present application, an RGB video sequence is used as the input, and the overall process of using the human pose estimation algorithm and the human motion prediction model for action normalization evaluation is as Figure 7 shown as follows:
[0102] Step 1: Input the RGB video sequence, and sample and segment the video frames into a historical patient motion image sequence and a real future patient motion image sequence.
[0103] Step 2: For the historical patient motion image sequence, obtain the historical patient motion poses through voxelPose, and then input the historical patient motion poses into the human motion prediction model to obtain K kinds of predicted future patient motion poses;
[0104] For the real future patient motion image sequence, obtain the real future patient motion poses through voxelPose.
[0105] Step 3: Calculate the minimum error err between the predicted future patient motion poses and the real future patient motion poses.
[0106] Step 4: Judge whether the rehabilitation training actions are standard according to err. Specifically: Judge whether the rehabilitation training actions are accurate according to err. If so, point out the problems and give adjustment suggestions. Otherwise, determine that the rehabilitation training actions are standard;
[0107] Perform rehabilitation status detection according to err. Specifically: Judge whether there are problems with the patient's daily motion poses according to err. If so, point out the problems and give diagnostic suggestions. Otherwise, determine that the human body has recovered well.
[0108] In summary, in the encoder part of the embodiments of the present application, a structure based on the attention mechanism is used to fuse multi-scale spatial information. At the same time, the discrete cosine transform (DCT) is used to convert the time-scale information into the frequency domain for processing, effectively reducing the high-frequency noise in the original data. In the decoder part, the gated recurrent unit (GRU) in the recurrent neural network (RNN) and graph convolution are combined to generate time series prediction results. In the prediction method, the existing motion prediction technology is improved so that it can predict multiple potential future motion trends based on the observed motion sequence. The embodiments of the present application also propose a new diversity sampling method. Based on the pre-trained variational autoencoder (VAE), a new encoder is trained to generate different latent variables. These latent variables correspond to different future motion patterns, thus achieving richer prediction results.
[0109] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application also provide an action normalization evaluation system based on the human motion prediction algorithm, as Figure 8 shown, the system includes:
[0110] The human body pose detection module 801 is configured to input the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human body pose estimation algorithm respectively, so as to obtain the historical patient motion pose and the real future patient motion pose;
[0111] The human body pose prediction module 802 is configured to input the historical patient motion pose into a pre-trained human body motion prediction model, so as to obtain multiple predicted future patient motion poses;
[0112] The specification evaluation module 803 is configured to calculate the average joint error for each frame between the predicted future patient motion pose and the real future patient motion pose, and determine whether the patient's motion conforms to the specification according to the error.
[0113] In the embodiment of the present application, the historical patient motion image sequence and the corresponding real future patient motion image sequence are respectively input into a preset human body pose estimation algorithm to obtain the historical patient motion pose and the real future patient motion pose. Without wearing devices, the convenience and popularity of use are greatly improved, and the cost is reduced. Then, the historical patient motion pose is input into a pre-trained human body motion prediction model to obtain multiple predicted future patient motion poses. Finally, the average joint error is calculated for each frame between the predicted future patient motion pose and the real future patient motion pose, and it is determined whether the patient's motion conforms to the specification according to the error. By predicting the motion pose of each patient through the human body motion prediction model and comparing it with the real future patient motion pose, accurate capture and personalized evaluation of the patient's actions are realized. The embodiment of the present application can capture and analyze the action information of the patient in real time, and can realize the standardized evaluation and guidance of the human body's actions under various medical scenarios such as rehabilitation training conditions.
[0114] The action standardization evaluation system based on the human body motion prediction algorithm provided by the embodiment of the present application can implement Figures 1 to 7 each process implemented in the method embodiment. To avoid repetition, it will not be elaborated here.
[0115] The action standardization evaluation system based on the human body motion prediction algorithm of the embodiment of the present application can execute the action standardization evaluation method based on the human body motion prediction algorithm provided by the embodiment of the present application. Its implementation principle is similar. The actions performed by each module and unit in the action standardization evaluation system based on the human body motion prediction algorithm in each embodiment of the present application correspond to the steps in the action standardization evaluation method based on the human body motion prediction algorithm in each embodiment of the present application. For the detailed function description of each module of the action standardization evaluation system based on the human body motion prediction algorithm, reference can be specifically made to the description in the corresponding action standardization evaluation method based on the human body motion prediction algorithm shown above, which will not be elaborated here.
[0116] Based on the same principle as the method shown in the embodiments of the present application, the embodiments of the present application further provide an electronic device, which may include, but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the action normalization evaluation method based on the human motion prediction algorithm shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, the action normalization evaluation method based on the human motion prediction algorithm provided by the present application inputs the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose, without the need for wearable devices, greatly improving the convenience and popularity of use and reducing the cost. Then, the historical patient motion pose is input into a pre-trained human motion prediction model to obtain multiple predicted future patient motion poses. Finally, the average joint error is calculated for each frame between the predicted future patient motion pose and the real future patient motion pose, and it is judged whether the patient's motion conforms to the specification according to the error. By predicting the motion pose of each patient through the human motion prediction model and comparing it with the real future patient motion pose, accurate capture and personalized evaluation of the patient's actions are realized. The embodiments of the present application can capture and analyze the patient's action information in real time, and can realize the normalization evaluation and guidance of the human body's actions in various medical scenarios, such as under rehabilitation training conditions.
[0117] In an optional embodiment, an electronic device is further provided, as Figure 9 shown Figure 9 The electronic device 900 shown may be a server, including: a processor 901 and a memory 903. Among them, the processor 901 and the memory 903 are connected, such as through a bus 902. Optionally, the electronic device 900 may further include a transceiver 904. It should be noted that in practical applications, the transceiver 904 is not limited to one, and the structure of the electronic device 900 does not constitute a limitation to the embodiments of the present application.
[0118] The processor 901 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 901 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0119] The bus 902 may include a path for transmitting information between the above components. The bus 902 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 902 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0120] The memory 903 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0121] The memory 903 is used to store the application program code for executing the solution of this application, and is controlled by the processor 901 for execution. The processor 901 is used to execute the application program code stored in the memory 903 to implement the content shown in the foregoing method embodiments.
[0122] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0123] The server provided by this application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0124] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.
[0125] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0126] It should be noted that the above-mentioned computer-readable storage medium in the present application can also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0127] The above computer-readable medium can be included in the above electronic device; it can also exist separately without being assembled into the electronic device.
[0128] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0129] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the action normalization evaluation method and system based on the human motion prediction algorithm provided in the above various optional implementation manners.
[0130] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0132] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the human pose detection module can also be described as "a human pose detection module for respectively inputting a historical patient motion image sequence and a corresponding real future patient motion image sequence into a preset human pose estimation algorithm to obtain the historical patient motion pose and the real future patient motion pose".
[0133] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.
Claims
1. An action normalization evaluation method based on a human motion prediction algorithm, characterized in that The method includes: Inputting the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose; Inputting the historical patient motion pose into a pre-trained human motion prediction model to obtain multiple predicted future patient motion poses; Calculating the average joint error between each frame of the predicted future patient motion pose and the real future patient motion pose, and judging whether the patient's motion conforms to the specification according to the error.
2. The action normalization evaluation method based on the human motion prediction algorithm according to claim 1, wherein The human motion prediction model adopts a variational autoencoder, including an encoder and a decoder; The step of inputting the historical patient motion pose into the pre-trained human motion prediction model to obtain multiple predicted future patient motion poses includes: Inputting the historical patient motion pose into the encoder to generate multiple different latent variables; Inputting the multiple different latent variables into the decoder for prediction to obtain corresponding multiple different predicted future patient motion poses.
3. The action normalization evaluation method based on the human motion prediction algorithm according to claim 2, wherein The step of inputting the historical patient motion pose into the encoder to generate multiple different latent variables includes: For each time frame feature in the historical patient motion pose, extracting multi-scale features on the spatial scale; Using a graph convolutional neural network to extract features from the multi-scale features to obtain multi-scale graph convolutional features; Performing discrete cosine transform on the multi-scale graph convolutional features to obtain multi-scale discrete transform features; Using a cross-attention mechanism to fuse the multi-scale discrete transform features to obtain discrete fusion features; Performing inverse discrete cosine transform on the discrete fusion features and compressing the time scale through a convolutional neural network to obtain the latent variable of each time frame.
4. The action normalization evaluation method based on the human motion prediction algorithm according to claim 3, characterized in that The step of inputting the multiple different latent variables into the decoder for prediction to obtain corresponding multiple different predicted future patient motion poses includes: Fusing the latent variables of different time scales and inputting them into a graph convolutional layer to obtain graph convolutional features; Performing frame-by-frame prediction on the graph convolutional features through a gated recurrent unit to obtain the predicted future patient motion pose of each time frame; Among them, the predicted future patient motion pose of the previous time frame is connected in residual with the current gated recurrent unit after passing through the graph convolutional layer to predict the predicted future patient motion pose of the current time frame.
5. The action normalization evaluation method based on the human motion prediction algorithm according to claim 2, wherein Before the step of inputting the historical patient motion pose into the pre-trained human motion prediction model to obtain multiple predicted future patient motion poses, the method further includes: In the first stage, training the human motion prediction model using KL divergence and mean square loss; In the second stage, training the human motion prediction model using mean square loss; Among them, the first stage is: for the input data, generating a mean and variance through the encoder, generating multiple latent variables through resampling, inputting them into the decoder, and generating a prediction result through the decoder; The second stage is: for the input data, generating multiple latent variables through a new encoder and inputting them into the decoder trained in the first stage, and generating a prediction result through the decoder.
6. The action normalization evaluation method based on the human motion prediction algorithm according to claim 1, characterized in that, Inputting the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose, including: Taking the historical patient motion image sequence and the corresponding real future patient motion image sequence as inputs respectively, using a two-dimensional image backbone network to predict the two-dimensional key point heat maps of the image frames, and projecting the two-dimensional key point heat maps into a discrete first three-dimensional space to construct a first 3D feature volume; the key points include the center points; Inputting the first 3D feature volume into a first three-dimensional convolutional neural network to locate the center point positions of all human bodies; Centering on the center point positions, projecting the two-dimensional key point heat maps into a second three-dimensional space to construct a second 3D feature volume; the second 3D feature volume has a smaller space and finer grids than the first 3D feature volume; Inputting the second 3D feature volume into a second three-dimensional convolutional neural network to predict the three-dimensional key point positions of the image frames; Determining the historical patient motion pose based on the three-dimensional key point positions of the image frames in the historical patient motion image sequence, and determining the real future patient motion pose based on the three-dimensional key point positions of the image frames in the real future patient motion image sequence.
7. The action normalization evaluation method based on the human motion prediction algorithm according to claim 1, characterized in that Before inputting the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose, the method further includes: Obtaining the RGB video sequence of the patient's motion; Sampling and segmenting the RGB video sequence to obtain the historical patient motion image sequence and the corresponding real future patient motion image sequence.
8. An action normalization evaluation system based on a human motion prediction algorithm, characterized in that, The system includes: A human pose detection module, configured to input the historical patient motion image sequence and the corresponding real future patient motion image sequence into a preset human pose estimation algorithm respectively to obtain the historical patient motion pose and the real future patient motion pose; A human pose prediction module, configured to input the historical patient motion pose into a pre-trained human motion prediction model to obtain multiple predicted future patient motion poses; A specification evaluation module, configured to calculate the average joint error between each frame of the predicted future patient motion pose and the real future patient motion pose, and determine whether the patient's motion conforms to the specification according to the error.
9. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method according to any one of claims 1-7 when executing the program.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-7.
Citation Information
Cited By
Human motion function evaluation method and human motion function evaluation system
TWI934715B