A method and device for predicting human motion intention during rehabilitation training
Through deep learning network models, the problem of movement delay error in traditional rehabilitation training equipment is solved, the synchronization between the rehabilitation training platform and the patient's posture is improved, and the training effect is improved.
Patent Information
- Application Number
- CN202410891683.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-07-04
AI Technical Summary
Traditional rehabilitation training equipment lags behind the human body's movement intention because the force sensor collects signals, resulting in lag in the training platform, reducing the training effect.
The deep learning network model is adopted, combining sole pressure data and joint posture data, and through data block embedding, encoder layer, time series decomposition layer and decoder layer, the human body's motion intention is predicted, the motion delay error is eliminated, and the synchronization is improved.
Through the prediction of the deep learning network model, the synchronization between the position of the rehabilitation training platform and the position of the patient is achieved, and the training effect is improved.
Smart Images

Figure CN118861765B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of action recognition and prediction, and particularly to a method and device for predicting human motion intention in rehabilitation training. Background Art
[0002] With the aggravation of the problem of population aging, there are more and more patients with impaired cognitive and balance coordination caused by stroke, hemiplegia, etc. Traditional rehabilitation methods have problems such as low cure rate and shortage of professional rehabilitation personnel.
[0003] In view of the above problems, equipment research based on robot-assisted rehabilitation has emerged, such as balance rehabilitation platforms, exoskeleton robots, etc., enabling patients to control the platform according to their own motion intentions, which can better mobilize the enthusiasm of patients for training. Such rehabilitation equipment often judges the patient's motion intention through force sensors to achieve motion tracking, and then controls the motion state of the platform according to the patient's motion intention, so as to achieve training under human-computer interaction. However, since the signals collected by the force sensors lag behind the human motion intention, and coupled with the time delay of the training platform control system, the motion of the training platform will lag behind the human motion intention, weakening the synchronization between the pose of the training platform and the pose of the patient, thus reducing the training effect. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present invention provides a method and device for predicting human motion intention in rehabilitation training to eliminate the delay error of the motion of the rehabilitation training platform relative to the human motion intention to the greatest extent, improve the synchronization between the pose of the training platform and the pose of the patient, and thus improve the training effect.
[0005] In one aspect of the present invention, there is provided a method for predicting human motion intention in rehabilitation training, including: collecting plantar pressure data and joint pose data of a human body; inputting the obtained plantar pressure data and joint pose data into a trained deep learning network model in the form of a time series to obtain a prediction result of the human motion intention; wherein, the deep learning network model includes: a data block embedding module for processing a discrete original input time series into a data block of a predetermined length, performing position encoding and value encoding on the data block, and outputting the encoded data block to an encoder layer; an encoder layer for extracting first periodic features of the original input time series from the encoded data block and outputting the first periodic features to a decoder layer; a first time series decomposition layer for decomposing the original input time series into initial periodic features and initial trend features and inputting the initial periodic features and initial trend features into the decoder layer; a decoder layer for calculating total trend features and total periodic features according to the first periodic features, initial trend features and initial periodic features, and adding the total trend features and total periodic features as a final prediction sequence.
[0006] Further, the encoder layer includes a plurality of serially connected encoders, and each encoder includes a multi-head self-attention module, a first residual and layer normalization module, a second time series decomposition layer, a first feed-forward layer, a second residual and layer normalization module, and a third time series decomposition layer that are connected in sequence; wherein, the output end of the data block embedding module is connected to the first residual and layer normalization module, and the output end of the second time series decomposition layer is connected to the second residual and layer normalization module; the third time series decomposition layer outputs the first periodic feature; the first residual and layer normalization module is used to form a skip connection between the input and output of the multi-head self-attention module; the second residual and layer normalization module is used to form a skip connection between the input and output of the first feed-forward layer; the multi-head self-attention module is used to calculate the correlation between the plantar pressure data and the joint pose data in the original input time series; the second time series decomposition layer and the third time series decomposition layer are used to decompose the trend feature and the periodic feature of their respective input data.
[0007] Further, the decoder layer includes a plurality of serially connected decoders, and each decoder includes a first branch and a second branch; the first branch includes a masked multi-head self-attention module, a third residual and layer normalization module, a fourth time series decomposition layer, an encoder-decoder self-attention module, a fourth residual and layer normalization module, a fifth time series decomposition layer, a second feed-forward layer, and a fifth residual and layer normalization module that are connected in sequence; wherein, the first periodic feature output by the encoder layer is input to the encoder-decoder self-attention module, the initial periodic feature output by the first time series decomposition layer is input to the masked multi-head self-attention module and the third residual and layer normalization module, the periodic feature component output by the fourth time series decomposition layer is input to the encoder-decoder self-attention module and the fourth residual and layer normalization module, and the periodic feature component output by the fifth time series decomposition layer is input to the second feed-forward layer and the fifth residual and layer normalization module; the second branch includes a sixth residual and layer normalization module and a seventh residual and layer normalization module that are connected in sequence; wherein, the initial trend feature output by the first time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fourth time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fifth time series decomposition layer is input to the seventh residual and layer normalization module, and the output of the seventh residual and layer normalization module is input to the fifth residual and layer normalization module; the fifth residual and layer normalization module adds the output of the second feed-forward layer, the output of the fifth time series decomposition layer, and the output of the seventh residual and layer normalization module as the final prediction sequence.
[0008] Further, the first time series decomposition layer, the second time series decomposition layer, the third time series decomposition layer, the fourth time series decomposition layer, and the fifth time series decomposition layer use a weighted moving average function to decompose the periodic features and trend features of the input data.
[0009] Further, it further includes: a preprocessing module for performing a stationary processing on the original input time series, eliminating non-stationary components, and outputting the processed data to the data block embedding module.
[0010] On the other hand, the present invention further provides a human motion intention prediction device, including: a data acquisition module for acquiring the plantar pressure data and the joint pose data of the human body; a prediction module for inputting the acquired plantar pressure data and joint pose data into a trained deep learning network model in the form of a time series to obtain a prediction result of the human motion intention; wherein, the deep learning network model includes: a data block embedding module for processing the discrete original input time series into data blocks of a predetermined length, performing positional encoding and value encoding on the data blocks, and outputting the encoded data blocks to the encoder layer; an encoder layer for extracting the first periodic features of the original input time series from the encoded data blocks and outputting the first periodic features to the decoder layer; a first time series decomposition layer for decomposing the original input time series into initial periodic features and initial trend features and inputting the initial periodic features and initial trend features into the decoder layer; a decoder layer for calculating the total trend features and total periodic features according to the first periodic features, initial trend features, and initial periodic features, and adding the total trend features and total periodic features as the final prediction sequence.
[0011] Further, the encoder layer includes a plurality of serially connected encoders, and each encoder includes a multi-head self-attention module, a first residual and layer normalization module, a second time series decomposition layer, a first feed-forward layer, a second residual and layer normalization module, and a third time series decomposition layer connected in sequence; wherein, the output end of the data block embedding module is connected to the first residual and layer normalization module, and the output end of the second time series decomposition layer is connected to the second residual and layer normalization module; the third time series decomposition layer outputs the first periodic features; the first residual and layer normalization module is used to form a skip connection between the input and output of the multi-head self-attention module; the second residual and layer normalization module is used to form a skip connection between the input and output of the first feed-forward layer; the multi-head self-attention module is used to calculate the correlation between the plantar pressure data and joint pose data in the original input time series; the second time series decomposition layer and the third time series decomposition layer are used to decompose the trend features and periodic features of their respective input data.
[0012] Further, the decoder layer includes a plurality of decoders connected in series, and each decoder includes a first branch and a second branch; the first branch includes a masked multi-head self-attention module, a third residual and layer normalization module, a fourth time series decomposition layer, an encoder-decoder self-attention module, a fourth residual and layer normalization module, a fifth time series decomposition layer, a second feed-forward layer, and a fifth residual and layer normalization module connected in sequence; wherein, the first periodic feature output by the encoder layer is input to the encoder-decoder self-attention module, the initial periodic feature output by the first time series decomposition layer is input to the masked multi-head self-attention module and the third residual and layer normalization module, the periodic feature component output by the fourth time series decomposition layer is input to the encoder-decoder self-attention module and the fourth residual and layer normalization module, and the periodic feature component output by the fifth time series decomposition layer is input to the second feed-forward layer and the fifth residual and layer normalization module; the second branch includes a sixth residual and layer normalization module and a seventh residual and layer normalization module connected in sequence; wherein, the initial trend feature output by the first time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fourth time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fifth time series decomposition layer is input to the seventh residual and layer normalization module, and the output of the seventh residual and layer normalization module is input to the fifth residual and layer normalization module; the fifth residual and layer normalization module adds the output of the second feed-forward layer, the output of the fifth time series decomposition layer, and the output of the seventh residual and layer normalization module as the final predicted sequence.
[0013] Further, the first time series decomposition layer, the second time series decomposition layer, the third time series decomposition layer, the fourth time series decomposition layer, and the fifth time series decomposition layer use a weighted moving average function to decompose the input data into periodic features and trend features.
[0014] Further, it further includes: a preprocessing module for performing a stationary processing on the original input time series, eliminating non-stationary components, and outputting the processed data to the data block embedding module.
[0015] A method and device for predicting human motion intention during rehabilitation training provided by the present invention construct a prediction model for pose time series based on deep learning. Aiming at the trend, period, and short-term dependence characteristics of pose time series prediction, a deep decomposition architecture and a time series decomposition module are designed to decompose the original time series, and then feature learning and prediction are carried out. The motion intention prediction model of the present invention realizes the information fusion of pose information and force information in motion intention prediction, and based on this, extracts the cross-dimensional dependence of different variables. Through the unique time series decomposition module, multiple decompositions of various motion intention information are realized, reducing the prediction difficulty of the model. Description of the Drawings
[0016] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present application will become more obvious:
[0017] Figure 1 is a flowchart of a method for predicting human motion intention during rehabilitation training provided by an embodiment of the present application;
[0018] Figure 2 is a schematic structural diagram of a deep learning network model provided by an embodiment of the present application;
[0019] Figure 3 is a schematic structural diagram of a device for predicting human motion intention during rehabilitation training provided by an embodiment of the present application. Detailed Embodiments
[0020] To make the purposes, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0022] It should be understood that although the terms first, second, third, etc. may be used to describe the acquisition modules in the embodiments of the present invention, these acquisition modules should not be limited to these terms. These terms are only used to distinguish the acquisition modules from each other.
[0023] Depending on the context, as used herein, the term "if" may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0024] It should be noted that the orientation terms such as "upper", "lower", "left", and "right" described in the embodiments of the present invention are described from the angles shown in the drawings, and should not be construed as limiting the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that one element is formed "on" or "under" another element, it can not only be directly formed "on" or "under" another element, but also be indirectly formed "on" or "under" another element through an intermediate element.
[0025] In order to minimize the delay error of the movement of the rehabilitation training platform relative to the human movement intention and improve the synchronization between the pose of the training platform and the pose of the patient, the present application proposes to predict the movement pose of the human body through an improved deep learning network model.
[0026] See Figure 1 , an embodiment of the present invention provides a method for predicting human movement intention in rehabilitation training, including the following steps:
[0027] Step S101, collect the plantar pressure data and the pose data of each joint of the human body.
[0028] Specifically, the force signal of the plantar pressure of the human body is collected through a plantar pressure sensor, and the pose signals of each joint of the human body are collected through inertial sensors (IMUs, also known as attitude sensors) fixed on the arm, thigh, and / or calf. Subsequently, in this embodiment, an improved deep learning network model will be used to compare the plantar pressure data and the pose data of each joint to predict the future trend of the human pose.
[0029] Step S102, input the obtained plantar pressure data and the pose data of each joint into the trained deep learning network model in the form of a time series to obtain the prediction result of the human movement intention.
[0030] Specifically, the above-collected plantar pressure data and the pose data of each joint are input into the trained deep learning network model in the form of a time series, and the model is used to predict the future movement pose of the human body, so as to give the advance amount for controlling the movement of the rehabilitation training platform according to the predicted human pose, so as to eliminate the delay error of the movement of the rehabilitation training platform relative to the human movement intention and improve the synchronization between the pose of the training platform and the pose of the patient.
[0031] In this embodiment, to solve the problem of insufficient ability of the long short-term memory network to extract global information, signal feature information is extracted and future signals are predicted based on the Transformer technology. However, it is difficult for the ordinary Transformer network structure to predict complex signals. Therefore, this application proposes a time series decomposition module composed of a time series decomposition layer, and develops an improved deep learning network model based on this.
[0032] See Figure 2 , the improved deep learning network model of this embodiment includes:
[0033] (1) Preprocessing module
[0034] The preprocessing module is used to perform stationary processing on the original input time series, eliminate non-stationary components, and output the processed data to the data block embedding module.
[0035] Specifically, for stationary processing of the original input time series data, first calculate the mean u z and standard deviation σ z of the input time series Z:
[0036]
[0037] where n is the length of the input time series, the subscript z represents the sequence, and the value of i ranges from 1 to n.
[0038] Perform stationary processing through the following formula to eliminate the non-stationary components of the original sequence:
[0039]
[0040] Output the processed data to the data block embedding module.
[0041] (2) Data block embedding module
[0042] The data block embedding module is used to process the discrete input time series into data blocks of a predetermined length, perform positional encoding and value encoding on the data blocks, and output the encoded data blocks to the encoder layer.
[0043] Specifically, the data block embedding module simultaneously adopts value embedding and positional embedding. Value embedding uses a linear layer to map the data from the original dimension to the embedding feature dimension, and the embedding feature dimension is an adjustable hyperparameter, generally 512. The dimension of the positional embedding is the same as that of the value embedding. The positional encoding of adjacent two dimensions is calculated using sine and cosine functions, which has periodicity. The vector after positional embedding and the vector after value embedding are added together to be the data input to the network.
[0044] The data block embedding module is used to encode the sequential information and the forward and backward dependencies of the pose sequence, enhance the model's perception ability of the pose sequence, and capture the non-linear complex relationships therein. At the same time, the input data is mapped to a high-dimensional vector space, so that the low-dimensional features of the input data can be converted into more abstract and richer high-dimensional features, which helps the model to capture more potential information and patterns in the high-dimensional space.
[0045] (3) Encoder layer
[0046] The encoder is used to extract the first periodic features of the input time series from the encoded data blocks and output the first periodic features to the decoder layer.
[0047] Specifically, the encoder layer is composed of multiple encoder modules, and the modules are in series connection. The effective extraction of the deep feature information of the image is achieved through N encoder modules. Each encoder includes a multi-head self-attention module, a first residual and layer normalization module, a second time series decomposition layer, a first feed-forward layer, a second residual and layer normalization module, and a third time series decomposition layer connected in sequence; wherein, the output end of the data block embedding module is connected to the first residual and layer normalization module, and the output end of the second time series decomposition layer is connected to the second residual and layer normalization module; the third time series decomposition layer outputs the first periodic features.
[0048] The first residual and layer normalization module is used to form a skip connection between the input and output of the multi-head self-attention module. The skip connection enables the data after the attention calculation and the data without calculation to perform a pointwise addition operation. Among them, the second residual and layer normalization module is used to form a skip connection between the input and output of the first feed-forward layer. The skip connection enables the data after the first feed-forward layer calculation and the data without calculation to perform a pointwise addition operation. The above skip connection structure accelerates the convergence speed and prediction effect of the network.
[0049] The multi-head self-attention module in the encoder is used to calculate the correlation between the plantar pressure data and the joint pose data in the input time series. This multi-head self-attention module adopts a cross-dimensional attention calculation module. The shape of the data input into the network is batch, n_vars, d_model. The rearrange function is used to merge the batch*n_vars dimension, and the attention calculation of a single variable in the time dimension is performed. Subsequently, the calculated variable is restored to the shape of batch, n_vars, d_model, and the cross-dimensional attention score is calculated. Among them, batch is the number of samples input into the network at one time, which is an adjustable hyperparameter. n_vars is the number of dimensions of the variables input into the network, and d_model is the dimension after embedding representation, generally 512, which is adjustable. In this embodiment, by continuously and alternately adopting the cross-dimensional attention calculation module and the time series decomposition layer, the information interaction of different variables in the time dimension and the signal dimension can be realized.
[0050] The second time series decomposition layer and the third time series decomposition layer in the encoder are used to decompose the trend features and periodic features in their respective input data. Specifically, the time series decomposition layer of the present invention decomposes the input time series Z composed of the original human joint pose sequences X, Y, Z, RX, RY, RZ (joint coordinates and joint rotation angles) and force signals into a trend feature Z t , a periodic feature Z p and a residual feature Z R into three subsequences, which are processed separately:
[0051] Z = Z t + Z p + Z R ;
[0052] The time series decomposition layer uses a weighted moving average function, giving higher weights to the closer data. After obtaining a relatively stable periodic term, the trend term is calculated. This time series decomposition layer first pads the data at the beginning and end of the input sequence to ensure that the periodic features and trend features can also be calculated at the head and tail of the sequence, that is, the padding operation in the following formula.
[0053] The trend feature Z t is obtained through the following formula:
[0054] Z t = WeightPool(padding(Z));
[0055] The periodic feature Z p is then obtained through the following formula:
[0056] Z p = Z - Z t
[0057] In this embodiment, the input time series is first filled at the beginning and the end, and then a weighted moving average is performed using a moving average window with a length of kernel_size to decompose its periodic features and trend features. The encoder decomposes the sequence features through two consecutive second time series decomposition layers and third time series decomposition layers, obtaining deep feature information.
[0058] (4) First time series decomposition layer
[0059] The first time series decomposition layer is used to decompose the input time series into initial periodic features and initial trend features, and input the initial periodic features and initial trend features into the decoder layer.
[0060] As Figure 2 shown, the data input to the network undergoes a time series decomposition through the first time series decomposition layer to obtain the initial periodic features and initial trend features input to the decoder layer. The subsequent decoder layer will calculate the prediction sequence based on the initial periodic features, initial trend features, and the first periodic features input by the encoder layer.
[0061] (5) Decoder layer
[0062] The decoder layer is used to calculate the total trend features and total periodic features based on the first periodic features, initial trend features, and initial periodic features, and add the total trend features and total periodic features as the final prediction sequence.
[0063] The structure of the decoder layer is as Figure 2 shown. Each decoder in this embodiment is designed with a dual-branch structure. The left branch processes the periodic features, and the right branch processes the trend features. The initial values of the periodic features and trend features input to the decoder are the components after the original input time series undergoes a one-layer time series decomposition through the first time series decomposition layer. Among them, the left branch passes through the backbone network and fuses the first periodic features extracted by the encoder layer to obtain the total periodic features in the prediction sequence. The right branch continuously accumulates the trend features decomposed by the left time series decomposition module based on the input initial trend features to obtain the total trend features. Finally, the sum of the total periodic features and total trend features of the left branch and the right branch is the final prediction sequence.
[0064] Specifically, the decoder layer includes multiple decoders connected in series, and each decoder includes a first branch and a second branch.
[0065] The first branch includes a masked multi-head self-attention module, a third residual and layer normalization module, a fourth time series decomposition layer, an encoder-decoder self-attention module, a fourth residual and layer normalization module, a fifth time series decomposition layer, a second feed-forward layer, and a fifth residual and layer normalization module connected in sequence; wherein, the first periodic feature output by the encoder layer is input to the encoder-decoder self-attention module, the initial periodic feature output by the first time series decomposition layer is input to the masked multi-head self-attention module and the third residual and layer normalization module, the periodic feature component output by the fourth time series decomposition layer is input to the encoder-decoder self-attention module and the fourth residual and layer normalization module, and the periodic feature component output by the fifth time series decomposition layer is input to the second feed-forward layer and the fifth residual and layer normalization module;
[0066] The second branch includes a sixth residual and layer normalization module and a seventh residual and layer normalization module connected in sequence; wherein, the initial trend feature output by the first time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fourth time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fifth time series decomposition layer is input to the seventh residual and layer normalization module, and the output of the seventh residual and layer normalization module is input to the fifth residual and layer normalization module;
[0067] The fifth residual and layer normalization module adds the output of the second feed-forward layer, the output of the fifth time series decomposition layer, and the output of the seventh residual and layer normalization module as the final prediction sequence.
[0068] It should be noted that the structures of all the time series decomposition layers in the decoder layer and the encoder layer are the same. The attention mechanisms adopted by the decoder layer and the encoder layer are both cross-dimensional attention mechanisms. The residual and layer normalization modules in the decoder layer and the encoder layer are Figure 2 collectively represented by the "+" symbol.
[0069] In this embodiment, a prediction model for pose time series is built based on an improved Transformer network structure. A deep decomposition architecture and a time series decomposition module are designed for the trend, periodic, and short-term dependence characteristics of pose time series prediction, so as to decompose the original input time series and then perform feature learning and prediction. Since the time series has the characteristics of short-term dependence and information sparsity, the present invention stacks discrete time points into data blocks (patches), retaining the short-term dependence. In order to reduce the computational complexity, the original attention mechanism is replaced with a cross-dimensional self-attention mechanism to extract the temporal features of the pose. Compared with the local receptive fields of convolutional neural networks, RNNs, LSTMs, etc., the prediction model of the present invention has a global receptive field, can directly capture the correlation between pose signal data blocks (Patches), and at the same time, based on the improved Transformer architecture, parallel computing can be realized, and the training and testing efficiency is higher.
[0070] The following are the measured prediction results of this embodiment:
[0071] The collected pose sequence data is divided into a training set, a validation set, and a test set according to a ratio of 6:2:2. The training and testing of the network model are carried out on a computer with a single NVIDIA RTX 4060 card. The model is built based on Python 3.9 and Pytorch 1.7.1. The initial learning rate for training is 0.0001, and a semi-decay training strategy of reducing the learning rate by 1 / 2 every 25 training rounds is adopted. The network model is trained for a total of 200 rounds, and stochastic gradient descent is performed using the Adam optimizer. The selected BatchSize is 32, and num_works is set to 4 to asynchronously load data. Each time, prediction data with a length of 1600 is input, and the prediction length is 400 time steps. Through actual measurement, the network model of the present invention is significantly superior to other models in terms of prediction accuracy. By transmitting the predicted value to the control software earlier, the synchronization of the position of the rehabilitation training platform and the patient's pose is improved, thus better exerting the exercise effect of the platform.
[0072] See Figure 3 , another embodiment of the present invention also provides a human motion intention prediction device 200 in rehabilitation training, including: a data acquisition module 201 and a prediction module 202. The human motion intention prediction device 200 can execute the human motion intention prediction method in the method embodiment.
[0073] Specifically, the human motion intention prediction device 200 includes:
[0074] A data acquisition module 201, configured to acquire the plantar pressure data and the pose data of each joint of the human body;
[0075] The prediction module 202 is configured to input the acquired plantar pressure data and each joint pose data in the form of a time series into a trained deep learning network model together to obtain a prediction result of the human motion intention;
[0076] Among them, the deep learning network model includes:
[0077] The data block embedding module is configured to process the discrete original input time series into data blocks of a predetermined length, perform position encoding and value encoding on the data blocks, and output the encoded data blocks to the encoder layer;
[0078] The encoder layer is configured to extract the first periodic features of the original input time series from the encoded data blocks and output the first periodic features to the decoder layer;
[0079] The first time series decomposition layer is configured to decompose the original input time series into initial periodic features and initial trend features, and input the initial periodic features and the initial trend features into the decoder layer;
[0080] The decoder layer is configured to calculate the total trend features and the total periodic features according to the first periodic features, the initial trend features and the initial periodic features, and add the total trend features and the total periodic features as the final prediction sequence.
[0081] It should be noted that the human motion intention prediction device 200 provided in this embodiment can be used to execute the technical solutions of the method embodiments. The implementation principle and technical effects are similar to those of the method, and will not be elaborated here.
[0082] The above description is only a preferred embodiment of the present invention. Those skilled in the art should understand that the disclosed scope of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with other technical features (but not limited to) having similar functions disclosed in the present invention.
Claims
1. A method for predicting human motion intention during rehabilitation training, characterized in that, Including: Collecting plantar pressure data and joint pose data of the human body; Inputting the obtained plantar pressure data and joint pose data into a trained deep learning network model in the form of a time series to obtain a prediction result of the human motion intention; Wherein, the deep learning network model includes: A data block embedding module for processing a discrete original input time series into data blocks of a predetermined length, performing positional encoding and value encoding on the data blocks, and outputting the encoded data blocks to an encoder layer; An encoder layer for extracting first-period features of the original input time series from the encoded data blocks and outputting the first-period features to a decoder layer; the encoder layer includes a plurality of serially connected encoders, and each encoder includes a multi-head self-attention module, a first residual and layer normalization module, a second time series decomposition layer, a first feed-forward layer, a second residual and layer normalization module, and a third time series decomposition layer connected in sequence; wherein, the output end of the data block embedding module is connected to the first residual and layer normalization module, and the output end of the second time series decomposition layer is connected to the second residual and layer normalization module; the third time series decomposition layer outputs the first-period features; the first residual and layer normalization module is used to form a skip connection between the input and output of the multi-head self-attention module; the second residual and layer normalization module is used to form a skip connection between the input and output of the first feed-forward layer; the multi-head self-attention module is used to calculate the correlation between the plantar pressure data and joint pose data in the original input time series; the second time series decomposition layer and the third time series decomposition layer are used to decompose the trend features and periodic features of their respective input data through a weighted moving average function; A first time series decomposition layer for decomposing the original input time series into initial periodic features and initial trend features through a weighted moving average function and inputting the initial periodic features and initial trend features into the decoder layer; A decoder layer for calculating total trend features and total periodic features according to the first-period features, initial trend features, and initial periodic features, and adding the total trend features and total periodic features as a final prediction sequence.
2. The method for predicting human motion intention in rehabilitation training according to claim 1, wherein: The decoder layer includes a plurality of serially connected decoders, and each decoder includes a first branch and a second branch; The first branch includes a masked multi-head self-attention module, a third residual and layer normalization module, a fourth time series decomposition layer, an encoder-decoder self-attention module, a fourth residual and layer normalization module, a fifth time series decomposition layer, a second feed-forward layer, and a fifth residual and layer normalization module connected in sequence; wherein, the first periodic feature output by the encoder layer is input into the encoder-decoder self-attention module, the initial periodic feature output by the first time series decomposition layer is input into the masked multi-head self-attention module and the third residual and layer normalization module, the periodic feature component output by the fourth time series decomposition layer is input into the encoder-decoder self-attention module and the fourth residual and layer normalization module, and the periodic feature component output by the fifth time series decomposition layer is input into the second feed-forward layer and the fifth residual and layer normalization module; The second branch includes a sixth residual and layer normalization module and a seventh residual and layer normalization module connected in sequence; wherein, the initial trend feature output by the first time series decomposition layer is input into the sixth residual and layer normalization module, the trend feature component output by the fourth time series decomposition layer is input into the sixth residual and layer normalization module, the trend feature component output by the fifth time series decomposition layer is input into the seventh residual and layer normalization module, and the output of the seventh residual and layer normalization module is input into the fifth residual and layer normalization module; The fifth residual and layer normalization module adds the output of the second feed-forward layer, the output of the fifth time series decomposition layer, and the output of the seventh residual and layer normalization module as the final prediction sequence.
3. A method for predicting human motion intention in rehabilitation training according to claim 2, characterized in that, The fourth time series decomposition layer and the fifth time series decomposition layer use a weighted moving average function to decompose the input data into periodic features and trend features.
4. A method for predicting human motion intention in rehabilitation training according to claim 3, characterized in that It further includes: A preprocessing module for performing stationary processing on the original input time series, eliminating non-stationary components, and outputting the processed data to the data block embedding module.
5. A device for predicting human motion intention during rehabilitation training, characterized in that, It includes: A data acquisition module for acquiring plantar pressure data and joint pose data of the human body; A prediction module for inputting the acquired plantar pressure data and joint pose data in the form of a time series into a trained deep learning network model to obtain a prediction result of the human motion intention; Wherein, the deep learning network model includes: A data block embedding module for processing the discrete original input time series into data blocks of a predetermined length, performing positional encoding and value encoding on the data blocks, and outputting the encoded data blocks to the encoder layer; The encoder layer is used to extract the first periodic features of the original input time series from the encoded data blocks and output the first periodic features to the decoder layer. The encoder layer includes multiple cascaded encoders, and each encoder includes a multi-head self-attention module, a first residual and layer normalization module, a second time series decomposition layer, a first feed-forward layer, a second residual and layer normalization module, and a third time series decomposition layer connected in sequence. Among them, the output end of the data block embedding module is connected to the first residual and layer normalization module, and the output end of the second time series decomposition layer is connected to the second residual and layer normalization module. The third time series decomposition layer outputs the first periodic features. The first residual and layer normalization module is used to form a skip connection between the input and output of the multi-head self-attention module. The second residual and layer normalization module is used to form a skip connection between the input and output of the first feed-forward layer. The multi-head self-attention module is used to calculate the correlation between the plantar pressure data and each joint pose data in the original input time series. The second time series decomposition layer and the third time series decomposition layer are used to decompose the trend features and periodic features of their respective input data through a weighted moving average function. The first time series decomposition layer is used to decompose the original input time series into initial periodic features and initial trend features through a weighted moving average function, and input the initial periodic features and initial trend features into the decoder layer. The decoder layer is used to calculate the total trend features and total periodic features according to the first periodic features, initial trend features and initial periodic features, and add the total trend features and total periodic features as the final prediction sequence.
6. The human motion intention prediction device in rehabilitation training according to claim 5, wherein: The decoder layer includes multiple cascaded decoders, and each decoder includes a first branch and a second branch. The first branch includes a masked multi-head self-attention module, a third residual and layer normalization module, a fourth time series decomposition layer, an encoder-decoder self-attention module, a fourth residual and layer normalization module, a fifth time series decomposition layer, a second feed-forward layer, and a fifth residual and layer normalization module connected in sequence. Among them, the first periodic features output by the encoder layer are input to the encoder-decoder self-attention module, the initial periodic features output by the first time series decomposition layer are input to the masked multi-head self-attention module and the third residual and layer normalization module, the periodic feature components output by the fourth time series decomposition layer are input to the encoder-decoder self-attention module and the fourth residual and layer normalization module, and the periodic feature components output by the fifth time series decomposition layer are input to the second feed-forward layer and the fifth residual and layer normalization module. The second branch includes a sixth residual and layer normalization module and a seventh residual and layer normalization module connected in sequence; wherein, the initial trend feature output by the first time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fourth time series decomposition layer is input to the sixth residual and layer normalization module, the trend feature component output by the fifth time series decomposition layer is input to the seventh residual and layer normalization module, and the output of the seventh residual and layer normalization module is input to the fifth residual and layer normalization module; The fifth residual and layer normalization module adds the output of the second feedforward layer, the output of the fifth time series decomposition layer, and the output of the seventh residual and layer normalization module as the final prediction sequence.
7. The human motion intention prediction device in rehabilitation training according to claim 6, characterized in that, The fourth time series decomposition layer and the fifth time series decomposition layer use a weighted moving average function to decompose the input data into periodic features and trend features.
8. A device for predicting human motion intention during rehabilitation training according to claim 7, characterized in that, It further includes: A preprocessing module for performing stationary processing on the original input time series, eliminating non-stationary components, and outputting the processed data to the data block embedding module.
Citation Information
Patent Citations
Attitude prediction method and device
CN108664122A
Three-dimensional human body posture estimation method and system based on human body topology sensing network
CN115908497A