A sustainable learning method for human motion prediction

By introducing sustainable learning human motion prediction methods in human-computer interaction, using Bayesian sequence-sequence neural network and Monte Carlo Dropout technology, the problem of uncertainty and continuous learning in the existing technology is solved, and higher prediction accuracy and smaller knowledge forgetting are achieved.

CN114758195BActive Publication Date: 2025-05-23XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210505137.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2025-05-23
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

The existing technology cannot effectively model uncertainty in human-computer interaction, resulting in robots performing dangerous behaviors against unfamiliar human motion patterns, while being unable to continuously learn and adapt to new human motion patterns, resulting in insufficient performance in online human-computer interaction scenarios.

Method used

A sustainable learning method for human motion prediction is proposed, which manages and updates motion data through memory managers, policy samplers and parameter updaters, and uses Bayesian sequence-sequence neural networks and Monte Carlo Dropout technology to model and predict human motion to achieve uncertainty perception and continuous learning.

Benefits of technology

This method can make decisions more securely in human-computer interaction, have the ability to learn continuously, and will hardly forget existing knowledge, and show higher prediction accuracy and less knowledge forgetting in real environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114758195B_ABST
    Figure CN114758195B_ABST
Patent Text Reader

Abstract

The present invention discloses a sustainable learning method for predicting human motion. The method uses the motion trajectory of human joints captured by sensors as input, uses a recurrent neural network to give motion predictions for the next few seconds and their cognitive uncertainty and random uncertainty, and saves the captured motion trajectory so that the model can complete continuous learning training. The Bayesian neural network is used to model the various uncertainties of observed human motion to achieve safe online collection of interactive data. The memory management module maintains a fixed-size knowledge sample library in a limited memory space, the sample acquisition module performs data sampling in the knowledge sample library and the data stream, and the parameter update module enables the algorithm to have the ability to continuously learn based on the knowledge distillation algorithm. The present invention enables the robot to have online independent and continuous learning capabilities, and continuously improves the ability to predict human motion in interaction with people, so as to improve the safety and reliability of intelligent robot operations and interactions with people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a human motion prediction algorithm in human-computer interaction, and in particular to a sustainable learning human motion prediction method. Background Art

[0002] With the rapid development of deep learning, some deep learning-based methods have achieved remarkable performance in human motion prediction. In detail, some methods take advantage of recurrent neural networks and their variants (such as Seq2Seq), while others adopt generative models to generate future motions. However, the above methods may cause problems in real human-computer interaction scenarios due to the lack of the ability to provide uncertainty in their predictions. For the task of probabilistic human motion prediction, some traditional methods utilize methods such as Gaussian processes, hidden Markov models, and dynamic forest models. However, it is very difficult to apply them to large-scale datasets. Recently, some probabilistic methods based on deep learning learn multiple motion patterns by utilizing variational autoencoders and adversarial generative networks. These existing methods all assume that all possible human motion patterns are known. This assumption may lead to disastrous consequences in human-computer interaction scenarios because it ignores the diversity of human behavior. For example, when a robot observes an unfamiliar motion pattern, it may behave dangerously. In addition, the pre-collected training data is usually non-interactive and does not take into account the real-time reactions of humans, and the robot cannot reasonably respond to the interactive actions. Therefore, these methods cannot meet the requirements of online human-computer interaction scenarios.

[0003] Summary: (1) Existing methods are unable to model uncertainty: they are unable to recognize unfamiliar human motion patterns, causing the robot to make potentially dangerous interactive actions; (2) Existing methods are unable to conduct continuous / lifelong learning: they can only use existing pre-collected datasets for training. On the one hand, existing methods cannot use interaction data with humans to improve model accuracy; on the other hand, pre-collected data are all non-interactive, which is significantly different from the data collected in online interactions, resulting in insufficient performance of existing methods. Summary of the invention

[0004] In order to overcome the above-mentioned deficiencies in the prior art, the present invention provides a human motion prediction method with sustainable learning.

[0005] The technical solution of the present invention is achieved in this way:

[0006] A human motion prediction method with continuous learning includes a memory manager (cache management stage), a policy sampler (policy sampling stage), and a parameter updater (parameter updating stage).

[0007] First, save the human motion data stream collected interactively, at intervals of 10 minutes, at 25fps, it is expected that there will be a maximum of 15k time steps of data. Then use the 50+25 window sliding described in (1) to obtain a large amount of sample data) and send the data collected during this period to the strategy sampler. The strategy weight sampler calculates the response of the current neural network model to an internally maintained database and the newly collected data respectively. It uses the cognitive uncertainty of the response (Reflecting the model's familiarity with the sample) Determine the sampling weight of the sample. Epistemic uncertainty can also be understood as the sample variance of different predictions given by the model for the same input. The sampling weight is obtained by normalizing the epistemic uncertainty to the maximum and minimum.

[0008] Next, the present invention performs conventional neural network parameter updates according to the given sampling weight sampling samples, and the specific network structure is described in detail in the next paragraph. The commonly used gradient descent method is used in the parameter update stage (the present invention uses the AdamW optimization algorithm for optimization). In order to enable the neural network to retain existing knowledge without forgetting, the present invention adopts knowledge distillation technology. Specifically, the knowledge distillation loss function minimizes the distribution distance (KL divergence metric) between the output of the current neural network and the neural network before the update. The objective function of the present invention consists of two parts: regression loss and knowledge distillation loss. The specific calculation formula for the above loss is as follows:

[0009]

[0010]

[0011]

[0012] Finally, in the parameter update phase, a portion of the data is randomly selected from the packaged data to randomly replace the data in the internally maintained database. The number of replacement samples = cache size × the number of samples collected this time ÷ the total number of samples collected. The purpose of this is to maintain the consistency between the distribution of the internally maintained database and the distribution of human motion data. The above three steps complete a continuous learning process. When facing a human motion interaction data stream of infinite length, you only need to repeat the above three steps.

[0013] The composition of the probabilistic human motion prediction neural network model is described below. The probabilistic human motion prediction model of the present invention is based on a sequence-to-sequence recurrent neural network framework. The encoder unit first performs a temporal difference operation on the input part, and then combines the original data, the first-order difference and the second-order difference (concat) as the input of the recurrent neural network GRU, and the state output of the GRU is connected to the next recurrent unit after a Dropout operation. The decoder unit is similar to the encoder unit, except that the state output of the GRU will pass through another Dropout operation and then pass through a fully connected layer to make it output the predicted speed of each key point of the human body, and finally superimpose it on the position of the current time step input to obtain the predicted position of the key point of the human body in the next time step. The number of time steps of the encoder is related to the length of the input data. The present invention uses 2s*25fps for a total of 50 time steps. The number of time steps of the decoder is related to the expected prediction time length. The present invention uses 1s*25fps for a total of 25 time steps. It is worth noting that the Dropout operation used here is different from the conventional Dropout, which is called Monte Carlo Dropout. The present invention uses Bayesian neural network technology to transform it into a neural network whose weights obey a certain distribution. In engineering, the Bayesian neural network technology keeps each "Dropout" layer added to the original network during the test phase, which mathematically makes the network parameters obey the Bernoulli distribution. At the same time, for recurrent neural networks, the discarding method of this "Dropout" layer should remain the same at each time step. The engineering implementation usually uses the same mask / mask to perform the dropout operation at different time steps.

[0014] The present invention proposes a probabilistic human motion prediction method with continuous learning. (1) The present invention has uncertainty perception ability, which can help robots make decisions more safely. At the same time, this allows deep neural network models to be safely deployed in online scenarios using random initialization parameters. Compared with previous methods, the present invention makes human-computer interaction safer. (2) The present invention has continuous learning ability. This makes the model almost never forget existing knowledge and even perform better in previous tasks, which is difficult to achieve for previous human motion prediction methods. A major problem currently faced by deep learning methods is the cost of data collection and annotation. The present invention can automatically collect samples and learn, which can achieve zero data cost. In addition, compared with joint training where the training cost increases with the amount of data, the present invention processes infinitely long human motion data streams with a fixed training cost. In summary: (1) The present invention proposes a continuous learning method for probabilistic human motion prediction. The proposed method not only makes corresponding uncertainty predictions for the observed motion sequence, but also has the ability to continuously adapt to new human motion patterns. (2) The present invention has conducted experiments in a continuous learning setting on the existing dataset Human3.6m and in a real environment. The results show that our method performs better than other baseline methods: higher prediction accuracy and less knowledge forgetting. In the real environment, the present invention can learn the human kinematic model from scratch. This interaction method is effective and safe. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The uncertainty-aware probabilistic human motion prediction neural network of the present invention;

[0016] Figure 2 It is the continuous learning algorithm architecture of the present invention;

[0017] Figure 3 It is a comparison of the prediction error and forgetting after the continuous learning of the present invention and the related methods. The mean term is the average prediction error within 1 second, and BWT is a measure of the forgetting under the continuous learning setting. The larger the value, the greater the forgetting.

[0018] Figure 4 It is a comparison of the changes in prediction error and negative log-likelihood of the present invention and related methods during the continuous learning process. The numbers in brackets represent the additional cache usage of the present invention. DETAILED DESCRIPTION

[0019] The following is a specific implementation mode, that is, a detailed description of the present invention.

[0020] The invention is a human motion prediction method with sustainable learning, which consists of two parts: a human motion prediction model and a sustainable learning algorithm.

[0021] Figure 1 is the human motion prediction model of the present invention, which is composed of a Bayesian sequence-sequence neural network. Assume that in the past period of time, the robot captures human postures at fixed time intervals. Assume that during this period of time, the robot captures a total of T p The captured T p The overall human posture represents the human motion sequence in the past period of time, denoted by xT p : 0 . Assume that the past motion sequence of a human body collected at a certain time is x i , -T p :0, then the model inputs this past human motion sequence and predicts the future human motion sequence, recorded as x 1 :T f (i.e. T f future human movement posture).

[0022] The Bayesian sequence-sequence neural network mainly consists of an encoder and a decoder, supplemented by a Dropout layer. Both the encoder and the decoder are based on the gated recurrent unit (GRU). i , -T p :0After inputting the Bayesian sequence-sequence neural network, the encoder receives the hidden variable of the previous state (left arrow) and x i , -T p :0 in a posture x i,t , and output the hidden variable of the next state. Repeat T p After that, the encoder completes the encoding of the input motion sequence and outputs the hidden variables of the input sequence for decoding by the decoder. The decoder first receives the hidden variables from the encoder (middle arrow) and the last observed human posture x 0 , output the hidden variable of the next state (arrow on the right in the figure) and predict the distribution of the human body motion posture in the next frame In the next iteration, the decoder inputs the latent variables obtained in the previous iteration and the distribution of the predicted human motion posture (arrow on the right in the figure), and obtain the new distribution of predicted human motion postures Repeat T f After decoding, the distribution of the predicted human motion sequence is obtained Since the Bayesian neural network independently samples the weights from the distribution of parameters each time it makes a prediction, repeating the sampling J times will give the variance of the model's J-time prediction, i.e., epistemic uncertainty, denoted as

[0023] In addition, the encoder and decoder perform differential (DIFF) on the input human posture sequence to obtain the position, velocity and acceleration of the human posture. The output part of the GRU of the decoder is first transformed by the linear layer, and then the residual connection is made with the position, velocity and acceleration of the input human posture to obtain the distribution of the predicted human posture.

[0024] Figure 2 It is the continuous learning algorithm framework of the present invention, which is used to enable the previously described Bayesian sequence-sequence neural network to have the ability of continuous learning. The algorithm maintains a cache (Upper left block). Suppose the robot obtains a human motion sequence data X from the streaming data (lower line block) s In the strategy sampling phase (Sampling Weight), X s and cache All observed human motion data in the first pass through the weight allocation algorithm (SaWe) to obtain the sampling weight (lower left block). The SaWe algorithm uses the normalized data after calculating the model's cognitive uncertainty about the sample as the sampling weight of the sample. Then, when updating the parameters of the Bayesian sequence-sequence neural network, the data will be weighted sampled according to the weight to obtain the training sample (right side).

[0025] In the Bayesian sequence-sequence neural network parameter update phase (Update Parameter), the algorithm first saves the network model parameters before the update, denoted as θ′. The updated network parameters are denoted as θ. In the update phase, the model uses θ′ and θ to calculate the distillation loss Use θ to calculate the prediction loss The distillation loss measures how well the model learns X s The prediction loss measures the model's learning of X. s The final loss function is obtained by combining the distillation loss and the prediction loss. By optimizing the final loss function, the model can s Try to minimize the forgetting of previously learned knowledge.

[0026] After completing the X s After learning, in order to prevent the model from forgetting the knowledge learned in the current learning, the algorithm needs to go to the third stage, namely the buffer management stage. In the buffer management stage, the algorithm randomly selects from X s . Extract a certain amount of data and then randomly replace the original cache The samples in .

[0027] After the above three stages, the continuous learning algorithm completes one run. The Bayesian sequence-sequence neural network can not only predict the previously learned human motion patterns, but also predict the newly learned human motion patterns.

[0028] Figure 3 It is a comparison of the prediction error and forgetting after the continuous learning is completed between the present invention and the related methods, where the mean term is the average prediction error within 1 second, and BWT is a measure of the forgetting under the continuous learning setting, and the larger the value, the greater the forgetting. Figure 3 The purpose is to illustrate that compared with other related continuous learning methods, the method of the present invention has the highest prediction accuracy in the human motion prediction task and has the smallest forgetting index. Figure 3 As shown, it reports the prediction error and its average value within 1000ms of the present invention and related methods, followed by the forgetting indicator BWT after learning 15 tasks. "Rejection rate" is the percentage of predictions rejected by the uncertainty detector. Compared with other methods (iCaRL, A-GEM and GPM), the method without uncertainty detector reported the smallest prediction error (0.061). More importantly, the method of the present invention achieved the smallest BWT indicator (-0.006) among all continuous learning algorithms. Negative BWT indicates that the learning of subsequent tasks helps previous tasks. After the uncertainty detector rejected 5% (7.02% in the table) of the predictions, the present invention reported better prediction accuracy after using the uncertainty detector. There is a gap between all continuous learning algorithms and joint training. It is precisely because it is impossible to obtain all data at the same time in practical applications that joint learning cannot be performed, so joint training is considered to be the upper limit of continuous learning settings. However, the method of the present invention has the smallest performance gap (0.057 vs. 0.061). In summary, compared with other related continuous learning methods, the method of the present invention has the highest prediction accuracy in the human motion prediction task and the smallest forgetting indicator.

[0029] Figure 4 It is a comparison of the prediction error and negative log-likelihood of the present invention and related methods in the continuous learning process, and the numbers in brackets represent the cache size of the present invention. The present invention shows robustness to the cache size. Figure 4 It is shown that the present invention has the lowest prediction error and the highest uncertainty prediction accuracy compared with other methods in the entire continuous learning process.

Claims

1. A sustainable learning method for human motion prediction, It is characterized in that It includes cache management phase, strategy sampling phase and parameter update phase; it calculates samples based on the model’s cognitive uncertainty, including the data in the cache and the sampling weights collected online; The model uses knowledge distillation technology to update parameters for continuous learning; Update the cache to maintain data distribution of human motion patterns; Save human motion data in human-computer interaction; The cache management phase includes: Save the human motion data stream collected interactively. Every 10 minutes, at 25fps, it is expected that there will be a maximum of 15K time steps of data. Slide the window with time step 75, where the input time step is 50 and the corresponding real future time step is 25, and you can get several sample data. Send the data collected during this period to the strategy sampler. The policy weight sampler calculates the response of the current neural network model to an internally maintained database and newly collected data, respectively, using the epistemic uncertainty of the response. It reflects the model's familiarity with the sample and determines the sampling weight of the sample. Epistemic uncertainty is the sample variance of different predictions given by the model for the same input. The sampling weight is obtained by normalizing the epistemic uncertainty to the maximum and minimum values. In the parameter update phase, a portion of data is randomly selected from the packaged data to randomly replace the data in the internally maintained database. The number of replacement samples = cache size × the number of samples collected this time / the total number of samples collected; The above completes a continuous learning process. When faced with a human motion interaction data stream of infinite length, you only need to repeat the above steps.

2. A human motion prediction method for continuous learning according to claim 1, It is characterized in that In the strategy sampling stage, samples are sampled according to the sampling weights to perform regular neural network parameter updates. The parameter update stage adopts the commonly used gradient descent method, uses the AdamW optimization algorithm for optimization, and adopts the knowledge distillation technology.

3. A human motion prediction method for sustainable learning according to claim 2, It is characterized in that Specifically, the knowledge distillation loss function minimizes the distribution distance between the output of the current neural network and the neural network before the update, measured by KL divergence. The objective function consists of two parts: regression loss and knowledge distillation loss.

4. A human motion prediction method with sustainable learning according to claim 3, It is characterized in that The specific calculation formulas for regression loss and knowledge distillation loss are as follows: For regression loss, and is the expectation and variance of the Gaussian distribution of the human key points predicted by the model, x i,t is the real future human key point position corresponding to the sample, and is the expectation and variance of the Gaussian distribution of the human key points predicted by the model before updating, T f is the set prediction time length. When the sample belongs to the cache hour, is 1, otherwise it is 0. λ is a hyperparameter used to adjust the weights of the two parts of the loss function.

Citation Information

Patent Citations

  • Robot motion decision-making method, system and device introducing emotion regulation and control mechanism

    CN110119844A

  • Knowledge distillation optimization RNN short-term power failure prediction method, storage medium and equipment

    CN110232203A