Diversified human motion prediction method and device based on dynamic interactive hybrid anchor points
By adopting a dynamic interactive hybrid anchor method in human motion prediction, the problem of insufficient diversity and efficiency in the prior art is solved, and more accurate and diverse motion prediction effects are achieved, which is suitable for real-time applications.
Patent Information
- Application Number
- CN202510267045.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
The existing human motion prediction methods have insufficient diversity and efficiency, resulting in pattern collapse and high computing costs, making it difficult to adapt to real-time applications such as autonomous driving and virtual reality.
Using a dynamic interactive hybrid anchor method, the collected human motion data is transformed into skeleton viewpoints and base space generation, combined with sampling strategies and graph convolution network, the latent variables are decomposed into anchors and dynamically interacted to form interactive hybrid anchor information to improve the diversity and accuracy of the motion mode.
It realizes the diversified distribution of human motion prediction modes, improves the accuracy and efficiency of prediction, is suitable for real-time applications, and avoids the problems of pattern collapse and high computing costs.
Smart Images

Figure CN120093284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human motion analysis, and in particular to a method and device for predicting diversified human motion based on dynamic interactive hybrid anchor points. Background Art
[0002] At present, in the research of human motion recognition and motion prediction, machine learning methods include hidden Markov models and conditional random fields. However, deep learning methods are usually better than machine learning methods. Commonly used deep learning models include recurrent neural networks (RNN), generative adversarial networks (GAN), variational autoencoders (VAN), graph attention mechanisms (GAT) and graph convolutional neural networks (GCN). Most of these methods focus on the design of generative models to effectively learn distributed data. However, the model training data cannot cover all possible different movements. Therefore, during testing, sampling often focuses on the main data distribution mode, while ignoring the secondary data distribution mode, resulting in mode collapse and limiting the diversity of the output. Random sampling has poor sample efficiency, which means that a large number of samples need to be drawn to cover all modes, which is computationally expensive and may cause high latency, making it unsuitable for real-time applications such as autonomous driving and virtual reality. Therefore, in order to improve the diversity of human motion prediction, an efficient model can be designed to fully utilize historical motion information to improve technical applications.
[0003] CN202110902812.5 proposes a method and system for synthesizing human motion sequences based on motion graphs. First, key frames are extracted from human motion sequences to construct motion graphs, transition edges are added to the motion graphs, and transition subsequences corresponding to the transition edges are generated; at the same time, based on the motion graphs, a graph path search algorithm related to the number of transfers is designed for the task of completing human motion sequences, which is used to find multiple feasible paths from the starting node to the ending node, thereby generating a set of diverse motion sequence completion results; finally, for the task of predicting the motion sequence of the human body, a method of walking forward on the motion graph starting from the starting node is used, which can return multiple possible ending nodes, thereby reducing the task of predicting human motion to a task of completing human motion with some different ending nodes. The invention can generate real and diverse human motion sequences, and has excellent generalization ability between different data.
[0004] CN202210505137.7 discloses a sustainable learning method for predicting human motion, which takes the motion trajectory of human joints captured by sensors as input, uses a recurrent neural network to give motion predictions for the next few seconds and their cognitive uncertainty and random uncertainty, and saves the captured motion trajectory so that the model completes continuous learning training. Then the Bayesian neural network is used to model the various uncertainties of observed human motion to achieve safe online collection of interactive data. The memory management module maintains a fixed-size knowledge sample library in a limited memory space, the sample acquisition module performs data sampling in the knowledge sample library and the data stream, and the parameter update module is based on the knowledge distillation algorithm to enable the algorithm to have the ability to continuously learn. The invention enables the robot to have online independent and continuous learning capabilities, and continuously improves the ability to predict human motion in interaction with people, so as to improve the safety and reliability of intelligent robot operations and interactions with people.
[0005] CN202210817302.2 discloses a 3D human motion prediction method based on multi-scale feature fusion. 3D human motion prediction belongs to the problem of time series prediction. The transformer-based method has improved this problem and uses the attention mechanism to dynamically learn from the time series data to achieve good prediction results. The method was experimentally verified on human3.6M and CMP Mocap data, and the data prediction results generated were relatively ideal. The transformer-based multi-scale 3D human motion prediction method disclosed in this invention can better obtain the results of the motion prediction sequence.
[0006] At present, a lot of research has been carried out on human motion prediction at home and abroad, and good results have been achieved in human motion generation and prediction. However, most studies use deterministic methods to simulate human motion and regress a single future motion from past postures or video frames. They usually study the prediction and generation of the most likely motion patterns, and lack diversified generation pattern methods. In the research work that has been carried out, these methods cannot simulate the multimodal nature of human motion, resulting in motion pattern collapse, which is crucial for safety-critical applications. Therefore, the present invention proposes a method based on dynamic interactive hybrid anchor points, which aims to make the predicted pattern no longer focus on the most likely motion pattern, but to diversify the predicted motion pattern. At the same time, the method proposed by the present invention makes the predicted motion pattern more accurate and effective. Summary of the invention
[0007] Based on this, it is necessary to provide a diversified human motion prediction method and device based on dynamic interactive hybrid anchor points to address the existing problems.
[0008] In a first aspect, an embodiment of the present application provides a diversified human motion prediction method based on dynamic interactive hybrid anchor points, comprising the following steps:
[0009] S1: Collect human motion data;
[0010] S2: Perform skeleton perspective transformation on the collected human motion data to obtain transformed skeleton sequence data;
[0011] S3: generating a human joint information base space according to the skeleton sequence data;
[0012] S4: Use sampling strategy to obtain weight coefficient matrix;
[0013] S5: Combine the basis space and coefficient matrix, input the graph convolutional network to extract features and generate multiple latent variables;
[0014] S6: Decompose the latent variables into anchor points, and dynamically mix the anchor points to form interactive mixed anchor point information;
[0015] S7: Diversify motion prediction for the initial motion sequence, latent variables, and interactive hybrid anchor point information, and output multiple prediction sequence results.
[0016] Preferably, step S2 comprises:
[0017] The human skeleton sequence X with H frames is rotated counterclockwise around the x-axis, y-axis and z-axis in the original coordinate system to obtain three different angles to form a new observation angle α t , β t , γ t ;
[0018] Among them, the j-th skeleton joint point v′ of the t-th frame under the new observation angle t,j =[x′ t,j ,y′ t,j ,z′ t,j ] is expressed as:
[0019] v′ t,j =[x′ t,j ,y′ t,j ,z′ t,j ] T =R t ×v t,j (1);
[0020] Among them, v t,j is the jth skeleton joint point of the tth frame of the skeleton sequence X, and the joint set of the tth frame of the skeleton sequence X is represented by V t = {v t,1 ,...,v t,j};in, Represents β t Coordinate transformation around the y-axis, and Respectively represent the coordinate transformation of the original coordinate system around the x-axis and z-axis, which are expressed by the following formulas:
[0021]
[0022] Preferably, step S3 comprises:
[0023] In the network N with α as parameter α The joint set v′ after the input network skeleton perspective transformation t,i The spatial position of the i-th joint in the t-th frame Obtain the basis matrix of human joint spatiotemporal information
[0024] The calculation of the basis matrix B is as follows:
[0025] B=N α (P) (3);
[0026] Among them, each row of the matrix B is of dimension n b There are M basis vectors in total; these basis vectors together constitute the basis space of human joint information. Any point in the basis space can be obtained by linearly combining the basis vectors.
[0027] Preferably, step S4 comprises:
[0028] The human joint information base space is processed by a Gumbel-softmax sampling method to obtain a weight coefficient matrix W;
[0029] Among them, the weight coefficient matrix W is expressed by the following formula:
[0030]
[0031] Among them, U(0,1) is uniform distribution, g ij is a sample of the Gumbel distribution, π and τ are the parameters of the Gumbel distribution; the distribution coefficient matrix W∈R is generated in Gumbel-Softmax K×M In the process, W i , i∈[1,K] is each row of the weight coefficient matrix W, W i Contains M weights for combining basis vectors, and the M weights satisfy
[0032] Preferably, step S5 comprises:
[0033] Multiply the weight coefficient matrix W and the basis space matrix B to obtain a matrix
[0034] Using a multi-layer perceptron MLP, The receptive field in is mapped to The receptive field in
[0035] Generate K Gaussian distributions based on the observation sequence X;
[0036] Use the random variable ε to multiply the variance of the Gaussian distribution to achieve pattern diversification distribution and generate K latent variables:
[0037] Among them, the latent variable z k It is expressed by the following formula:
[0038] z k =A k ε+b k ,1≤k≤K (5);
[0039] in, are the variance and mean of k Gaussian distributions; n z For network N γ The dimension is expressed by the following formula:
[0040]
[0041] Each row in the matrix WB represents a sample sampled from the basis space, and a total of K samples are sampled from the basis space; X = [X 1 ,X 2 ,…,X h ]∈R H×J×C , H is the sequence length, J is the number of joint points, C is the number of features of each joint, and a new feature sequence X′=[X 1 ,X 2 ,…,X h ]∈R H×J×F , F is the number of features after joint point extraction.
[0042] Preferably, step S6 comprises:
[0043] The latent variable is decomposed into a random component sampled from a Gaussian prior distribution p(z) and a set of K learnable anchor points The deterministic component of the representation;
[0044] Select two different anchor point sets and Perform interactive mixing to form interactive mixing anchor point information;
[0045] Among them, the interactive hybrid anchor point information is expressed by the following formula:
[0046]
[0047] Among them, G represents a graph convolutional neural network with θ as a parameter, is the output of the graph convolutional hybrid network, z k is the noise randomly sampled from Gaussian distribution; a k is the kth learned anchor point, represents the motion features of the spatial anchor s in the kth prediction output layer, It represents the motion feature of the time anchor point at the kth prediction output layer, which is expressed by the following formula:
[0048]
[0049] Among them, γ is the weighting factor, and the spatial anchor point and time anchor Where K = K s ×K t ; It is expressed by the following formula:
[0050]
[0051] in, is the original sampled anchor point a k The corresponding blend anchor point.
[0052] in, It is expressed by the following formula:
[0053]
[0054] Preferably, step S7 includes:
[0055] The observed motion sequence Copy the last pose of T p times and perform motion sequence filling to obtain the filling sequence
[0056] Using the predefined M basis Filling sequence Perform DCT transformation to obtain motion transformation
[0057] Calculate the DCT coefficients of the motion sequence X and generate a function to predict future motion The DCT coefficients are then used to restore the predicted future motion sequence through inverse DCT.
[0058] in, Motion Transformation It is expressed by the following formula:
[0059]
[0060] Among them, the future motion sequence It is expressed by the following formula:
[0061]
[0062] C is the DCT coefficient matrix of the motion sequence X, and the kth trajectory X of the motion sequence k , the corresponding lth DCT coefficient is expressed by the following formula:
[0063]
[0064] Among them, l∈{1,2,…,K},δ ij represents the Kronecker function.
[0065] Preferably, the method further comprises:
[0066] Apply a loss function on the K predicted results after the input sequence X to train the sampling model;
[0067] Among them, the loss function L is expressed as follows:
[0068] L=λ d L d +λ KL L′ KL (14);
[0069] Among them, L d is the first loss function, L′ KL is the second loss function, λ d is the first hyperparameter, λ KL is the second hyperparameter;
[0070] Among them, the first loss function L d It is expressed by the following formula:
[0071]
[0072] Among them, the second loss function L′KL is expressed by the following formula:
[0073] L′ KL =KL(r β,γ (z k |X)||p(z)),k∈[1,K] (16);
[0074] Where η is the threshold value, is the K results predicted after the input sequence X, KL is the divergence, r β,γ (z k |X) is a network model N with parameters β and γ β and N γ Code z k potential distribution.
[0075] In a second aspect, the embodiment of the present application provides a diversified human motion prediction device based on dynamic interactive hybrid anchor points, including:
[0076] A data acquisition unit, used for collecting human motion data;
[0077] A sequence transformation unit is used to transform the collected human motion data into skeleton viewpoints to obtain transformed skeleton sequence data;
[0078] A joint generation unit, used for generating a human joint information base space according to the skeleton sequence data;
[0079] A matrix generation unit, used to obtain a weight coefficient matrix using a sampling strategy;
[0080] A data combining unit is used to combine according to the basis space and the coefficient matrix, input the graph convolution network to extract features and generate multiple latent variables;
[0081] Anchor point generation unit, used to decompose latent variables into anchor points, and dynamically mix the anchor points to form interactive mixed anchor point information;
[0082] The result output unit is used to diversify the motion prediction of the initial motion sequence, latent variables, and interactive mixed anchor point information, and output multiple prediction sequence results.
[0083] Compared with the prior art, the present invention has the following beneficial effects:
[0084] The present invention provides a diversified human motion prediction method and device based on dynamic interactive hybrid anchor points, which processes the collected human motion data to obtain a transformed skeleton sequence data, and further generates a basis space about human joint information; uses a sampling strategy to obtain a weight coefficient matrix; combines the obtained basis space and coefficient matrix, inputs a graph convolution network to extract features and generate multiple latent variables, and decomposes the latent variables into anchor points, allowing each anchor point to dynamically interact to make full use of skeleton information to diversify the distribution of motion patterns; finally, combines the initial motion sequence, latent variables, and interactive hybrid anchor point information to output multiple prediction sequence results. The present invention allows the predicted mode to no longer focus on the most likely motion mode, but diversifies the predicted motion mode. At the same time, the method proposed by the present invention makes the predicted motion mode more accurate and effective. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] By referring to the following drawings, the exemplary embodiments of the present invention can be more completely understood. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present invention and do not constitute a limitation of the present invention. In the drawings, the same reference numerals generally represent the same parts or steps.
[0086] Figure 1 An overall flow chart of a diversified human motion prediction method based on dynamic interactive hybrid anchor points provided according to an exemplary embodiment of the present application;
[0087] Figure 2 A flowchart of graph convolution feature extraction according to an exemplary embodiment of the present application;
[0088] Figure 3 A dynamic interactive hybrid anchor point graph provided according to an exemplary embodiment of the present application;
[0089] Figure 4 A schematic diagram of overall diversification prediction provided according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0090] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0091] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by technicians in the field to which this application belongs.
[0092] In addition, the terms "first" and "second" etc. are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.
[0093] The present application embodiment provides a diversified human motion prediction method based on dynamic interactive hybrid anchor points. The overall flow chart is as follows: Figure 1 As shown. Includes:
[0094] S1: Collect human motion data;
[0095] Assume a human skeleton sequence X with H frames, the joint set of the tth frame is represented as V t = {v t,1 ,...,v t,j}.
[0096] S2: Perform skeleton perspective transformation on the collected human motion data to obtain transformed skeleton sequence data;
[0097] The human skeleton sequence X with H frames is rotated counterclockwise around the x-axis, y-axis and z-axis in the original coordinate system to obtain three different angles to form a new observation angle α t , β t , γ t ;
[0098] Among them, the j-th skeleton joint point v′ of the t-th frame under the new observation angle t,j =[x′ t,j ,y′ t,j ,z′ t,j ] is expressed as:
[0099] v′ t,j =[x′ t,j ,y′ t,j ,z′ t,j ] T =R t ×v i,j (1);
[0100] Among them, v t,j is the jth skeleton joint point of the tth frame of the skeleton sequence X; Represents β t Coordinate transformation around the y-axis, and Respectively represent the coordinate transformation of the original coordinate system around the x-axis and z-axis, which are expressed by the following formulas:
[0101]
[0102] The observed viewpoint change is actually a rigid transformation, where all skeleton nodes in the tth frame follow the same set of transformation parameters, namely α t , β t , γ t Based on these transformation parameters, the above formula can be used to derive the skeleton representation in the new view coordinate system.
[0103] S3: Generate human joint information base space based on skeleton sequence data;
[0104] In most of the existing diversified human motion prediction methods, the prediction results of the network model are usually closely related to its parameters. This connection inevitably makes the prediction performance of the model intertwined with the learning process. Nevertheless, the network model is trained on the entire training data set, which means that the probability of predicting the motion pattern is distributed over the entire training data set. Although this is comprehensive, it also weakens the model's ability to make extreme predictions on minor motion patterns to a certain extent. The method designed by the present invention can reduce the occurrence of this situation.
[0105] Specifically, in the network N with α as the parameter α The joint set v′ after the input network skeleton perspective transformation t,i The spatial position of the i-th joint in the t-th frame Obtain the basis matrix of human joint spatiotemporal information
[0106] The calculation of the basis matrix B is as follows:
[0107] B=N α (P) (3);
[0108] Among them, each row of the matrix B is of dimension n b There are M basis vectors in total; these basis vectors together constitute the basis space of human joint information. Any point in the basis space can be obtained by linearly combining the basis vectors.
[0109] S4: Use sampling strategy to obtain weight coefficient matrix;
[0110] For the random sampling strategy, after the generating function is learned, some traditional methods will generate samples from the learned data distribution. It first randomly samples a set of latent variables from the potential prior distribution, and then uses the generating function to decode the latent variables into a set of data samples. This sampling method is more objective and can represent the overall distribution, but there are some problems. One is that the sampling strategy cannot simulate the exclusivity between diverse samples well; the other is that this sampling is random, and many samples may be repeated, resulting in the loss of some important information. In order to solve this challenge, the present invention uses the Gumbel-softmax sampling method to obtain a weight coefficient matrix.
[0111] Specifically, the human joint information base space is processed by a Gumbel-softmax sampling method to obtain a weight coefficient matrix W;
[0112] Among them, the weight coefficient matrix W is expressed by the following formula:
[0113]
[0114] Among them, U(0,1) is uniform distribution, g ij is a sample of the Gumbel distribution, π and τ are the parameters of the Gumbel distribution; the distribution coefficient matrix W∈R is generated in Gumbel-Softmax K×M In the process, W i , i∈[1,K] is each row of the weight coefficient matrix W, W i Contains M weights for combining basis vectors, and the M weights satisfy
[0115] S5: Combine the basis space and coefficient matrix, input the graph convolutional network to extract features and generate multiple latent variables;
[0116] Specifically, multiply the weight coefficient matrix W and the basis space matrix B to obtain a matrix
[0117] Using a multi-layer perceptron MLP, The receptive field in is mapped to The receptive field in
[0118] Generate K Gaussian distributions based on the observation sequence X;
[0119] Use the random variable ε to multiply the variance of the Gaussian distribution to achieve pattern diversification distribution and generate K latent variables:
[0120] Among them, the latent variable z k It is expressed by the following formula:
[0121] z k =A k ε+b k ,1≤k≤K (5);
[0122] in, are the variance and mean of k Gaussian distributions; n z For network N γ The dimension is expressed by the following formula:
[0123]
[0124] Each row in the matrix WB represents a sample sampled from the basis space, and a total of K samples are sampled from the basis space; X = [X 1 ,X 2 ,…,X h ]∈R H×J×C , H is the sequence length, J is the number of joint points, C is the number of features of each joint, and a new feature sequence X′=[X 1,X 2 ,…,X h ]∈R H×J×F , F is the number of features after joint point extraction. Figure 2 The process of graph convolution feature extraction is given.
[0125] S6: Decompose the latent variables into anchor points, and dynamically mix the anchor points to form interactive mixed anchor point information;
[0126] Inspired by interactive graph convolution, the present invention decomposes the latent variables in the basis space generated by the feedforward network model into a random component sampled from a Gaussian prior distribution p(z) and a set of K learnable network parameters, called anchor points These decomposed anchor points distribute as many motion modes as possible, which is achieved through a carefully designed optimization, while random noise further specifies the motion variations in certain modes.
[0127] Specifically, the latent variable is decomposed into a random component sampled from a Gaussian prior distribution p(z) and a set of K learnable anchor points The deterministic component of the representation;
[0128] Select two different anchor point sets and Perform interactive mixing to form interactive mixing anchor point information;
[0129] Among them, the interactive hybrid anchor point information is expressed by the following formula:
[0130]
[0131] Among them, G represents a graph convolutional neural network with θ as a parameter, is the output of the graph convolutional hybrid network, z k is the noise randomly sampled from Gaussian distribution; a k is the kth learned anchor point, represents the motion features of the spatial anchor s in the kth prediction output layer, It represents the motion feature of the time anchor point at the kth prediction output layer, which is expressed by the following formula:
[0132]
[0133] Among them, γ is the weighting factor, and the spatial anchor point and time anchor Where K = K s ×K t ; It is expressed by the following formula:
[0134]
[0135] in, is the original sampled anchor point a k The corresponding blend anchor point.
[0136] in, It is expressed by the following formula:
[0137]
[0138] Specifically, Figure 3 A schematic diagram of the dynamic hybrid interaction process of anchor points is given.
[0139] S7: Diversify motion prediction for the initial motion sequence, latent variables, and interactive hybrid anchor point information, and output multiple prediction sequence results.
[0140] Inspired by the motion of non-rigid structures, the present invention uses the trajectory representation of discrete cosine transform (DCT) to represent human motion. The purpose of this is that DCT transform can effectively remove high-frequency sequences and provide a more compact representation of motion patterns, thereby capturing the smoothness of human motion well.
[0141] Specifically, given the kth trajectory X of the observed human motion sequence k , the corresponding lth DCT coefficient can be calculated as:
[0142]
[0143] In formula (12), δ ij represents the Kronecker function, l∈{1,2,…,K}. Given the observed motion sequence First you need to copy T for the last pose p Perform motion sequence filling again to obtain the filling sequence Then, using the predefined M basis Perform DCT transformation, and the motion transformation is expressed as:
[0144]
[0145] Secondly, by calculating the DCT coefficients of the past motion sequence X, the generating function is used to predict the future motion The DCT coefficients are then inverted to restore the predicted future motion sequence:
[0146]
[0147] Among them, the last T of the recovery sequence p A frame represents a sequence of predicted future motion.
[0148] At the same time, in order to predict the accuracy of the model, is the K results predicted after inputting sequence X. The following loss function is applied to train the proposed sampling model. In order to improve the diversity distribution of the prediction mode, the present invention adopts the following diversity loss:
[0149]
[0150] In formula (15), η is a customizable threshold value, which is calculated by L d The loss function of the present invention makes the distance between any pair of generated predicted motion patterns not less than η. Compared with the diversity loss functions of other methods, the loss function used in the present invention is more efficient in improving the diversity of prediction patterns.
[0151] We also define L′ KL The loss function ensures that our network model predicts realistic and reasonable motion patterns, rather than those with high diversity but physically invalid motion patterns. The loss function is defined as follows:
[0152] L′ KL =KL(r β,γ (z k |X)||p(z)),k∈[1,K] (16);
[0153] In formula (16), KL is a method to measure the difference between two probability distributions, called KL divergence, r β,γ (z k |X) is a network model N with parameters β and γ β and N γ Code z k The potential distribution of , the final total training loss function is:
[0154] L=λ d L d +λ KL L′ KL (14);
[0155] Here, λ is a hyperparameter used to balance these two terms.
[0156] Finally, according to the above method, the initial sequence is transformed in perspective and the basis space is generated. At the same time, the weight matrix generated by the sampling strategy is combined with the basis space to generate a graph convolution hybrid space. Then, the latent variables are decomposed into hybrid anchor points composed of spatial anchor points and temporal anchor points. The anchor points can interact and make full use of node information to make the prediction mode more accurate and the mode distribution more extensive. Finally, multiple motion modes are generated through inverse DCT transformation. Overall diversified prediction such as Figure 4 shown.
[0157] The present invention provides a diversified human motion prediction method based on dynamic interactive hybrid anchor points, the method comprising: (1) processing the collected human motion data to obtain a transformed skeleton sequence data, and further generating a basis space about human joint information; (2) using a sampling strategy to obtain a weight coefficient matrix; (3) combining the obtained basis space and the coefficient matrix, inputting a graph convolutional network to extract features and generate multiple latent variables, and decomposing the latent variables into anchor points, allowing each anchor point to dynamically interact to make full use of skeleton information to diversify the distribution of motion patterns; (4) finally combining the initial motion sequence, latent variables, and interactive hybrid anchor point information to output multiple prediction sequence results. The present invention allows the predicted mode to no longer focus on the most likely motion mode, but diversifies the predicted motion mode. At the same time, the method proposed by the present invention makes the predicted motion mode more accurate and effective.
[0158] In the above embodiment, a method is provided, and correspondingly, the present application also provides a device. The device provided in the embodiment of the present application can implement the above method, and the device can be implemented by software, hardware, or a combination of software and hardware. For example, the device may include integrated or separate functional modules or units to perform the corresponding steps in the above methods.
[0159] In some implementation modes of the embodiments of the present application, the device provided by the embodiments of the present application is based on the same inventive concept as the method provided by the aforementioned embodiments of the present application and has the same beneficial effects.
[0160] Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0161] The device includes:
[0162] A data acquisition unit, used for collecting human motion data;
[0163] A sequence transformation unit is used to transform the collected human motion data into skeleton viewpoints to obtain transformed skeleton sequence data;
[0164] A joint generation unit, used for generating a human joint information base space according to skeleton sequence data;
[0165] A matrix generation unit, used to obtain a weight coefficient matrix using a sampling strategy;
[0166] A data combining unit is used to combine according to the basis space and the coefficient matrix, input the graph convolution network to extract features and generate multiple latent variables;
[0167] Anchor point generation unit, used to decompose latent variables into anchor points, and dynamically mix the anchor points to form interactive mixed anchor point information;
[0168] The result output unit is used to diversify the motion prediction of the initial motion sequence, latent variables, and interactive mixed anchor point information, and output multiple prediction sequence results.
[0169] It should be noted that the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0170] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0171] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0172] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0173] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0174] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and they should all be included in the scope of the claims and specification of the present application.
Claims
1. A diversified human motion prediction method based on dynamic interactive hybrid anchor points, characterized in that: The steps include: S1: Collect human motion data; S2: Perform skeleton perspective transformation on the collected human motion data to obtain transformed skeleton sequence data; S3: generating a human joint information base space according to the skeleton sequence data; S4: Use sampling strategy to obtain weight coefficient matrix; S5: Combine the basis space and coefficient matrix, input the graph convolutional network to extract features and generate multiple latent variables; S6: Decompose the latent variables into anchor points, and dynamically mix the anchor points to form interactive mixed anchor point information; S7: Diversify motion prediction for the initial motion sequence, latent variables, and interactive hybrid anchor point information, and output multiple prediction sequence results.
2. The method according to claim 1, characterized in that Step S2 includes: The human skeleton sequence X with H frames is rotated counterclockwise around the x-axis, y-axis and z-axis in the original coordinate system to obtain three different angles to form a new observation angle α t , β t , γ t ; Among them, the j-th skeleton joint point v′ of the t-th frame under the new observation angle t,j =[x′ t,j ,y′ t,j ,z′ t,j ] T It is expressed as: v′ t,j =[x′ t,j ,y′ t,j ,z t,j ] T =R t ×v t,j (1); Among them, v t,j is the jth skeleton joint point of the tth frame of the skeleton sequence X, and the joint set of the tth frame of the skeleton sequence X is represented by V t = {v t,1 ,...,v t,j };in, Represents β t Coordinate transformation around the y-axis, and Respectively represent the coordinate transformation of the original coordinate system around the x-axis and z-axis, expressed by the following formula:
3. The method according to claim 2, characterized in that Step S3 includes: In the network N with α as parameter α The joint set v′ after the input network skeleton perspective transformation t,i The spatial position of the i-th joint in the t-th frame Obtain the basis matrix of human joint spatiotemporal information The calculation of the basis matrix B is as follows: B=N α (P) (3); Among them, each row of the matrix B is of dimension n b There are M basis vectors in total; these basis vectors together constitute the basis space of human joint information. Any point in the basis space can be obtained by linearly combining the basis vectors.
4. The method according to claim 3, characterized in that Step S4 includes: The human joint information base space is processed by a Gumbel-softmax sampling method to obtain a weight coefficient matrix W; Among them, the weight coefficient matrix W is expressed by the following formula: Among them, U(0,1) is uniform distribution, g ij is a sample of the Gumbel distribution, π and τ are the parameters of the Gumbel distribution; the distribution coefficient matrix W∈R is generated in Gumbel-Softmax K×M In the process, W i , i∈[1,K] is each row of the weight coefficient matrix W, W i Contains M weights for combining basis vectors, and the M weights satisfy 5. The method according to claim 4, characterized in that Step S5 includes: Multiply the weight coefficient matrix W and the basis space matrix B to obtain a matrix Using a multi-layer perceptron MLP, The receptive field in is mapped to The receptive field in Generate K Gaussian distributions based on the observation sequence X; Use the random variable ε to multiply the variance of the Gaussian distribution to achieve pattern diversification distribution and generate K latent variables: Among them, the latent variable z k It is expressed by the following formula: z k =A k ε+b k ,1≤k≤K (5); in, are the variance and mean of k Gaussian distributions; n z For network N γ The dimension is expressed by the following formula: Each row in the matrix WB represents a sample sampled from the basis space, and a total of K samples are sampled from the basis space; X = [X1, X2, ..., X h ]∈R H×J×C , H is the sequence length, J is the number of joint points, C is the number of features of each joint, and a new feature sequence X′=[X1,X2,…,X h ]∈R H×J×F , F is the number of features after joint point extraction.
6. The method according to claim 5, characterized in that Step S6 includes: The latent variable is decomposed into a random component sampled from a Gaussian prior distribution p(z) and a set of K learnable anchor points The deterministic component of the representation; Select two different anchor point sets and Perform interactive mixing to form interactive mixing anchor point information; Among them, the interactive hybrid anchor point information is expressed by the following formula: Among them, G represents a graph convolutional neural network with θ as a parameter, is the output of the graph convolutional hybrid network, z k is the noise randomly sampled from Gaussian distribution; a k is the kth learned anchor point, represents the motion features of the spatial anchor s in the kth prediction output layer, It represents the motion feature of the time anchor point at the kth prediction output layer, which is expressed by the following formula: Among them, γ is the weighting factor, and the spatial anchor point and time anchor Where K = K s ×K t ; It is expressed by the following formula: in, is the original sampled anchor point a k The corresponding blend anchor point. in, It is expressed by the following formula:
7. The method according to claim 6, characterized in that Step S7 includes: The observed motion sequence Copy the last pose of T p times and perform motion sequence filling to obtain the filling sequence Using the predefined M basis Filling sequence Perform DCT transformation to obtain motion transformation Calculate the DCT coefficients of the motion sequence X and generate a function to predict future motion The DCT coefficients are then used to recover the predicted future motion sequence through inverse DCT. in, Motion Transformation It is expressed by the following formula: Among them, the future motion sequence It is expressed by the following formula: C is the DCT coefficient matrix of the motion sequence X, and the kth trajectory X of the motion sequence k , the corresponding lth DCT coefficient is expressed by the following formula: Among them, l∈{1,2,…,K},δ ij represents the Kronecker function.
8. The method according to claim 7, characterized in that The method further comprises: Apply a loss function on the K predicted results after the input sequence X to train the sampling model; Among them, the loss function L is expressed as follows: L=λ d L d +λ KL L′ KL (14); Among them, L d is the first loss function, L′ KL is the second loss function, λ d is the first hyperparameter, λ KL is the second hyperparameter; Among them, the first loss function L d It is expressed by the following formula: Among them, the second loss function L′ KL It is expressed by the following formula: L′ KL =KL(r β,γ (from k |X)||p(z)),k∈[1,K] (16); Where η is the threshold value, is the K results predicted after the input sequence X, KL is the divergence, r β,γ (z k |X) is a network model N with parameters β and γ β and N γ Code z k potential distribution.
9. A diversified human motion prediction device based on dynamic interactive hybrid anchor points, characterized in that: include: A data acquisition unit, used for collecting human motion data; A sequence transformation unit is used to transform the collected human motion data into skeleton viewpoints to obtain transformed skeleton sequence data; A joint generation unit, used for generating a human joint information base space according to the skeleton sequence data; A matrix generation unit, used to obtain a weight coefficient matrix using a sampling strategy; A data combining unit is used to combine according to the basis space and the coefficient matrix, input the graph convolution network to extract features and generate multiple latent variables; Anchor point generation unit, used to decompose latent variables into anchor points, and dynamically mix the anchor points to form interactive mixed anchor point information; The result output unit is used to diversify the motion prediction of the initial motion sequence, latent variables, and interactive mixed anchor point information, and output multiple prediction sequence results.
Citation Information
Patent Citations
Human motion prediction method capable of realizing continuous learning
CN114758195A
3D human body motion prediction method based on multi-scale feature fusion
CN115188075A
Human motion sequence synthesis method and system based on motion diagram
CN115705755A