A Human Motion Prediction Method Based on Multi-Level Spatiotemporal Information
By constructing binary joint groups with father-son relationships and multi-scale time evolution information processing, the problem of low prediction accuracy of human motion in the prior art is solved, and higher prediction accuracy and performance are achieved.
Patent Information
- Application Number
- CN202510413009.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing human motion prediction algorithms are difficult to fully obtain information related to joint space, which makes it difficult for the model to accurately understand human motion and generate inaccurate prediction results.
By constructing a binary joint group with a father-son relationship, aggregating the relevant features of the joints and joints and binary joint groups, a spatial information matrix is obtained; a parallel convolution layer is used to obtain temporal evolution information at different scales, and adaptive fusion and multi-head attention coding are performed to generate future human motion posture sequences.
It effectively enhances the richness of space-time related information, improves the accuracy of human movement prediction, and especially shows high performance in the prediction of long-term and high-random movements.
Smart Images

Figure CN119942650B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a human motion prediction method based on multi-level spatio-temporal information. Background Art
[0002] Human motion prediction is an important task in computer vision, and its purpose is to predict future human motions based on the observed human motions. This technology can endow intelligent devices with the ability to understand human motions and perceive future motion states. Currently, this technology has been widely applied to many fields such as autonomous driving, human-machine cooperation, and motion synthesis.
[0003] Currently, the methods for modeling the spatial dependence relationship between joints in human motion algorithms mainly include those based on convolutional networks, relational graphs, physical constraints, and graph convolutional networks, etc. Specifically, the model based on a convolutional network captures the relevant information between joints through a convolutional kernel. However, such algorithms are limited by the receptive field size of the convolutional kernel when capturing joint relationships, and it is difficult to reasonably model joint dependencies and will introduce noise, resulting in a decrease in accuracy. The method based on a relational graph predefines the explicit relationship between joints before model training according to prior knowledge such as human kinematics and anatomy. The method based on physical constraints encodes spatial features by setting geometric constraints between joint pairs. The above methods for predefining joint relationships can avoid the interference of negative information, but they also make the model lose the ability to flexibly capture implicit joint relationships that are difficult to detect. Most of the latest models are based on graph convolutional networks, and the model uses a relational matrix to flexibly encode joint dependencies from human motion data, so as to capture sufficient spatially relevant information. However, the way of completely capturing dependent features by the model itself may ignore some prior knowledge, bringing instability to training. Researchers have proposed to introduce some prior knowledge into the generation process of the relational matrix, including i) designing a multi-stage model, first setting the known physical relationships between joints, and then the model mines other joint dependency information from the data, ii) constructing a semi-constrained set of relational matrices, respectively defining the physical relationships, semantic relationships, and joint relationships captured by the model itself between joints. However, the above methods using relational matrices only consider the dependency relationships between binary joint pairs and fail to model the hierarchical structure existing between joints, resulting in the neglect of the multi-level dependency relationships existing between joints. Therefore, the existing methods are difficult to fully obtain the spatially relevant information of joints, resulting in the model being difficult to accurately understand human motions and generating inaccurate prediction results.
[0004] There is a strong temporal evolution relationship between future human movements and historical movements. Most current human movement prediction algorithms encode the temporal evolution relationship based on generative adversarial networks, temporal convolutional networks, recurrent neural networks, and Transformers. Among them, the model based on generative adversarial networks is mainly used to infer the possible temporal evolution relationship to generate relatively natural and reasonable movements, and can also generate various possible future movements and long-term movement predictions. However, the purpose of such models is to synthesize sequences that conform to human movement characteristics, and the accuracy of the generated prediction results is relatively low. The model based on temporal convolution effectively models the long-distance temporal evolution relationship by constructing a hierarchical model and has the ability to synthesize relatively accurate long-term movement prediction results. However, the model based on temporal convolution usually exhibits poor short-term prediction performance. The model based on recurrent neural networks captures the implicit temporal evolution information by reading the movement characteristics frame by frame in sequence and updating the hidden state of neurons. They circulate historical information in the entire memory framework and can model the movement context characteristics relatively accurately. However, the way of reading data by such algorithms is single and it is difficult to capture information at a long time distance, resulting in a relatively low long-term movement prediction accuracy of the model. Most of the latest human movement prediction models are based on Transformers to encode the temporal evolution relationship. It can evaluate the correlation between each frame in the entire movement sequence and the current frame and establish a temporal dependence relationship according to the importance of each frame of information. However, the current model based on Transformers can only evaluate the correlation of two frames of information at a time and aggregate features. It ignores the strong correlation temporal evolution relationship and multiple temporal evolution patterns in the local movement sequence of human movement. Therefore, the model that defines a single temporal evolution relationship only between two frames usually exhibits poor long-term prediction performance and serious inter-frame discontinuity in the results it generates. Summary of the Invention
[0005] The object of the present invention is to improve and innovate in view of the shortcomings and problems in the background technology, and provide a human movement prediction method based on multi-level spatio-temporal information.
[0006] According to the first aspect of the present invention, there is provided a human movement prediction method based on multi-level spatio-temporal information, specifically including the following steps:
[0007] Step S101, construct a binary joint group with a parent-child relationship between joints in each frame;
[0008] Step S102, aggregate the relevant features between joints and between joints and their binary joint groups in each frame, and obtain the spatial information matrix of each joint in each frame to capture complementary movement information from multiple spatial structure layers;
[0009] Step S103, splicing the spatial information matrices of each joint in each frame, and fusing the spliced spatial information matrices of each frame together to obtain the spatial information matrix corresponding to the human body motion posture sequence;
[0010] Step S104, using parallel convolution layers with different receptive fields to perform convolution operations on the spatial information matrix corresponding to the human motion posture sequence to obtain time evolution information of different scales; and mapping the time evolution information of different scales into time evolution information of the same scale as the spatial information matrix of the human motion posture sequence;
[0011] Step S105, adaptively fusing the time evolution information of uniform scale obtained after mapping;
[0012] Step S106: using a multi-head attention module to encode the adaptively fused time evolution information to obtain a human motion posture sequence containing time-space related information;
[0013] Step S107: predict and generate a future human body motion posture sequence based on the acquired spatiotemporal related information.
[0014] A further solution is that the specific operation formula of step S101 is as follows:
[0015] ,
[0016] In the formula, represents the binary joint group consisting of joint j and J joints of the human body posture in the Tth frame; , and The elements in are the coordinate values of joint 1, joint j and joint J on the X-axis, Y-axis and Z-axis respectively; represents the information correlation between the fusion feature and joint j after joint j and joint 1 form a binary joint group in the Tth frame, and is a learnable weight, is also a learnable weight, Represents the multiplication of corresponding matrix elements.
[0017] A further solution is that the operation formula of step S102 is as follows:
[0018] Tanh + + ),
[0019] in, The binary joint group used to evaluate the relevance of joint j to joint j, Used to evaluate the correlation between joints, Denotes the spatial information matrix of joint j in the T-th frame obtained after aggregating the relevant features of joint j with the J joints of the human pose and the binary joint group of joint j. Is the activation function. And Are both learnable weights. Denotes the matrix product.
[0020] A further solution is that the specific operation formula of step S104 is as follows:
[0021] ;
[0022] ;
[0023] ;
[0024] Denotes the spatial information matrix corresponding to the T-frame human motion pose sequence; and Is the mapping function. Is the convolution operation of the one-dimensional convolutional layer of the first receptive field. Is the convolution operation of the one-dimensional convolutional layer of the second receptive field. Is the convolution operation of the one-dimensional convolutional layer of the third receptive field. Denotes the time evolution information corresponding to the one-dimensional convolutional layer of the first receptive field. Denotes the time evolution information corresponding to the one-dimensional convolutional layer of the second receptive field. Denotes the time evolution information corresponding to the one-dimensional convolutional layer of the third receptive field.
[0025] A further solution is that step S104 specifically includes:
[0026] Transpose the spatial information matrix corresponding to the T-frame human motion pose sequence into a two-dimensional matrix, so that each row represents the eigenvalue of each joint on the same coordinate axis from the first frame to the T-th frame.
[0027] Perform convolution operations using three one-dimensional convolutional layers with different receptive fields to obtain multi-scale time evolution information.
[0028] Map the multi-scale two-dimensional matrix back to a two-dimensional matrix of a unified scale through a linear layer and transpose it into time evolution information of the same scale as the spatial information matrix corresponding to the human motion pose sequence.
[0029] A further solution is that the specific operation formula of step S105 is as follows:
[0030]
[0031] Among them, represents the T-frame motion posture sequence fused with various time evolution information, is the learnable weight corresponding to the time evolution information.
[0032] A further solution is that the specific operation formula of the step S106 is as follows:
[0033]
[0034] Among them, is the human motion posture sequence encoded with rich spatio-temporal related information, represents the s-th self-attention operation.
[0035] A further solution is that the specific operation formula of the step S107 is as follows:
[0036]
[0037] Among them, represents the predicted future t-frame human motion posture sequence.
[0038] According to the second aspect of the present invention, there is provided an electronic device, including: a memory and a processor;
[0039] The memory is used for storing programs;
[0040] The processor is used for calling the program stored in the memory to execute a human motion prediction method based on multi-level spatio-temporal information as described in any one of the above.
[0041] According to the third aspect of the present invention, there is provided a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, a human motion prediction method based on multi-level spatio-temporal information as described in any one of the above is implemented.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a human motion prediction method based on multi-level spatio-temporal information. By constructing a binary joint group with a parent-child relationship, the spatial structure of the binary joint group is modeled, and then the correlation between joints and between joints and the binary joint group is evaluated to capture complementary motion information from multiple spatial structure layers. Secondly, by encoding time context features at different scales, local-global time evolution information is obtained. Thirdly, by learning time information in multiple subspaces, rich time context is effectively obtained; finally, a gated recurrent unit is used to generate predictions for the future. The present invention captures spatially relevant information in various spatial structure levels such as joints and binary joint groups, and constructs time evolution information at multiple scales, effectively enhancing the richness of the obtained spatio-temporal relevant information, thereby improving the prediction accuracy. A two-stage multi-level spatio-temporal representation learning is adopted to perform the human motion prediction task, which includes two core stages: 1) Modeling the hierarchical structure of various human joints to assist the model in understanding the spatial cooperation of human motion and capturing complementary motion information from structures with spatial correlation relationships; 2) Modeling the human evolution features from multiple time scales to obtain rich temporal context information and capturing temporal evolution relationships therein to generate predictions for the future. A large number of experiments show that the present invention exhibits high performance on two large human motion benchmark datasets, namely Human 3.6M and CMU Mocap. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a schematic flowchart of a human motion prediction method based on multi-level spatio-temporal information provided by the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the detailed description of the specific embodiments of the present invention will be given below with reference to the drawings.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0047] Example 1
[0048] Please refer to Figure 1 , the present invention provides a human motion prediction method based on multi-level spatio-temporal information, which specifically includes the following steps:
[0049] Step S101: Construct a binary joint group with a parent-child relationship between joints in each frame;
[0050] Construct a binary joint group by setting a joint relationship matrix, and its definition is shown in the following formula:
[0051] ,
[0052] In this embodiment, when constructing a binary joint group where joint 1 has a parent-child relationship with other joints, the above formula is transformed into ,
[0053] It should be noted that the present invention takes the Human3.6M dataset as an example. In the Human3.6M dataset, the human motion posture is divided into 25 joint points; therefore, to construct a binary joint group where each joint has a parent-child relationship with other joints, 25 pairs of binary joint groups with parent-child relationships need to be constructed. In the formula, ∈ represents the binary joint group (including joint j itself) formed by joint j and the other 25 joints in the T-th frame; , and ∈ , , and The three elements in are the coordinate values of joint 1, joint j, and joint J on the X-axis, Y-axis, and Z-axis respectively; in this embodiment, J takes the value of 25; represents the information correlation between the fusion feature of joint j and joint 1 after forming a binary joint group and the information of joint j in the T-th frame, and ∈ is a learnable weight. ∈ is also a learnable weight, which is used to adjust the importance of the joint group feature in the information transmission process, further improving the flexibility of the model. represents the element-wise multiplication of the corresponding elements of the matrices, requiring the two matrices to have the same scale, also known as element-wise difference, and the symbol can also be ⊙.
[0054] In this embodiment, the above formula can be used to sequentially obtain the binary joint groups (including joint j itself) formed by joint j and the other 25 joints in the T-th frame, and obtain the binary joint groups with parent-child relationships between each joint and the other 25 joints in other frames. In this embodiment, binary joint groups with parent-child relationships between each joint and other joints in each frame of the 10-frame human body pose motion sequence can be constructed.
[0055] Step S102: Aggregate the relevant features of joints with joints and joints with their binary joint groups in each frame, and obtain the spatial information matrix of each joint in each frame to capture complementary motion information from multiple spatial structure layers;
[0056] To enable the model to obtain more sufficient spatially relevant features, by aggregating the relevant features of joints with joints and joints with their binary joint groups in each frame, complementary motion information is captured from multiple spatial structure layers, and the operation formula is as follows:
[0057] Tanh( + + ),
[0058] where, ∈ is used to evaluate the correlation between the binary joint group of joint j and joint j, represents the association between joints, represents the spatial information matrix of joint j in the T-th frame obtained after aggregating the relevant features of joint j and the other 25 joints and the binary joint group of joint j in the T-th frame, is the activation function.
[0059] In the formula, ∈ is a learnable weight, is also a learnable weight, represents the matrix product.
[0060] In this embodiment, for the spatial information matrix of joint 1 in the T-th frame obtained, the spatial information matrices of the 25 joints in the T-th frame can be sequentially obtained using the above formula, and the spatial information matrices of each joint in each frame can be obtained.
[0061] Step S103: Concatenate the spatial information matrices of each joint in each frame, and fuse these concatenated spatial information matrices of each frame together to obtain the spatial information matrix corresponding to the human motion pose sequence.
[0062] Since the spatial information matrices of each joint are a spatial information matrix of 3, and the human postures in the Human3.6M dataset used in the present invention include 25 joints; therefore, after splicing the spatial information matrices of each joint in each frame, 25 will be obtained a spatial information matrix of 3, where 25 each row of the information matrix of 3 corresponds to a spatial information matrix of a joint. After fusing the spatial information matrices of each frame together, a spatial information matrix corresponding to the human motion posture sequence will be obtained ,
[0063] In this embodiment, T is taken as 10 frames, then .
[0064] Step S104: Perform a convolution operation on the spatial information matrix corresponding to the human motion posture sequence by using convolutional layers with different receptive fields in parallel to obtain time evolution information of different scales; and map the time evolution information of different scales into time evolution information of the same scale as the spatial information matrix of the human motion posture sequence
[0065] Specifically, use three one-dimensional convolutional layers with different receptive fields in parallel to perform a convolution operation on the spatial information matrix corresponding to the human motion posture sequence, obtain time evolution information of multiple scales, and map the time evolution information of different scales into time evolution information of the same scale as the spatial information matrix of the human motion posture sequence. The specific operations are as follows:
[0066] ;
[0067] ;
[0068] ;
[0069] As described above, represents a T-frame human motion posture sequence that fuses the spatial information matrices of each frame; and is a mapping function, is the convolution operation of the one-dimensional convolutional layer, represents the time evolution information corresponding to the one-dimensional convolutional layer with the first receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer with the second receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer with the third receptive field.
[0070] It should be noted that before the convolution operation, first transpose the spatial information matrix corresponding to the T-frame human motion pose sequence into a two-dimensional matrix, so that each row represents the eigenvalue of each joint on the same coordinate axis from the first frame to the T-th frame; then perform the convolution operation using three one-dimensional convolutional layers with different receptive fields. Since the sizes of the one-dimensional convolutional layers are different, multi-scale temporal evolution information will be obtained after the convolution operation; then, map the multi-scale two-dimensional matrix back to a two-dimensional matrix of a unified scale through a linear layer, and finally transpose it into temporal evolution information of the same scale as the spatial information matrix corresponding to the original human motion pose sequence.
[0071] Since T takes the value of 10 frames, the above formula is transformed into:
[0072] ;
[0073] ;
[0074] ;
[0075] Among them, represents the spatial information matrix corresponding to the 10-frame human motion pose sequence by fusing each frame from the first frame to the tenth frame, 、 and all represent the obtained temporal evolution information.
[0076] For example: The three one-dimensional convolutional layers with different receptive fields respectively use 、 and convolution kernels of sizes. Before performing the convolution operation, first transpose to a 10 75 two-dimensional matrix, and then transpose it to a 75 10 two-dimensional matrix. Each row of the transposed represents the eigenvalue corresponding to each joint on the same coordinate axis from the first frame to the tenth frame; since there are a total of 25 joint points, and the coordinate information of each joint point includes the X-axis, Y-axis, and Z-axis, therefore, the transposed includes a total of 75 rows. Since three one-dimensional convolutional layers with different receptive fields are used for the convolution operation, the obtained convolution results will contain temporal evolution information of different scales, corresponding to 75 8, 75 6, and 75 4 two-dimensional matrices respectively; therefore, we can use a linear layer to convert the temporal evolution information of different scales into a 75 10 two-dimensional matrix; finally, we will 75 The two-dimensional matrix of 10 is mapped back to the time-evolution information with the same scale as the spatial information matrix of the human body motion posture sequence, that is, mapped into 10 25 a three-dimensional matrix of 3.
[0077] Step S105: Perform adaptive fusion on the time-evolution information with the unified scale obtained after mapping;
[0078] In order to better capture the time features closely related to prediction and better cooperate with the accurate extraction of relevant features, an adaptive fusion module is used to perform adaptive fusion on the time-evolution information with the unified scale obtained after mapping. The specific operations are as follows:
[0079]
[0080] Among them, represents the T-frame motion sequence fused with various time-evolution information. In this embodiment, are the learnable weights corresponding to various time-evolution information.
[0081] Step S106: Use the multi-head attention module to encode the time-evolution information after adaptive fusion to obtain the human body motion posture sequence containing spatio-temporal related information;
[0082] In order to more effectively obtain time information from a single human body posture motion sequence, five parallel self-attention heads are used to model time information in different subspaces respectively and perform an addition operation to obtain the human body sequence features with various complementary time series features and motion contexts. The specific operations are as follows:
[0083]
[0084] Among them, is the human body motion posture sequence containing rich spatio-temporal related information obtained by encoding, represents the s-th self-attention operation.
[0085] Step S107: Predict and generate the future human body motion posture sequence according to the obtained spatio-temporal related information;
[0086] The human body motion posture sequence containing rich spatio-temporal related information obtained through processing is input into the gated recurrent unit to generate a prediction for the future. The specific operations are as follows:
[0087]
[0088] Among them, represents the future t-frame human body motion posture sequence predicted and generated.
[0089] Since the present invention predicts and generates a future human body motion posture sequence of t frames, the number of frames of the output human body motion posture sequence is not necessarily equal to that of the input human body motion posture sequence. Therefore, in this embodiment, the GRU network adopts an N vs M network structure. The encoding module uses a number of gated recurrent units corresponding to the number of frames of the input human body motion posture sequence, and the decoding module uses a number of gated recurrent units corresponding to the number of frames of the output human body motion posture sequence. The encoding module outputs a context vector C, and the decoding module decodes the context vector C to predict and generate a future human body motion posture sequence of t frames.
[0090] For example: Input a human body motion posture sequence of 10 frames containing rich spatio-temporal correlation information into the GRU network to predict and generate a future human body motion posture sequence of 25 frames. Specifically, the first frame of the human body motion posture containing spatio-temporal correlation information is flattened into a vector of 75 × 1 and then input into the first gated recurrent unit. The second frame of the human body motion posture containing spatio-temporal correlation information is flattened into a vector of 75 × 1 and then input into the second gated recurrent unit, and so on. The tenth frame of the human body motion posture containing spatio-temporal correlation information is flattened into a vector of 75 × 1 and then input into the tenth gated recurrent unit. The gated recurrent unit corresponding to the previous frame outputs a hidden vector to the gated recurrent unit corresponding to the next frame; finally, the gated recurrent unit corresponding to the tenth frame outputs the context vector C to the decoding module, and the decoding module decodes the context vector C, so that the gated recurrent units of the decoding module sequentially output the human body motion postures of the 11th to 35th frames, thereby generating a future human body motion posture sequence of 25 frames.
[0091] In summary, the present invention provides a human motion prediction method based on multi-level spatio-temporal information. The present invention constructs a binary joint group with a parent-child relationship, models the spatial structure of the binary joint group, and then evaluates the correlation between joints and between joints and the binary joint group to capture complementary motion information from multiple spatial structure layers. Secondly, by encoding time context features at different scales, local-global time evolution information is obtained. Thirdly, by learning time information in multiple subspaces, rich time context is effectively obtained; finally, a gated recurrent unit is used to generate predictions for the future. The present invention effectively enhances the richness of the obtained spatio-temporal related information by capturing spatial related information in various spatial structure levels such as joints and binary joint groups, and constructing time evolution information at multiple scales, thereby improving the prediction accuracy. A two-stage multi-level spatio-temporal representation learning is adopted to perform the human motion prediction task, which includes two core stages: 1) Modeling the hierarchical structure of various human joints to assist the model in understanding the spatial cooperation of human motion and capturing complementary motion information from structures with spatial related relationships; 2) Modeling the human evolution features from multiple time scales to obtain rich temporal context information and capturing temporal evolution relationships therein to generate predictions for the future. A large number of experiments show that the present invention exhibits high performance on two large human motion benchmark datasets, Human 3.6M and CMU Mocap.
[0092] The present invention uses the Mean Per Joint Position Error (MPJPE) as an evaluation index to test the performance of different algorithms.
[0093] By detecting the model performance on the S5 file of the Human3.6M dataset, the prediction accuracy of each action obtained is summarized in Table 1. In comparison with existing human motion prediction algorithms such as LSTM3LR, Res-GRU, HP-GAN, Bi-GAN, ConSeq2Seq, HMR, LTD, DM-GNN, AVG, Traj-Net, and MMA, it is found that the prediction accuracy of the present invention is significantly better than the above algorithms in terms of 15 different actions and average accuracy. The present invention exhibits high performance on highly random actions. For example, the MPJPE of the present invention is 6.6 mm at 80 ms for the "posing" action and 9.5 mm at 160 ms for the "smoking" action. In addition, the average prediction error of the present invention at 1000 ms is 89.7 mm, which indicates that the present invention can also achieve high long-term prediction accuracy.
[0094] Table 1 Quantitative comparison results of the Human3.6M dataset
[0095]
[0096] To further evaluate the performance of the algorithm, CMU Mocap was used to detect the model performance. A total of six methods were used for evaluation, including LTD, Res-GRU, DMGNN, STSGCN, LPJP, and TIM. From the experimental results in Table 2, the average prediction accuracy of the present invention exceeds that of the current state-of-the-art methods. It is worth mentioning that for those challenging actions that are difficult to predict (e.g., playing basketball alone, walking, and wiping windows), the present invention also achieves a significant performance improvement.
[0097] Table 2 Quantitative comparison structure of CMU Mocap dataset
[0098]
[0099] Example 2
[0100] The present invention also provides an electronic device, including: a memory and a processor;
[0101] The memory is used to store programs;
[0102] The processor is used to call the programs stored in the memory to execute a human motion prediction method based on multi-level spatio-temporal information as described in Example 1.
[0103] Example 3
[0104] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes a human motion prediction method based on multi-level spatio-temporal information as described in Example 1.
[0105] Although the embodiments of the present invention have been disclosed as above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the embodiments shown and described herein.
Claims
1. A human motion prediction method based on multi-level spatiotemporal information, characterized in that: The specific steps include: Step S101, constructing a binary joint group with a parent-child relationship between joints in each frame; Step S102, aggregating the relevant features of the joints and the joints and their binary joint groups in each frame, obtaining the spatial information matrix of each joint in each frame, so as to capture complementary motion information from multiple spatial structure layers; Step S103, splicing the spatial information matrices of each joint in each frame, and fusing the spliced spatial information matrices of each frame together to obtain the spatial information matrix corresponding to the human body motion posture sequence; Step S104, using parallel convolutional layers with different receptive fields to perform convolution operations on the spatial information matrix corresponding to the human motion posture sequence to obtain time evolution information of different scales; And the time evolution information of different scales is mapped into the time evolution information of the same scale as the spatial information matrix of the human motion posture sequence; Step S105, adaptively fusing the time evolution information of uniform scale obtained after mapping; Step S106: using a multi-head attention module to encode the adaptively fused time evolution information to obtain a human motion posture sequence containing time-space related information; Step S107: predict and generate a future human body motion posture sequence based on the acquired spatiotemporal related information.
2. The method for predicting human motion based on multi-level spatiotemporal information according to claim 1, characterized in that: The specific operation formula of step S101 is as follows: , In the formula, represents the binary joint group consisting of joint j and J joints of the human body posture in the Tth frame; , and The elements in are the coordinate values of joint 1, joint j and joint J on the X-axis, Y-axis and Z-axis respectively; represents the information correlation between the fusion feature and joint j after joint j and joint 1 form a binary joint group in the Tth frame, and is a learnable weight, is also a learnable weight, Represents the multiplication of corresponding matrix elements.
3. The method for predicting human motion based on multi-level spatiotemporal information according to claim 2, characterized in that: The operation formula of step S102 is as follows: Fishy + + ), in, The binary joint group used to evaluate the relevance of joint j to joint j, Used to evaluate the correlation between joints, represents the spatial information matrix of joint j in the Tth frame obtained by aggregating the related features of joint j in the Tth frame and J joints in the human body posture and joint j and its binary joint group, is the activation function, and are all learnable weights, Represents matrix product.
4. The method for predicting human motion based on multi-level spatiotemporal information according to claim 3, characterized in that: The specific operation formula of step S104 is as follows: ; ; ; represents the spatial information matrix corresponding to the T-frame human motion posture sequence; and is the mapping function, is the convolution operation of the one-dimensional convolution layer of the first receptive field, is the convolution operation of the one-dimensional convolution layer of the second receptive field, is the convolution operation of the one-dimensional convolution layer of the third receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer of the first receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer of the second receptive field, Represents the time evolution information corresponding to the one-dimensional convolutional layer of the third receptive field.
5. The method for predicting human motion based on multi-level spatiotemporal information according to claim 4, characterized in that: The step S104 specifically includes: Transpose the spatial information matrix corresponding to the T-frame human motion posture sequence into a two-dimensional matrix so that each row represents the eigenvalue of each joint on the same coordinate axis from the first frame to the T-th frame; Three one-dimensional convolutional layers with different receptive fields are used for convolution operations to obtain multi-scale time evolution information; The multi-scale two-dimensional matrix is mapped back to a two-dimensional matrix of uniform scale through a linear layer and transposed into the time evolution information of uniform scale of the spatial information matrix corresponding to the human motion posture sequence.
6. A method for predicting human motion based on multi-level spatiotemporal information according to claim 4 or 5, characterized in that: The specific operation formula of step S105 is as follows: in, represents a T-frame motion posture sequence that integrates multiple time evolution information. is the learnable weight corresponding to the time evolution information.
7. The method for predicting human motion based on multi-level spatiotemporal information according to claim 6, characterized in that: The specific operation formula of step S106 is as follows: in, It is the encoded sequence of human motion postures containing rich spatiotemporal information. represents the s-th self-attention operation.
8. The method for predicting human motion based on multi-level spatiotemporal information according to claim 7, characterized in that: The specific operation formula of step S107 is as follows: in, Represents the predicted future t-frame human motion posture sequence.
9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to call the program stored in the memory to execute the human motion prediction method based on multi-level spatiotemporal information as described in any one of claims 1-8.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a human motion prediction method based on multi-level spatiotemporal information as described in any one of claims 1-8.
Citation Information
Patent Citations
Pedestrian detection method based on global-local features
CN118314606A
Human motion sequence prediction method and system based on skeleton enhanced Transform
CN119540279A