Human motion prediction method based on multilevel spatio-temporal information
By constructing binary joint groups and multi-scale time evolution information processing, the problem of low prediction accuracy of human motion in the prior art is solved, and higher prediction accuracy and performance are achieved.
Patent Information
- Application Number
- CN202510413009.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing human motion prediction algorithms are difficult to fully obtain information related to joint space, which makes it difficult for the model to accurately understand human motion and generate inaccurate prediction results.
By constructing a binary joint group with a father-son relationship, aggregating the relevant features of the joints and joints and binary joint groups, a spatial information matrix is obtained; then a parallel convolutional layer is used to obtain temporal evolution information at different scales, and adaptive fusion and multi-head attention coding are performed, and future motion predictions are generated using gated loop units.
It effectively enhances the richness of space-time related information, improves the accuracy of human movement prediction, and especially shows high performance in the prediction of long-term and high-random movements.
Smart Images

Figure CN119942650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and in particular to a method for predicting human motion based on multi-level spatiotemporal information. Background Art
[0002] Human motion prediction is an important task in computer vision, and its purpose is to predict future human motion based on observed human motion. This technology can give smart devices the ability to understand human motion and perceive future motion states. Currently, this technology has been widely used in many fields such as autonomous driving, human-machine collaboration, and motion synthesis.
[0003] At present, the methods for modeling spatial dependencies between joints in human motion algorithms mainly include those based on convolutional networks, relationship graphs, physical constraints, and graph convolutional networks. Specifically, models based on convolutional networks use convolution kernels to capture relevant information between joints. However, such algorithms are limited by the size of the receptive field of the convolution kernel when capturing joint relationships, making it difficult to reasonably model joint dependencies and introducing noise, resulting in reduced accuracy. The relationship graph-based method predefines explicit relationships between joints before model training based on prior knowledge such as human kinematics and anatomy. The physical constraint-based method encodes spatial features by setting geometric constraints between joint pairs. The above-mentioned method of predefines joint relationships can avoid the interference of negative information, but it also makes the model lose the ability to flexibly capture imperceptible implicit joint relationships. Most of the latest models are based on graph convolutional networks, in which the model uses relationship matrices to flexibly encode joint dependencies from human motion data, thereby capturing sufficient spatially relevant information. However, the method of capturing dependency features entirely by the model itself may ignore some prior knowledge, causing instability in training. The researchers proposed to introduce some prior knowledge into the generation process of the relationship matrix, including i) designing a multi-stage model, first setting the known physical relationship between joints, and then the model mining other joint dependency information from the data, ii) constructing a set of semi-constrained relationship matrices to define the physical relationship, semantic relationship and joint relationship captured by the model. However, the above method of using the relationship matrix only considers the dependency between binary joint pairs, fails to model the hierarchical structure between joints, and causes the multi-level dependency between joints to be ignored. Therefore, it is difficult for existing methods to fully obtain information related to the joint space, which makes it difficult for the model to accurately understand human motion and generate inaccurate prediction results.
[0004] There is a strong temporal evolution relationship between future human motion and historical motion. Most current human motion prediction algorithms encode temporal evolution relationships based on generative adversarial networks, temporal convolutional networks, recurrent neural networks, and transformers. Among them, models based on generative adversarial networks are mainly used to infer possible temporal evolution relationships to generate more natural and reasonable motions, and can also generate a variety of possible future motions and long-term motion predictions. However, the purpose of such models is to synthesize sequences that conform to human motion characteristics, and the generated prediction results are of low accuracy. Models based on temporal convolutions effectively model long-distance temporal evolution relationships by constructing hierarchical models, and have the ability to synthesize more accurate long-term motion prediction results. However, models based on temporal convolutions usually show poor short-term prediction performance. Models based on recurrent neural networks capture implicit temporal evolution information by sequentially reading motion features frame by frame and updating the hidden states of neurons. They circulate historical information throughout the memory frame, and can more accurately model motion context features. However, such algorithms have a single way of reading data and are difficult to capture information over a long time distance, resulting in low long-term motion prediction accuracy of the model. Most of the latest human motion prediction models are based on Transformer to encode the temporal evolution relationship. It can evaluate the degree of correlation between each frame in the entire motion sequence and the current frame, and establish a temporal dependency relationship based on the importance of each frame's information. However, the current Transformer-based model can only evaluate the correlation between two frames of information at a time and aggregate features. It ignores the characteristics of human motion, such as the strongly correlated temporal evolution relationship and multiple temporal evolution patterns in the local motion sequence. Therefore, the results generated by the model that only defines a single temporal evolution relationship between two frames usually show poor long-term prediction performance and severe inter-frame discontinuity. Summary of the invention
[0005] The purpose of the present invention is to improve and innovate the shortcomings and problems existing in the background technology and provide a human body motion prediction method based on multi-level spatiotemporal information.
[0006] According to a first aspect of the present invention, a method for predicting human motion based on multi-level spatiotemporal information is provided, which specifically comprises the following steps: Step S101, constructing a binary joint group with a parent-child relationship between joints in each frame; Step S102, aggregating the relevant features of the joints and the joints and their binary joint groups in each frame, obtaining the spatial information matrix of each joint in each frame, so as to capture complementary motion information from multiple spatial structure layers; Step S103, splicing the spatial information matrices of each joint in each frame, and fusing the spliced spatial information matrices of each frame together to obtain the spatial information matrix corresponding to the human body motion posture sequence; Step S104, using parallel convolution layers with different receptive fields to perform convolution operations on the spatial information matrix corresponding to the human motion posture sequence to obtain time evolution information of different scales; and mapping the time evolution information of different scales into time evolution information of the same scale as the spatial information matrix of the human motion posture sequence; Step S105, adaptively fusing the time evolution information of uniform scale obtained after mapping; Step S106: using a multi-head attention module to encode the adaptively fused time evolution information to obtain a human motion posture sequence containing time-space related information; Step S107: predict and generate a future human body motion posture sequence based on the acquired spatiotemporal related information.
[0007] A further solution is that the specific operation formula of step S101 is as follows: , In the formula, represents the binary joint group consisting of joint j and J joints of the human body posture in the Tth frame; , and The elements in are the coordinate values of joint 1, joint j and joint J on the X-axis, Y-axis and Z-axis respectively; represents the information correlation between the fusion feature and joint j after joint j and joint 1 form a binary joint group in the Tth frame, and is a learnable weight, is also a learnable weight, Represents the multiplication of corresponding matrix elements.
[0008] A further solution is that the operation formula of step S102 is as follows: Tanh + + ), in, The binary joint group used to evaluate the relevance of joint j to joint j, Used to evaluate the correlation between joints, represents the spatial information matrix of joint j in the Tth frame obtained by aggregating the related features of joint j in the Tth frame and J joints in the human body posture and joint j and its binary joint group, is the activation function, and are all learnable weights, Represents matrix product.
[0009] A further solution is that the specific operation formula of step S104 is as follows: ; ; ; represents the spatial information matrix corresponding to the T-frame human motion posture sequence; and is the mapping function, is the convolution operation of the one-dimensional convolution layer of the first receptive field, is the convolution operation of the one-dimensional convolution layer of the second receptive field, is the convolution operation of the one-dimensional convolution layer of the third receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer of the first receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer of the second receptive field, Represents the time evolution information corresponding to the one-dimensional convolutional layer of the third receptive field.
[0010] A further solution is that step S104 specifically includes: Transpose the spatial information matrix corresponding to the T-frame human motion posture sequence into a two-dimensional matrix so that each row represents the eigenvalue of each joint on the same coordinate axis from the first frame to the T-th frame; Three one-dimensional convolutional layers with different receptive fields are used for convolution operations to obtain multi-scale time evolution information; The multi-scale two-dimensional matrix is mapped back to a two-dimensional matrix of uniform scale through a linear layer and transposed into the time evolution information of uniform scale of the spatial information matrix corresponding to the human motion posture sequence.
[0011] A further solution is that the specific operation formula of step S105 is as follows:
[0012] in, represents a T-frame motion posture sequence that integrates multiple time evolution information. is the learnable weight corresponding to the time evolution information.
[0013] A further solution is that the specific operation formula of step S106 is as follows:
[0014] in, It is the encoded sequence of human motion postures containing rich spatiotemporal information. represents the s-th self-attention operation.
[0015] A further solution is that the specific operation formula of step S107 is as follows:
[0016] in, Represents the predicted future t-frame human motion posture sequence.
[0017] According to a second aspect of the present invention, there is provided an electronic device, comprising: a memory and a processor; The memory is used to store programs; The processor is used to call the program stored in the memory to execute a human motion prediction method based on multi-level spatiotemporal information as described in any one of the above items.
[0018] According to a third aspect of the present invention, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a method for predicting human motion based on multi-level spatiotemporal information as described in any one of the above items is implemented.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention provides a method for predicting human motion based on multi-level spatiotemporal information. The present invention constructs a binary joint group with a parent-child relationship, models the spatial structure of the binary joint group, and then evaluates the correlation between joints and binary joint groups to capture complementary motion information from multiple spatial structural layers. Secondly, local-global time evolution information is obtained by encoding temporal context features of different scales. Thirdly, rich temporal context is effectively obtained by learning temporal information in multiple subspaces; finally, a gated recurrent unit is used to generate predictions for the future. The present invention effectively enhances the richness of the acquired spatiotemporal related information by capturing spatially related information in multiple spatial structural levels such as joints and binary joint groups, and constructing temporal evolution information at multiple scales, thereby improving the prediction accuracy. A two-stage multi-level spatiotemporal representation learning is used to perform human motion prediction tasks, which includes two core stages: 1) Modeling the hierarchical structure of multiple human joints to assist the model in understanding the spatial coordination of human motion and capturing complementary motion information from structures with spatial correlation; 2) Modeling the human evolution characteristics from multiple time scales to obtain rich temporal context information, and capturing temporal evolution relationships to generate future predictions. A large number of experiments show that the present invention has shown high performance on two large-scale human motion benchmark datasets, Human 3.6M and CMU Mocap. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 It is a flowchart of a method for predicting human motion based on multi-level spatiotemporal information provided by the first embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0024] Example 1 See also Figure 1 The present invention provides a method for predicting human motion based on multi-level spatiotemporal information, which specifically comprises the following steps: Step S101, constructing a binary joint group with a parent-child relationship between joints in each frame; A binary joint group is constructed by setting the joint relationship matrix, which is defined as follows: , In this embodiment, when constructing a binary joint group in which joint 1 has a parent-child relationship with other joints, the above formula is converted into , It should be noted that the present invention takes the Human3.6M dataset as an example. In the Human3.6M dataset, human motion postures are divided into 25 joint points; therefore, to construct a binary joint group in which each joint has a parent-child relationship with other joints, 25 pairs of binary joint groups with a parent-child relationship need to be constructed. In the formula, ∈ represents the binary joint group consisting of joint j and the other 25 joints (including joint j itself) in the Tth frame; , and ∈ , , and The three elements in are the coordinate values of joint 1, joint j and joint J on the X-axis, Y-axis and Z-axis respectively; in this embodiment, the value of J is 25; represents the information correlation between the fusion feature and joint j after joint j and joint 1 form a binary joint group in the Tth frame, and ∈ is a learnable weight. ∈ It is also a learnable weight that is used to adjust the importance of joint group features in the information transmission process and further enhance the flexibility of the model. It represents the multiplication of corresponding matrix elements. The two matrices must have the same scale. It is also called element-by-element difference. The symbol can also be ⊙.
[0025] In this embodiment, the above formula can be used to sequentially obtain the binary joint group consisting of joint j and other 25 joints in the Tth frame (including joint j itself), and obtain the binary joint group in which each joint has a parent-child relationship with the other 25 joints in other frames. In this embodiment, a binary joint group in which each joint has a parent-child relationship with other joints in each frame of a 10-frame human posture motion sequence can be constructed.
[0026] Step S102: Aggregate the relevant features of the joints and joints and the joints and their binary joint groups in each frame to obtain the spatial information matrix of each joint in each frame, so as to capture complementary motion information from multiple spatial structure layers; In order to enable the model to obtain more sufficient spatial correlation features, the relevant features of joints and joints and joints and their binary joint groups in each frame are aggregated to capture complementary motion information from multiple spatial structure layers. The operation formula is as follows: Tanh + + ), in, ∈ The binary joint group used to evaluate the relevance of joint j to joint j, Indicates the relationship between joints. represents the spatial information matrix of joint j in the Tth frame obtained by aggregating the relevant features of joint j and other 25 joints and joint j and its binary joint group in the Tth frame, is the activation function.
[0027] In the formula, ∈ is a learnable weight, is also a learnable weight, Represents matrix product.
[0028] In this embodiment, the spatial information matrix of joint 1 in the Tth frame is obtained, and the above formula can be used to sequentially obtain the spatial information matrices of 25 joints in the Tth frame, and obtain the spatial information matrix of each joint in each frame.
[0029] Step S103, splicing the spatial information matrices of each joint in each frame, and fusing the spliced spatial information matrices of each frame together to obtain the spatial information matrix corresponding to the human body motion posture sequence.
[0030] Since the spatial information matrix of each joint for 3, and the human body posture in the Human3.6M data set used in the present invention includes 25 joints; therefore, after splicing the spatial information matrices of each joint in each frame, 25 3 spatial information matrix, where 25 Each row of the information matrix of 3 corresponds to the spatial information matrix of a joint. After fusing the spliced spatial information matrices of each frame together, the spatial information matrix corresponding to the human body motion posture sequence will be obtained. , In this embodiment, T is set to 10 frames. .
[0031] Step S104, using parallel convolution layers with different receptive fields to perform convolution operations on the spatial information matrix corresponding to the human motion posture sequence to obtain time evolution information of different scales; and mapping the time evolution information of different scales into time evolution information of the same scale as the spatial information matrix of the human motion posture sequence; Specifically, three parallel one-dimensional convolutional layers with different receptive fields are used to perform convolution operations on the spatial information matrix corresponding to the human motion posture sequence to obtain multi-scale time evolution information, and the time evolution information of different scales is mapped into time evolution information of the same scale as the spatial information matrix of the human motion posture sequence. The specific operations are as follows: ; ; ; As mentioned above, represents a T-frame sequence of human motion postures that fuses the spatial information matrix of each frame; and is the mapping function, is the convolution operation of the one-dimensional convolution layer, represents the time evolution information corresponding to the one-dimensional convolutional layer of the first receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer of the second receptive field, Represents the time evolution information corresponding to the one-dimensional convolutional layer of the third receptive field.
[0032] It should be noted that before the convolution operation, the spatial information matrix corresponding to the T-frame human motion posture sequence is first transposed into a two-dimensional matrix, so that each row represents the eigenvalue of each joint from the first frame to the T-th frame on the same coordinate axis; then, three one-dimensional convolution layers with different receptive fields are used for convolution operation. Due to the different sizes of the one-dimensional convolution layers, multi-scale time evolution information will be obtained after the convolution operation; then, the multi-scale two-dimensional matrix is mapped back to a two-dimensional matrix of uniform scale through a linear layer, and finally it is transposed into the time evolution information of the same scale as the spatial information matrix corresponding to the original human motion posture sequence.
[0033] Since the value of T is 10 frames, the above formula is converted to: ; ; ; in, It represents the spatial information matrix corresponding to the 10-frame human motion posture sequence fused from each frame to the tenth frame, , and Both represent the acquired time evolution information.
[0034] For example: three one-dimensional convolutional layers with different receptive fields use , and Before performing the convolution operation, first Transpose to 10 75, and then transpose it to 75 10 two-dimensional matrix, transposed Each row represents the eigenvalues corresponding to the same coordinate axis of each joint from the first frame to the tenth frame; since there are 25 joint points in total, and the coordinate information of each joint point includes the X-axis, Y-axis and Z-axis, the transposed There are 75 lines in total. Since three one-dimensional convolutional layers with different receptive fields are used for convolution operation, the convolution result will contain time evolution information of different scales, corresponding to 75 8, 75 6 and 75 4 two-dimensional matrix; therefore, we can use the linear layer to transform the time evolution information of different scales into 75 10; finally, we will 75 The two-dimensional matrix of 10 is mapped back to the time evolution information of the same scale as the spatial information matrix of the human motion posture sequence, that is, it is mapped into 10 25 3-dimensional matrix.
[0035] Step S105, adaptively fusing the time evolution information of uniform scale obtained after mapping; In order to better capture the time features that are closely related to prediction and better cooperate with the accurate extraction of relevant features, the adaptive fusion module is used to adaptively fuse the time evolution information of uniform scale obtained after mapping. The specific operations are as follows:
[0036] in, represents a T-frame motion sequence integrating multiple time evolution information. In this embodiment, It is the learnable weight corresponding to various time evolution information.
[0037] Step S106: using a multi-head attention module to encode the adaptively fused time evolution information to obtain a human motion posture sequence containing time-space related information; In order to more effectively obtain the temporal information from a single human posture motion sequence, five parallel self-attention heads are used to model the temporal information in different subspaces and perform addition operations to obtain a variety of complementary temporal features and human sequence features of motion context. The specific operations are as follows:
[0038] in, It is the encoded sequence of human motion postures containing rich spatiotemporal information. represents the s-th self-attention operation.
[0039] Step S107, predicting and generating a future human body motion posture sequence based on the acquired time-space related information; The human motion posture sequence obtained after processing and containing rich spatiotemporal information is input into the gated recurrent unit to generate future predictions. The specific operations are as follows:
[0040] in, Represents the predicted future t-frame human motion posture sequence.
[0041] Since the present invention predicts and generates a human motion posture sequence of t frames in the future, the number of frames of the output human motion posture sequence is not necessarily equal to the number of frames of the input human motion posture sequence. Therefore, in this embodiment, the GRU network adopts an N vsM network structure, the encoding module adopts a number of gated recurrent units corresponding to the number of frames of the input human motion posture sequence, the decoding module adopts a number of gated recurrent units corresponding to the number of frames of the output human motion posture sequence, the encoding module outputs a context vector C, and the decoding module decodes the context vector C to predict and generate a human motion posture sequence of t frames in the future.
[0042] For example, a 10-frame sequence of human motion postures containing rich spatiotemporal information is input into the GRU network to predict and generate a 25-frame sequence of human motion postures in the future. Specifically, the first frame of human motion postures containing spatiotemporal information is flattened to 75 1 is then input into the first gated recurrent unit, and the second frame of human motion posture containing temporal and spatial related information is flattened to 75 1 is then input into the second gated recurrent unit, and so on. The human motion posture containing temporal and spatial related information in the tenth frame is flattened to 75 1 is input into the tenth gated recurrent unit, and the gated recurrent unit corresponding to the previous frame outputs the hidden vector to the gated recurrent unit corresponding to the next frame; finally, the gated recurrent unit corresponding to the tenth frame outputs the context vector C to the decoding module, and the decoding module decodes the context vector C, so that the gated recurrent unit of the decoding module sequentially outputs the human body motion postures of the 11th to 35th frames, thereby generating a human body motion posture sequence of the next 25 frames.
[0043] In summary, the present invention provides a method for predicting human motion based on multi-level spatiotemporal information. The present invention constructs a binary joint group with a parent-child relationship, models the spatial structure of the binary joint group, and then evaluates the correlation between joints and binary joint groups to capture complementary motion information from multiple spatial structure layers. Secondly, local-global time evolution information is obtained by encoding temporal context features of different scales. Thirdly, by learning temporal information in multiple subspaces, rich temporal context is effectively obtained; finally, a gated recurrent unit is used to generate predictions for the future. The present invention effectively enhances the richness of the acquired spatiotemporal related information by capturing spatially related information in multiple spatial structure levels such as joints and binary joint groups, and constructing temporal evolution information at multiple scales, thereby improving the prediction accuracy. A two-stage multi-level spatiotemporal representation learning is used to perform human motion prediction tasks, which includes two core stages: 1) Modeling the hierarchical structure of multiple human joints to assist the model in understanding the spatial coordination of human motion and capturing complementary motion information from structures with spatial correlation; 2) Modeling the human evolution characteristics from multiple time scales to obtain rich temporal context information, and capturing temporal evolution relationships to generate future predictions. A large number of experiments show that the present invention has shown high performance on two large-scale human motion benchmark datasets, Human 3.6M and CMU Mocap.
[0044] The present invention uses the mean per joint position error (MPJPE) as an evaluation index to test the performance of different algorithms.
[0045] By testing the model performance on the Human3.6M dataset S5 file, the prediction accuracy of each action is summarized in Table 1. Compared with the existing human motion prediction algorithms LSTM3LR, Res-GRU, HP-GAN, Bi-GAN, ConSeq2Seq, HMR, LTD, DM-GNN, AVG, Traj-Net and MMA, it is found that the prediction accuracy of the present invention is significantly better than the above algorithms in terms of 15 different actions and average accuracy. The present invention shows higher performance in highly random actions. For example, the MPJPE of the present invention is 6.6mm at 80ms for the "posing" action, and an MPJPE of 9.5mm is achieved at 160ms for the "smoking" action. In addition, the average prediction error of the present invention at 1000ms is 89.7mm, which shows that the present invention can also achieve higher long-term prediction accuracy.
[0046] Table 1 Quantitative comparison results of Human3.6M dataset
[0047] In order to further evaluate the performance of the algorithm, CMU Mocap was used to test the model performance. A total of 6 methods were evaluated, including LTD, Res-GRU, DMGNN, STSGCN, LPJP and TIM. From the experimental results in Table 2, the average prediction accuracy of the present invention exceeds the most advanced methods. It is worth mentioning that for those challenging actions that are difficult to predict (for example, playing basketball alone, walking, and wiping windows), the present invention also achieves significant performance improvement.
[0048] Table 2 Quantitative comparison structure of CMU Mocap dataset
[0049] Example 2 The present invention also provides an electronic device, comprising: a memory and a processor; The memory is used to store programs; The processor is used to call the program stored in the memory to execute the human motion prediction method based on multi-level spatiotemporal information as described in Example 1.
[0050] Example 3 The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a human body motion prediction method based on multi-level spatiotemporal information as described in Example 1 is implemented.
[0051] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation modes. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the embodiments shown and described herein.
Claims
1. A human motion prediction method based on multi-level spatiotemporal information, characterized in that: The specific steps include: Step S101, constructing a binary joint group with a parent-child relationship between joints in each frame; Step S102, aggregating the relevant features of the joints and the joints and their binary joint groups in each frame, obtaining the spatial information matrix of each joint in each frame, so as to capture complementary motion information from multiple spatial structure layers; Step S103, splicing the spatial information matrices of each joint in each frame, and fusing the spliced spatial information matrices of each frame together to obtain the spatial information matrix corresponding to the human body motion posture sequence; Step S104, using parallel convolutional layers with different receptive fields to perform convolution operations on the spatial information matrix corresponding to the human body motion posture sequence to obtain time evolution information of different scales; And the time evolution information of different scales is mapped into the time evolution information of the same scale as the spatial information matrix of the human motion posture sequence; Step S105, adaptively fusing the time evolution information of uniform scale obtained after mapping; Step S106: using a multi-head attention module to encode the adaptively fused time evolution information to obtain a human motion posture sequence containing time-space related information; Step S107: predict and generate a future human body motion posture sequence based on the acquired spatiotemporal related information.
2. The method for predicting human motion based on multi-level spatiotemporal information according to claim 1, characterized in that: The specific operation formula of step S101 is as follows: , In the formula, represents the binary joint group consisting of joint j and J joints of the human body posture in the Tth frame; , and The elements in are the coordinate values of joint 1, joint j and joint J on the X-axis, Y-axis and Z-axis respectively; represents the information correlation between the fusion feature and joint j after joint j and joint 1 form a binary joint group in the Tth frame, and is a learnable weight, is also a learnable weight, Represents the multiplication of corresponding matrix elements.
3. The method for predicting human motion based on multi-level spatiotemporal information according to claim 2, characterized in that: The operation formula of step S102 is as follows: Fishy + + ), in, The binary joint group used to evaluate the relevance of joint j to joint j, Used to evaluate the correlation between joints, represents the spatial information matrix of joint j in the Tth frame obtained by aggregating the relevant features of joint j in the Tth frame and J joints in the human body posture and joint j and its binary joint group, is the activation function, and are all learnable weights, Represents matrix product.
4. The method for predicting human motion based on multi-level spatiotemporal information according to claim 3, characterized in that: The specific operation formula of step S104 is as follows: ; ; ; represents the spatial information matrix corresponding to the T-frame human motion posture sequence; and is the mapping function, is the convolution operation of the one-dimensional convolution layer of the first receptive field, is the convolution operation of the one-dimensional convolution layer of the second receptive field, is the convolution operation of the one-dimensional convolution layer of the third receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer of the first receptive field, represents the time evolution information corresponding to the one-dimensional convolutional layer of the second receptive field, Represents the time evolution information corresponding to the one-dimensional convolutional layer of the third receptive field.
5. The method for predicting human motion based on multi-level spatiotemporal information according to claim 4, characterized in that: The step S104 specifically includes: Transpose the spatial information matrix corresponding to the T-frame human motion posture sequence into a two-dimensional matrix so that each row represents the eigenvalue of each joint on the same coordinate axis from the first frame to the T-th frame; Three one-dimensional convolutional layers with different receptive fields are used for convolution operations to obtain multi-scale time evolution information; The multi-scale two-dimensional matrix is mapped back to a two-dimensional matrix of uniform scale through a linear layer and transposed into the time evolution information of uniform scale of the spatial information matrix corresponding to the human motion posture sequence.
6. A method for predicting human motion based on multi-level spatiotemporal information according to claim 4 or 5, characterized in that: The specific operation formula of step S105 is as follows: in, represents a T-frame motion posture sequence that integrates multiple time evolution information. is the learnable weight corresponding to the time evolution information.
7. The method for predicting human motion based on multi-level spatiotemporal information according to claim 6, characterized in that: The specific operation formula of step S106 is as follows: in, It is the encoded sequence of human motion postures containing rich spatiotemporal information. represents the s-th self-attention operation.
8. The method for predicting human motion based on multi-level spatiotemporal information according to claim 7, characterized in that: The specific operation formula of step S107 is as follows: in, Represents the predicted future t-frame human motion posture sequence.
9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to call the program stored in the memory to execute the human motion prediction method based on multi-level spatiotemporal information as described in any one of claims 1-8.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a human motion prediction method based on multi-level spatiotemporal information as described in any one of claims 1-8.
Citation Information
Patent Citations
Pedestrian detection method based on global-local features
CN118314606A
Three-dimensional human motion prediction method and system based on improved graph convolutional network
CN118470115A
Human motion sequence prediction method and system based on skeleton enhanced Transform
CN119540279A
Action prediction method for unknown category
WO2023142552A1