Three-dimensional human body posture prediction method and device, medium and product
The three-dimensional human posture prediction model constructed through the LSTM network and the generative adversarial network solves the nonlinear characteristics and high-dimensional space-time dependence relationship problems of human motion prediction in the prior art, realizes high-precision human posture prediction, and improves the safety and efficiency of the human-computer collaboration system.
Patent Information
- Application Number
- CN202510363648.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively model the nonlinear characteristics and high-dimensional space-time dependence of human body movement in three-dimensional human posture prediction, resulting in a lack of coherence and adaptability in the prediction results, especially when handling actions at different speeds, showing poor adaptability, and neglecting the structural constraints of the human body skeleton, which cannot accurately reflect the spatial relationship between joints.
The LSTM network and the generative adversarial network are used to construct a three-dimensional human posture prediction model. Through the mixed timing interaction module, the timing enhancement maximum information pooling module and the skeleton graph network module, combined with multi-layer perceptron and graph neural network, a human skeleton structure diagram is constructed to enhance the model's adaptability and prediction coherence to different speed movements, reflecting the coordination of the overall skeleton structure.
High-precision prediction of human motion trajectory is achieved, and the safety and efficiency of human-computer collaboration system is improved, especially in complex action scenarios, with higher prediction accuracy and coherence.
Smart Images

Figure CN120299082A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of pose prediction, and in particular, to a three-dimensional human pose prediction method, device, medium, and product. Background Art
[0002] In modern human-robot collaboration systems, the accuracy of three-dimensional human trajectory prediction is crucial for ensuring safety and improving efficiency. With the continuous progress of robotics technology, how to effectively capture and predict the complexity of human motion has become a research hotspot. However, existing technologies still have limitations in dealing with the following problems: In terms of traditional methods, mainly referring to the basic methods for human motion prediction in the early stage, including physical deterministic models such as the social force model that constructs a mathematical description of motion behavior through mechanical principles, statistical models such as the extended Kalman filter and particle filter based on a probabilistic inference framework, and early simple sequence-to-sequence neural network models. Due to their inherent linear assumptions and limited feature representation capabilities, these traditional methods are difficult to effectively model the non-linear features and high-dimensional spatio-temporal dependence relationships in human motion, lack adaptability in dealing with actions at different speeds, and are difficult to effectively integrate multi-step and single-step time information, resulting in a lack of coherence and naturalness in the prediction results, and thus showing poor adaptability in dealing with actions at different speeds.
[0003] For existing human motion prediction models, they mainly include the Convseq2seq model based on convolutional sequence-to-sequence, the Learning Trajectory Dependencies (LTD) series of models using different historical frame lengths (such as LTD-10-25 and LTD-50-10), the Discrete Cosine Transform Recurrent Neural Network Graph Convolutional Network (DCT-RNN-GCN) that combines discrete cosine transform and graph convolutional network, the Multi-Scale Residual Graph Convolutional Networks (MSR-GCN) that adopts a multi-scale residual structure, as well as the Spatio-Temporal Dynamic Graph Convolutional Network (ST-DGCN) and the Spatial-Temporal Graph Convolutional Network (ST-GCN) based on spatio-temporal graph structure, etc. The above models all face challenges such as decreased accuracy and low feature extraction efficiency in long sequence prediction and complex action scenarios, and often ignore the structural constraints of the human skeleton, unable to accurately reflect the spatial relationship between joints, which affects the ability to capture complex action details.
[0004] Based on the above problems, there is an urgent need to provide a new three-dimensional human pose prediction method to achieve high-precision prediction of human motion trajectories, thereby improving the safety and efficiency of human-machine collaboration systems. Summary of the Invention
[0005] The purpose of this application is to provide a three-dimensional human pose prediction method, device, medium, and product, which can achieve high-precision prediction of human motion trajectories, thereby improving the safety and efficiency of human-machine collaboration systems.
[0006] To achieve the above purpose, this application provides the following solutions:
[0007] In the first aspect, this application provides a three-dimensional human pose prediction method, and the three-dimensional human pose prediction method includes:
[0008] Obtain three-dimensional human bone data to be predicted; the three-dimensional human bone data includes: the three-dimensional spatial coordinates of the bones;
[0009] Construct a three-dimensional human pose prediction model based on the LSTM network and the generative adversarial network; the three-dimensional human pose prediction model includes: a generator and a discriminator; the generator includes: an LSTM encoder, a hybrid temporal interaction module, a temporal enhanced maximum information pooling module, a bone graph network module, and an LSTM decoder; the LSTM encoder is used to perform feature encoding on the input three-dimensional human bone data; the hybrid temporal interaction module is used to fuse the multi-head attention mechanism and the Fourier attention mechanism, and perform spatio-temporal feature extraction on the encoded three-dimensional human bone data; the temporal enhanced maximum information pooling module is used to preliminarily aggregate spatio-temporal features using a multi-layer perceptron; and perform screening on the preliminarily aggregated spatio-temporal features using an improved max-pooling operation; then use the LSTM network layer to perform temporal enhancement processing on the pooled features; the bone graph network module is used to construct a human bone structure diagram in the form of a graph neural network according to the three-dimensional human bone data; the LSTM decoder is used to fuse a random noise vector with the human bone structure diagram and the temporally enhanced features to generate a prediction sequence; the discriminator is used to discriminate the prediction sequence;
[0010] According to the three-dimensional human bone data to be predicted, use the trained three-dimensional human pose prediction model to perform three-dimensional human pose prediction.
[0011] Optionally, the obtaining of the three-dimensional human bone data to be predicted specifically includes:
[0012] Use three ZED stereo cameras arranged in a spatial hexagonal layout to obtain the images to be predicted; the ZED stereo cameras are fixedly installed at a set interval;
[0013] According to the predicted images, use the SDK software built into the ZED stereo camera to determine the three-dimensional human bone data.
[0014] Optionally, after obtaining the three-dimensional human bone data to be predicted, it further includes:
[0015] Perform dimensionality reduction processing on the three-dimensional human bone data to obtain the three-dimensional human bone data of key points; the key points include: the tip of the nose point, the cervical vertebra point, the left shoulder joint point, the right shoulder joint point, the left elbow joint point, the right elbow joint point, the left wrist joint point, and the right wrist joint point.
[0016] Optionally, the LSTM encoder adopts a 128-dimensional hidden layer and a 32-dimensional embedding layer structure, and performs feature encoding through a deep learning network.
[0017] Optionally, the subsequent use of the LSTM network layer to perform temporal enhancement processing on the pooled features specifically includes:
[0018] Use the formula Perform temporal enhancement processing on the pooled features;
[0019] where H pred is the feature after temporal enhancement processing; and are the pooled features at time t - 1 and time t respectively, and W p is the weight parameter of the LSTM network layer.
[0020] Optionally, the loss function of the skeleton graph network module is:
[0021] L gn = λ1L bone + λ2L motion ;
[0022] where L gn is the loss function of the skeleton graph network module, L bone is the constraint loss of the skeleton length, L motion is the constraint loss of the motion smoothness, λ1 is the weight coefficient for balancing the constraint loss of the skeleton length, and λ2 is the weight coefficient for balancing the constraint loss of the motion smoothness.
[0023] Optionally, the loss function of the three - dimensional human pose prediction model is:
[0024] L total = L2 + λ1L bone + λ2L motion ;
[0025] where L total is the loss function of the three - dimensional human pose prediction model, L2 is the reconstruction loss, X i is the true skeleton coordinate at the i - th time step, is the predicted skeleton coordinate at the i - th time step, T obs is the length of the observed three - dimensional human pose data sequence, t pred is the length of the predicted three - dimensional human pose data sequence, L bone is the skeleton constraint loss, d ij is the distance between the i - th and j - th joint points in the predicted skeleton, is the standard skeleton length, E represents the set of edges connecting the skeletons, L motion is the motion smoothness loss, is the feature representation of node vi at the k - th iteration, is the feature representation of node vi at the (k - 1) - th iteration, V is the set of skeleton nodes, λ1, λ2 are balance factors, || ||2 is the 2 - norm, || || 2 is the square of the norm, is the square of the 2-norm.
[0026] In a second aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the three-dimensional human body posture prediction method.
[0027] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the three-dimensional human body posture prediction method when executed by a processor.
[0028] In a fourth aspect, the present application provides a computer program product, including a computer program, which implements the three-dimensional human posture prediction method when executed by a processor.
[0029] According to the specific embodiments provided in this application, this application has the following technical effects:
[0030] The present application provides a three-dimensional human posture prediction method, device, medium and product, which constructs a three-dimensional human posture prediction model based on LSTM (long short-term memory) network and generative adversarial network (GAN), and better predicts the human posture after a longer time step through LSTM network and temporal enhanced maximum information pooling (TEMIPM) module; the introduction of hybrid temporal interaction (HTSI) module in the generator and the generative adversarial of GAN network effectively enhance the adaptability of the model to actions of different speeds, and improve the coherence and naturalness of the prediction; GAN is used to make the prediction results more accurately reflect the coordination of the overall skeleton structure; the present application can achieve high-precision prediction of human motion trajectory, thereby improving the safety and efficiency of human-machine collaboration system; and provides reliable technical support for human-machine collaboration in practical applications, which can be widely used in human-machine collaboration systems in high-risk environments such as industrial production, construction sites and medical care. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 This is a schematic diagram of a flow chart of a three-dimensional human body posture prediction method in an embodiment of the present application;
[0033] Figure 2 Schematic diagram of the structure of a three-dimensional human posture prediction model in one embodiment of the present application. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0035] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0036] In an exemplary embodiment, as Figure 1 shown, a three-dimensional human pose prediction method is provided. The method includes the following S101 to S102. Among them:
[0037] S101, obtaining three-dimensional human bone data to be predicted; the three-dimensional human bone data includes: three-dimensional spatial coordinates (x, y, z) of the bones;
[0038] S101 specifically includes:
[0039] S11, using three ZED stereo cameras arranged in a spatial hexagonal pattern to obtain the images to be predicted; the ZED stereo cameras are fixedly installed at a set interval; wherein, the set interval is 3 meters;
[0040] S12, according to the predicted images, using the SDK software built in the ZED stereo cameras to determine the three-dimensional human bone data. Software Development Kit (SDK)
[0041] In an exemplary embodiment, when constructing the training set of the three-dimensional human pose prediction model, the ZED stereo cameras are used to record each action category (walking, walking together, discussing, pointing, object operation, product assembly, and tool use) for no less than 120 seconds, and the sampling frequency is 25Hz;
[0042] After S101, it further includes:
[0043] Performing dimensionality reduction processing on the three-dimensional human bone data to obtain the three-dimensional human bone data of key points; the key points include: the tip of the nose point, the cervical vertebra point, the left shoulder joint point, the right shoulder joint point, the left elbow joint point, the right elbow joint point, the left wrist joint point, and the right wrist joint point.
[0044] S102, constructing a three-dimensional human pose prediction model based on the LSTM network and the generative adversarial network;
[0045] AsFigure 2 As shown in Figure 2 , the three-dimensional human pose prediction model includes: a generator and a discriminator; the generator is responsible for predicting the human motion sequence, and the discriminator is responsible for evaluating the authenticity of the generated sequence. The generator consists of an encoder; the discriminator adopts an encoder-classifier structure.
[0046] The generator includes: an LSTM encoder, a hybrid temporal-spatial interaction module, a temporal enhancement maximum information pooling module, a skeletal graph network module, and an LSTM decoder;
[0047] The LSTM encoder is used to adopt a 128-dimensional hidden layer and a 32-dimensional embedding layer structure to perform feature encoding through a deep learning network; the LSTM encoder not only ensures a full understanding of the input sequence but also controls the computational complexity. To achieve effective encoding of the input sequence, the LSTM encoder represents the encoding process in the following mathematical form:
[0048]
[0049] where, represents the output state of the encoder hidden layer of the i-th skeletal point at time t; represents the hidden layer state at time t-1; represents the three-dimensional coordinate input [x, y, z] of the i-th skeletal point at time t; W en represents the weight parameter of the encoder.
[0050] The hybrid temporal-spatial interaction module is used to fuse the multi-head attention mechanism and the Fourier attention mechanism, and perform spatio-temporal feature extraction on the encoded three-dimensional human skeletal data; in the spatial dimension, the HTSI module configures 4 attention heads to calculate the spatial correlation weights between skeletal points from different angles respectively, and fuse the multi-head features to obtain a rich spatial representation. In the temporal dimension, the Fourier transform is introduced to analyze the periodic features of the action, effectively capturing the rhythmic patterns of human motion. Finally, through the feature fusion layer, the spatial and temporal information is organically combined to construct a complete spatio-temporal feature representation. In the multi-head attention mechanism, the core attention calculation process can be expressed by the following mathematical formula:
[0051]
[0052] where, A k represents the output of the k-th attention head; q t , k t , v t represent the Query, Key, and Value vectors respectively; d k represents the scaling factor of the attention mechanism, and the softmax function is used to normalize the attention scores into a probability distribution.
[0053] The temporal enhancement maximum information pooling module is used to initially aggregate spatio-temporal features using a multi-layer perceptron; and perform screening on the initially aggregated spatio-temporal features using an improved max-pooling operation; then perform temporal enhancement processing on the pooled features using an LSTM network layer, which improves the model's ability to capture long-term dependencies.
[0054] Traditional max-pooling decomposes the human skeleton space into an i×j grid structure, projects joint features onto the corresponding grid cells based on three-dimensional coordinates, and aggregates the features in these grids using simple operations such as summation or averaging. This method mainly focuses on the extraction of spatial features, but largely ignores the information in the time dimension, resulting in the weakening or even complete loss of the characteristics of joint dynamic motion during the feature aggregation process, and being unable to effectively capture the temporal changes in human motion. The improved max-pooling operation constructs an n×n local receptive field for each anatomical key point, focuses on the feature extraction in the area around each key point, and dynamically evaluates the features by quantifying the temporal importance of the features within the grid column vector. It retains the most representative joint motion features through a global sorting optimization algorithm and combines an LSTM layer to further enhance the learning ability of temporal features. The improved max-pooling operation can simultaneously maintain the integrity of the skeleton's spatial structure and capture the temporal dynamic characteristics of joint motion, providing a more comprehensive and accurate feature representation for human motion prediction.
[0055] To enhance the expressive ability of temporal features, this module processes the pooled features through an LSTM network, and its mathematical expression is:
[0056]
[0057] where, H pred is the feature after temporal enhancement processing; and are the pooled features at time t-1 and time t respectively, and W p is the weight parameter of the LSTM network layer.
[0058] The Skeleton Graph Network (SKGN) module is used to construct a human skeleton structure graph in the form of a graph neural network based on three-dimensional human skeleton data; in the design of node features, the SKGN module comprehensively considers position information (three-dimensional coordinates of skeleton points), temporal information (time step information), and motion features (description of motion state). In the design of edge features, it includes bone length and azimuth angle information, which is used to describe the geometric relationship between bones. Through the message passing mechanism of the graph network, the dynamic update of node and edge features is realized, which not only maintains the constraint of the skeleton structure but also realizes the effective transmission of motion information. To simultaneously ensure the rationality of the skeleton structure and the continuity of motion, the loss function of the SKGN module is:
[0059] Lgn = λ1L bone + λ2L motion ;
[0060] where L gn is the loss function of the bone graph network module, L bone is the constraint loss of the bone length, L motion is the constraint loss of the motion smoothness, λ1 is the weight coefficient for balancing the constraint loss of the bone length, and λ2 is the weight coefficient for balancing the constraint loss of the motion smoothness.
[0061] The LSTM decoder, as the last layer of the generator, is responsible for generating the final prediction sequence; the LSTM decoder is used to fuse the random noise vector with the human bone structure diagram and the features after temporal enhancement processing to generate a natural and continuous human motion sequence through a multi-layer network structure. The introduction of noise increases the diversity of the generation results and improves the robustness of the prediction. Finally, the decoder generates the prediction sequence through the following mapping function:
[0062]
[0063] where represents the prediction output at the i-th time step; represents the hidden state of the decoder at time t; μ(·) represents a multi-layer perceptron network, which is used to generate the final prediction coordinates;
[0064] The discriminator adopts a two-layer structure. The first layer is an LSTM encoder, which is responsible for extracting the temporal features of the input sequence; the second layer is a multi-layer perceptron classifier, which maps the extracted features to a authenticity score between 0 and 1. This design not only ensures the recognition ability of the discriminator but also avoids the training difficulties caused by overly complex structures.
[0065] The three-dimensional human pose prediction model trains the optimization strategy and the loss function respectively;
[0066] Among them, during network training, the Adam optimizer is used to update the parameters of the generator and the discriminator, and the learning rate is set to 0.0005. Considering the computational efficiency and training stability, the batch size is set to 64, and the total number of training epochs is 2000. In each training cycle, an asymmetric update strategy is adopted: the parameters of the generator are updated once for every two updates of the parameters of the discriminator to maintain the balance of the generative adversarial training. The training process can be expressed as:
[0067]
[0068] where represents the hidden state of the discriminator at time t; T real / fake represents the real or generated trajectory sequence; WD Represent the network parameters of the discriminator.
[0069] To achieve high-precision human motion prediction, use L total = L2 + λ1L bone + λ2L motion Determine the loss function of the 3D human pose prediction model;
[0070] Among them, L2 is the reconstruction loss, which is used to measure the difference between the predicted sequence and the real sequence. X i Is the real bone coordinate at the i-th time step. Is the predicted bone coordinate at the i-th time step, T obs Is the length of the observed 3D human pose data sequence, t pred Is the length of the predicted 3D human pose data sequence, L bone Is the bone constraint loss, which is used to ensure the stability of the bone structure. d ij Is the distance between the i-th and j-th joint points in the predicted bone. Is the standard bone length, E represents the set of edges connecting bones, L motion Is the motion smoothness loss, which is used to ensure the continuity of the motion sequence. Is the feature representation of node vi at the k-th iteration. Is the feature representation of node vi at the (k - 1)-th iteration, V is the set of bone nodes, λ1 and λ2 are balance factors, || ||2 is the 2-norm, || || 2 The square of the norm. Is the square of the 2-norm.
[0071] S103. According to the 3D human bone data to be predicted, use the trained 3D human pose prediction model to perform 3D human pose prediction.
[0072] Specifically, use the trained 3D human pose prediction model for the historical 10-frame bone sequence data to obtain the next 25-frame prediction sequence. Each frame contains the 3D coordinates of 8 key points. The prediction time span is 1 second (sampling at 25Hz).
[0073] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a three-dimensional human body pose prediction method.
[0074] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0075] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0076] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0077] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0078] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0079] In the present application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.
[0080] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0081] In this text, specific examples are used to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A three-dimensional human body pose prediction method, characterized in that, The three-dimensional human body pose prediction method includes: Obtaining three-dimensional human body bone data to be predicted; the three-dimensional human body bone data includes: the three-dimensional spatial coordinates of the bones; Constructing a three-dimensional human body pose prediction model based on the LSTM network and the generative adversarial network; the three-dimensional human body pose prediction model includes: a generator and a discriminator; the generator includes: an LSTM encoder, a hybrid temporal interaction module, a temporal enhancement maximum information pooling module, a bone graph network module, and an LSTM decoder; the LSTM encoder is used to perform feature encoding on the input three-dimensional human body bone data; the hybrid temporal interaction module is used to fuse the multi-head attention mechanism and the Fourier attention mechanism, and extract spatio-temporal features from the encoded three-dimensional human body bone data; the temporal enhancement maximum information pooling module is used to preliminarily aggregate spatio-temporal features using a multi-layer perceptron; and perform screening on the preliminarily aggregated spatio-temporal features using an improved max-pooling operation; then perform temporal enhancement processing on the pooled features using an LSTM network layer; the bone graph network module is used to construct a human body bone structure diagram in the form of a graph neural network according to the three-dimensional human body bone data; the LSTM decoder is used to fuse a random noise vector with the human body bone structure diagram and the temporally enhanced features to generate a prediction sequence; the discriminator is used to discriminate the prediction sequence; According to the three-dimensional human body bone data to be predicted, use the trained three-dimensional human body pose prediction model to predict the three-dimensional human body pose.
2. The three-dimensional human body pose prediction method according to claim 1, wherein, The obtaining of the three-dimensional human body bone data to be predicted specifically includes: Using three ZED stereo cameras arranged in a spatial hexagonal pattern to obtain the image to be predicted; the ZED stereo cameras are fixedly installed at a set interval; According to the predicted image, use the SDK software built into the ZED stereo camera to determine the three-dimensional human body bone data.
3. The three-dimensional human body posture prediction method according to claim 1, wherein, After obtaining the three-dimensional human body bone data to be predicted, it further includes: Performing dimensionality reduction processing on the three-dimensional human body bone data to obtain the three-dimensional human body bone data of key points; the key points include: the tip of the nose point, the cervical vertebra point, the left shoulder joint point, the right shoulder joint point, the left elbow joint point, the right elbow joint point, the left wrist joint point, and the right wrist joint point.
4. The three-dimensional human body posture prediction method according to claim 1, wherein The LSTM encoder adopts a 128-dimensional hidden layer and a 32-dimensional embedding layer structure, and performs feature encoding through a deep learning network.
5. The three-dimensional human body pose prediction method according to claim 1, wherein The subsequent use of the LSTM network layer to perform temporal enhancement processing on the pooled features specifically includes: Use the formula to perform temporal enhancement processing on the pooled features; Among them, H pred is the feature after temporal enhancement processing; P t-1 i and P t i are the pooled features at time t-1 and time t respectively, and W p is the weight parameter of the LSTM network layer.
6. The three-dimensional human body pose prediction method according to claim 1, wherein The loss function of the bone graph network module is: L gn = λ1L bone + λ2L motion ; Among them, L gn is the loss function of the bone graph network module, L bone is the constraint loss of the bone length, L motion is the constraint loss of the motion smoothness, λ1 is the weight coefficient for balancing the constraint loss of the bone length, and λ2 is the weight coefficient for balancing the constraint loss of the motion smoothness.
7. The three-dimensional human body posture prediction method according to claim 1, wherein The loss function of the three-dimensional human body pose prediction model is: L total = L2 + λ1L bone + λ2L motion ; Among them, L total is the loss function of the 3D human pose prediction model, L2 is the reconstruction loss, X i is the true bone coordinate at the i-th time step, is the predicted bone coordinate at the i-th time step, T obs is the length of the observed 3D human pose data sequence, t pred is the length of the predicted 3D human pose data sequence, L bone is the bone constraint loss, d ij is the distance between the i-th and j-th joint points in the predicted bone, is the standard bone length, E represents the set of edges connecting the bones, L motion is the motion smoothness loss, h vi (k) is the feature representation of node vi at the k-th iteration, h vi (k-1) is the feature representation of node vi at the (k - 1)-th iteration, V is the set of bone nodes, λ1, λ2 are balance factors, || ||2 is the 2-norm, || || 2 is the square of the norm, is the square of the 2-norm.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the three-dimensional human body pose prediction method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional human body pose prediction method according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional human body pose prediction method according to any one of claims 1-7.
Citation Information
Cited By
Joint angle prediction method, system and equipment for infant crawling and medium
CN120600226A
A method, system, device and medium for predicting joint angles of infant crawling
CN120600226B