Video three-dimensional human body posture estimation method and device based on discretization potential space
By constructing a synchronous network and 3D human pose reconstruction network, using the training data set to learn the correlation between 2D and 3D human poses, the problem that 3D pose estimation relies on 2D pose estimation network accuracy in the prior art is solved, and a more accurate 3D human pose estimation is achieved.
Patent Information
- Application Number
- CN202510043734.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-10
AI Technical Summary
In the prior art, 3D human pose estimation depends on the accuracy of the 2D human pose estimation network, resulting in the estimation error mostly from the error of the 2D pose estimation network rather than the 3D pose estimation process itself.
By obtaining the training data set, including 2D human pose sequences and 3D human pose sequences, a synchronous network, 3D human pose reconstruction network and 3D human pose estimation network are built, and the potential connections between 2D and 3D poses are learned using the synchronous network, and the 3D pose reconstruction network is trained to improve the accuracy of 3D pose estimation.
Through the synchronous network learning the correlation between 2D and 3D poses and injecting them into the 3D human pose estimation network, the 3D human pose estimation network can be more accurately estimated and the accuracy of the estimation can be improved.
Smart Images

Figure CN120047996A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of human pose estimation, and in particular, to a method and device for video three-dimensional human pose estimation based on a discretized latent space. Background Art
[0002] With the rapid development of computer vision and deep learning technologies, 3D human pose estimation, as one of the core tasks for understanding and analyzing human behaviors, has received extensive attention. It has broad application prospects in many fields such as intelligent fitness, virtual reality, motion capture, security monitoring, and human-computer interaction. The goal of 3D human pose estimation is to accurately capture the joint positions of a human body in three-dimensional space by analyzing the positions of human key points in images or videos through algorithms, providing a reliable basis for action evaluation and interaction.
[0003] In current technical solutions, a pre-trained 2D human pose estimation network is often used to extract 2D human poses from video images, and then a 3D human pose sequence is estimated based on the 2D human pose sequence, that is, the 2D-to-3D Lifting method. However, the above method is highly dependent on the accuracy of the 2D human pose estimation network, resulting in the estimation error possibly originating more from the error of the 2D human pose estimation network rather than from the process of estimating the 3D human pose from the 2D human pose.
[0004] Therefore, how to make full use of the connection between 2D human key points and 3D human key points to ensure the accuracy of the network's 3D human pose estimation results has become a technical problem to be solved urgently. Summary of the Invention
[0005] Embodiments of the present application provide a method and device for video three-dimensional human pose estimation based on a discretized latent space, which can, to at least a certain extent, make full use of the connection between 2D human key points and 3D human key points and ensure the accuracy of the network's 3D human pose estimation results.
[0006] Other features and advantages of the present application will become apparent through the following detailed description, or will be partially learned through the practice of the present application.
[0007] According to one aspect of the embodiments of the present application, a method for video three-dimensional human pose estimation based on a discretized latent space is provided, including:
[0008] Obtain a training data set, where the training data set contains a number of training data pairs, the number of training data pairs includes positive sample pairs and negative sample pairs, and each training data pair contains a 2D human pose sequence and a 3D human pose sequence;
[0009] Train a pre - constructed synchronization network according to the training data set, so that the synchronization network learns the potential relationship between the 2D human pose sequence and the 3D human pose sequence;
[0010] Construct a 3D human pose reconstruction network and train it based on the 3D human pose sequences in the training data set. The 3D human pose reconstruction network includes an encoder, a discretized latent space, and a decoder;
[0011] Construct a 3D human pose estimation network based on a sampler network, the discretized latent space, and the decoder in the trained 3D human pose reconstruction network. Train the sampler network in the 3D human pose estimation network according to the 2D human pose sequences in the training data set, and calculate an auxiliary loss through the synchronization network during training for weak supervision, so that the 3D human pose estimation network can estimate the corresponding 3D human pose based on the input 2D human pose sequence;
[0012] Perform three - dimensional human pose estimation according to the pre - trained 2D human pose estimation network and 3D human pose estimation network.
[0013] According to one aspect of the embodiments of the present application, a video three - dimensional human pose estimation device based on a discretized latent space is provided, including:
[0014] An acquisition module, configured to acquire a training data set, where the training data set contains a number of training data pairs, and the number of training data pairs includes positive sample pairs and negative sample pairs, and each training data pair contains a 2D human pose sequence and a 3D human pose sequence;
[0015] A first training module, configured to train a pre - constructed synchronization network according to the training data set, so that the synchronization network learns the potential relationship between the 2D human pose sequence and the 3D human pose sequence;
[0016] A second training module, configured to construct a 3D human pose reconstruction network and train it based on the 3D human pose sequences in the training data set. The 3D human pose reconstruction network includes an encoder, a discretized latent space, and a decoder;
[0017] A third training module, configured to construct a 3D human pose estimation network based on a sampler network, the discretized latent space, and the decoder in the trained 3D human pose reconstruction network. Train the sampler network in the 3D human pose estimation network according to the 2D human pose sequences in the training data set, and calculate an auxiliary loss through the synchronization network during training for weak supervision, so that the 3D human pose estimation network can estimate the corresponding 3D human pose based on the input 2D human pose sequence;
[0018] A processing module for performing three-dimensional human pose estimation based on a pre-trained 2D human pose estimation network and a 3D human pose estimation network.
[0019] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the method for video three-dimensional human pose estimation based on a discretized latent space as described in the above embodiments.
[0020] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for video three-dimensional human pose estimation based on a discretized latent space as described in the above embodiments.
[0021] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for video three-dimensional human pose estimation based on a discretized latent space provided in the above embodiments.
[0022] In the technical solutions provided in some embodiments of the present application, by obtaining a training data set, the training data set includes a number of training data pairs, the number of training data pairs includes positive sample pairs and negative sample pairs, and each training data pair includes a 2D human pose sequence and a 3D human pose sequence; training a pre-constructed synchronization network according to the training data set so that the synchronization network learns the latent connection between the 2D human pose sequence and the 3D human pose sequence; constructing a 3D human pose reconstruction network and training it based on the 3D human pose sequences in the training data set, the 3D human pose reconstruction network includes an encoder, a discretized latent space, and a decoder; constructing a 3D human pose estimation network based on a sampler network, the discretized latent space, and the decoder in the trained 3D human pose reconstruction network, training the sampler network in the 3D human pose estimation network according to the 2D human pose sequences in the training data set, and calculating an auxiliary loss through the synchronization network during the training process for weak supervision, so that the 3D human pose estimation network can estimate the corresponding 3D human pose based on the input 2D human pose sequence; performing three-dimensional human pose estimation based on a pre-trained 2D human pose estimation network and a 3D human pose estimation network.
[0023] In this way, by synchronously learning the correlation between the 2D human pose sequence and the 3D human pose sequence through the network, injecting it into the 3D human pose estimation network, and using the 3D human pose reconstruction network to perform reconstruction training on the 3D pose data to extract the prior knowledge of the data, the 3D human pose can be estimated more accurately, improving the accuracy of the estimation.
[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings
[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:
[0026] Figure 1 Shows a schematic diagram of the construction of positive and negative sample pairs according to an embodiment of this application;
[0027] Figure 2 Shows a schematic diagram of the structure of the synchronization network according to an embodiment of this application;
[0028] Figure 3 Shows a schematic diagram of the structure of the 3D human pose reconstruction network according to an embodiment of this application;
[0029] Figure 4 Shows a schematic diagram of the structure of the 3D human pose estimation network according to an embodiment of this application;
[0030] Figure 5 Shows a block diagram of a video three-dimensional human pose estimation device based on a discretized latent space according to an embodiment of this application;
[0031] Figure 6 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of this application. Detailed Embodiments
[0032] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0033] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0034] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0035] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0036] An embodiment of the present application provides a method for video three-dimensional human pose estimation based on a discretized latent space. This method can be applied to a terminal device or a server. Among them, the terminal device may include, but is not limited to, one or more of a smart phone, a tablet computer, a portable computer, and a desktop computer; the server may be a physical server or a cloud server.
[0037] In some embodiments of the present application, the method for video three-dimensional human pose estimation based on a discretized latent space includes the following steps:
[0038] S110. Obtain a training data set, where the training data set contains a number of training data pairs. The number of training data pairs includes positive sample pairs and negative sample pairs, and each training data pair contains a 2D human pose sequence and a 3D human pose sequence.
[0039] In one embodiment, obtaining the training data set includes:
[0040] Perform preprocessing on the video image sequence to obtain a 2D human pose sequence and a 3D human pose sequence;
[0041] Perform random dislocation according to the 2D human pose sequence and the 3D human pose sequence to construct a number of training data pairs, where the training data pairs include positive sample pairs and negative sample pairs.
[0042] In this embodiment, a video image sequence may include several video images. During preprocessing, for the video images in the video image sequence that do not contain people, they are screened and cropped to ensure that only the full body of one person is included in the video image. The image is scaled to a size of 384*288, and then the data is augmented, including 2D data augmentation and 3D data augmentation. Then, the 2D human key points are flipped. In one example, the number of human key points is 17, namely: hip, right hip, right knee, right foot, left hip, left knee, left foot, spine, chest, neck, head, left shoulder, left elbow, left wrist, right shoulder, right elbow, and right wrist.
[0043] Specifically, according to the camera parameters and the 3D data points (i.e., 3D human pose key points) in the world coordinate system, first, coordinate transformation is used to convert the 3D data points in the world coordinate system into 3D data points in the camera coordinate system, and the data is augmented. Then, according to the camera parameters, the 3D human pose key points are projected into 2D points, and then converted into the pixel coordinates of the image to obtain 2D human pose key points. According to the selected time window size, each segment of data is intercepted, and a video image sequence is divided into several video segments with the same number of frames. For the selected video segment demarcation points, the extra frames are filled, and the people in the filled frames are stationary. Finally, a reverse operation is performed to reverse the 2D data points and 3D data points once, so as to double the data again, obtaining a 2D human pose sequence and a 3D human pose sequence.
[0044] Next, according to the 2D human pose sequence and 3D human pose sequence obtained by preprocessing, it is necessary to further strengthen the connection between the two, and add this connection to the network to be trained later through a loss function. Specifically, the 2D data and 3D data with the same timestamp in the data can be regarded as positive sample pairs, and then the 3D data with a time offset and the 2D data with inconsistent timestamps are used to form negative sample pairs (as Figure 1 shown).
[0045] S120. Train a pre-constructed synchronization network according to the training data set, so that the synchronization network learns the potential connection between the 2D human pose sequence and the 3D human pose sequence.
[0046] In this embodiment, after the construction of the positive and negative sample pairs is completed, the pre-constructed synchronization network can be trained according to them. Specifically, as Figure 2As shown in the figure, in the synchronous network, a 6-layer Transformer Block is used to encode the 2D human posture sequence and the 3D human posture sequence respectively, and both are encoded into feature vectors of the same dimension (i.e., 2D posture features and 3D posture features). The similarity distance between the 2D posture features and the 3D posture features is measured by the contrast loss Contrastive Loss. The purpose is to shorten the similarity distance between the positive sample pairs and increase the similarity distance between the negative sample pairs. Then, the gradient descent algorithm is executed to optimize the synchronous network.
[0047] In one example, during the training of the synchronous network, the key parameter settings include: 100 training rounds, using the Adam optimizer, and a learning rate of 1e-4. The loss function of the synchronous network is defined as follows:
[0048]
[0049] Among them, y n Indicates whether it is a positive sample pair, 1 for a positive sample pair and 0 for a negative sample pair, N is the total amount of sample data, d is the cosine similarity distance between 2D posture features and 3D posture features, the goal is to make the distance between positive sample pair features as small as possible and the distance between negative sample pair features as large as possible. The definition of d is as follows:
[0050]
[0051] in, They represent the features of 2D human posture sequence and 3D human posture sequence after synchronous network encoding, and ∈ represents a very small number.
[0052] S130, constructing a 3D human posture reconstruction network, and training it based on the 3D human posture sequence in the training data set, the 3D human posture reconstruction network comprising an encoder, a discretized latent space, and a decoder.
[0053] In this embodiment, a person skilled in the art can pre-construct a 3D human posture reconstruction network, which is constructed based on a VQ-VAE network and includes an encoder, a discretized latent space, and a decoder. A VQ-VAE network is used to learn the distribution of 3D human postures in a training data set through network learning, compress network features into a discretized latent space, and use a multi-head self-attention structure to encode the 3D human posture sequence into multiple embedding vectors, and then select the embedding vector with the smallest similarity distance in the discretized latent space for replacement, and then input it into the decoder to complete decoding.
[0054] like Figure 3As shown, B is the batch size, T is the sequence length, and N is the number of key points. During the training process of the 3D human pose reconstruction network, the encoder encodes the key point coordinates in the 3D human pose sequence into 1024-dimensional feature vectors through a linear layer, uses a Temporal Convolutional Network (TCN) to analyze the temporal dimension of the feature vectors output by the linear layer, adds the temporal position information through positional encoding, and then converts them into embedding vectors through 6 Transformer Blocks;
[0055] The discretized latent space is designed as an [n, d] matrix, which contains n d-dimensional embedding vectors. Query the embedding vector with the highest similarity to the embedding vector output by the encoder in the discretized latent space to replace the embedding vector output by the encoder;
[0056] The decoder converts the replaced embedding vector into a 1024-dimensional feature, uses a temporal convolutional network to analyze the temporal dimension of the feature, adds the temporal position information through positional encoding, and then converts the feature into the original data dimension through 6 Transformer Blocks and a linear layer to complete the decoding. It should be noted that since the operation of querying and replacing features in the discretized latent space cannot perform gradient backpropagation, the gradient update of the encoder can be completed through the Stop Gradient technique. By comparing the difference between the input 3D human pose sequence and the 3D human pose sequence estimated by the network, the network parameters are optimized using the gradient descent method.
[0057] Specifically, during the training process, the key parameter settings include: setting the total number of embedding vectors n in the discretized latent space to 256, the feature dimension d to 64, the number of layers of Transformer Blocks in the encoder to 6, the number of network layers of Transformer Blocks in the decoder to 6, the number of iteration rounds to 80, using the AdamW optimizer, and the learning rate to 4e-5.
[0058] The total loss of the network consists of three parts: action loss, vector quantization loss, and embedding regression loss, which are defined as follows, where α and β are weight coefficients:
[0059] L = L motion + αL vq + βL reg
[0060] The action loss L motion in the total loss function is defined as follows:
[0061]
[0062] where, L w is the WMPJPE loss, Lt is the TC loss, L m is the MPJVE loss, are their respective weight coefficients.
[0063] The WMPJPE loss L w is defined as follows:
[0064]
[0065] where N is the number of human key points, W is the weight of each joint set, T is the number of frames in a sequence, p i,j and gt i,j are the predicted value and the ground truth of the 3D position corresponding to the i-th joint in the j-th frame, ||·|| 2 represents the L2 norm calculation, and the goal is to make the network predict the 3D human pose as accurately as possible.
[0066] The TC loss L t is defined as follows:
[0067]
[0068] where p i,j and p i,j-1 are the predicted values of the 3D positions corresponding to the i-th joint in the j-th frame and the (j - 1)-th frame, and the goal is to make the difference between each frame not too large, so that the result predicted by the network is smoother.
[0069] The MPJVE loss L m is defined as follows:
[0070]
[0071] The goal of this loss function is to focus on the movement speed of the joint points, ensure smoothness and continuity in the time dimension, and enable the model to better learn the dynamic changes of the action sequence.
[0072] The vector quantization loss L in the total loss function vq is defined as follows:
[0073] L vq = ||sg[z e (x)] - e k || 2
[0074] where z e (x) is the continuous feature vector generated by the encoder, e k is the discrete code vector stored in the discretized latent space, and sg[·] represents the stop gradient operation, indicating that z will not be updated during backpropagation e(x). The goal is to reduce the distance between the encoder output and the nearest code vector, while the code vector only updates itself.
[0075] The embedding regression loss L in the total loss function reg is defined as follows:
[0076] L reg = ||z e (x) - sg[e k || 2
[0077] where z e (x) is the continuous feature vector generated by the encoder, and e k is the discrete code vector stored in the discretized latent space. sg[·] represents the stop-gradient operation, indicating that e will not be updated during backpropagation. k The goal constrains the encoder embedding vector to prevent it from moving away from the nearest feature vector.
[0078] S140. Based on a sampler network, the discretized latent space in the trained 3D human pose reconstruction network, and the decoder, construct a 3D human pose estimation network. Train the sampler network in the 3D human pose estimation network according to the 2D human pose sequence in the training dataset, and calculate the auxiliary loss through the synchronization network during the training process for weak supervision, so that the 3D human pose estimation network can estimate the corresponding 3D human pose based on the input 2D human pose sequence.
[0079] In this embodiment, based on the trained 3D human pose reconstruction network, use its discretized latent space and decoder network, and train a sampler network to construct a 3D human pose estimation network. As Figure 4 shown, the sampler network encodes the 2D human pose sequence in the training dataset into a feature vector with the same dimension as the embedding vector in the discretized latent space, selects the vector with the smallest similarity distance in the discretized latent space for replacement, and after obtaining the replaced embedding vector, uses the decoder network for decoding to obtain the 3D human pose.
[0080] Specifically, during the training process of the 3D human pose estimation network, only train the sampler network, and freeze the parameters of the discretized latent space and the decoder of the 3D human pose reconstruction network. And use the frozen synchronization network after training to calculate the loss value as the auxiliary loss, and inject the correlation between the 2D human pose sequence and the 3D human pose sequence learned by the synchronization network into the 3D human pose estimation network, thereby enhancing the data connectivity.
[0081] Among them, the key parameter settings include: the iteration round is 100, the AdamW optimizer is used, and the learning rate is 4e-5. The total loss of the network consists of three parts, namely regression loss, action loss and synchronization loss, which are defined as follows, where α and β are weight coefficients:
[0082] L=L reg +γL motion +δL sync
[0083] Regression loss L reg is defined as follows:
[0084]
[0085] in, is the feature generated by the sampler network, e k is the discrete embedding vector stored in the discretized latent space, and sg[·] represents the stop gradient operation, indicating that e will not be updated during back propagation. k The goal is to close the distance between the sampler output and the nearest code vector while only updating the sampler network.
[0086] Action loss L motion The definition of is consistent with that in the 3D human pose VQ-VAE reconstruction network. is the weight coefficient, which will not be repeated here:
[0087]
[0088] Synchronization loss L sync is defined as follows:
[0089]
[0090] in, They respectively represent the features of the 2D human pose sequence and the 3D human pose data predicted and generated by the 3D human pose estimation network after being encoded by the synchronous network. Since the training process of the 3D human pose estimation network only includes positive sample targets, the second term in the loss function of training the synchronous network is discarded. The purpose is to inject the correlation prior knowledge of the synchronous network into the 3D human pose estimation network.
[0091] In the actual process of training the 3D human pose estimation network, a 2D human pose sequence is regarded as a matrix of [B, T, N, 2], where B is the batch size, T is the sequence length, N is the number of human key points, and the last dimension represents the 2D coordinates of the human key points in the image. Among them, the sampler network maps the data to high-dimensional features [B, T, N, 512] through a linear layer, and 512 represents the feature dimension. The Transformer Block is used to analyze the spatial dimension and time dimension of the data respectively. Specifically, when analyzing the spatial dimension, a 2D human pose sequence feature [B, T, N, 512] is regarded as multiple single-frame key points [B * T, N, 1024], aiming to learn the connection of body key points in each frame. When analyzing the time dimension, a 2D human pose sequence feature [B, T, N, 512] is regarded as the position change of multiple human key points in time [B * N, T, 512], aiming to learn the connection of key points between frames, as Figure 4 shown in the dashed box of
[0092] S150. Perform three-dimensional human pose estimation according to the pre-trained 2D human pose estimation network and 3D human pose estimation network.
[0093] In this embodiment, after the 3D human pose estimation network is trained, when three-dimensional human pose estimation is required, the input picture can be processed by the aforementioned preprocessing operation to be an image containing only one person, and then the trained 2D human pose estimation network is used to estimate it to obtain the 2D human pose sequence of this person. Then, the 2D human pose sequence is input into the 3D human pose estimation network so that it outputs the predicted 3D human pose estimation.
[0094] The following introduces the device embodiment of the present application, which can be used to execute the method for video three-dimensional human pose estimation based on a discretized latent space in the above embodiments of the present application. For details not disclosed in the device embodiment of the present application, please refer to the embodiment of the method for video three-dimensional human pose estimation based on a discretized latent space in the above of the present application.
[0095] Figure 5 The block diagram of a device for video three-dimensional human pose estimation based on a discretized latent space according to an embodiment of the present application is shown.
[0096] Referring to Figure 5 shown, a device for video three-dimensional human pose estimation based on a discretized latent space according to an embodiment of the present application includes:
[0097] An acquisition module for acquiring a training data set, where the training data set contains a number of training data pairs, the number of training data pairs including positive sample pairs and negative sample pairs, and each training data pair contains a 2D human body pose sequence and a 3D human body pose sequence;
[0098] A first training module for training a pre-constructed synchronization network according to the training data set, so that the synchronization network learns the potential relationship between the 2D human body pose sequence and the 3D human body pose sequence;
[0099] A second training module for constructing a 3D human body pose reconstruction network and training it based on the 3D human body pose sequence in the training data set, where the 3D human body pose reconstruction network includes an encoder, a discretized latent space, and a decoder;
[0100] A third training module for constructing a 3D human body pose estimation network based on a sampler network, the discretized latent space, and the decoder in the trained 3D human body pose reconstruction network, training the sampler network in the 3D human body pose estimation network according to the 2D human body pose sequence in the training data set, and calculating an auxiliary loss through the synchronization network during the training process for weak supervision, so that the 3D human body pose estimation network can estimate the corresponding 3D human body pose based on the input 2D human body pose sequence;
[0101] A processing module for performing three-dimensional human body pose estimation according to a pre-trained 2D human body pose estimation network and 3D human body pose estimation network.
[0102] Figure 6 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0103] It should be noted that Figure 6 The computer system of the shown electronic device is only an example and should not bring any restrictions to the functions and usage scopes of the embodiments of the present application.
[0104] Such as Figure 6As shown, the computer system includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 602 or a program loaded from a storage section 608 into a Random Access Memory (RAM) 603, such as executing the method described in the above embodiments. In the RAM 603, various programs and data required for system operations are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An Input / Output (I / O) interface 605 is also connected to the bus 604.
[0105] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.
[0106] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by a Central Processing Unit (CPU) 601, various functions defined in the system of the present application are executed.
[0107] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0109] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the unit itself.
[0110] As another aspect, this application also provides a computer-readable medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.
[0111] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0112] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of this application.
[0113] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of this application. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include common general knowledge or conventional technical means in the technical field not disclosed in this application.
[0114] It should be understood that this application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.
Claims
1. A method for estimating 3D human pose from video based on discretized latent space, characterized in that: include: Acquire a training data set, wherein the training data set includes a plurality of training data pairs, the plurality of training data pairs include positive sample pairs and negative sample pairs, and each of the training data pairs includes a 2D human body posture sequence and a 3D human body posture sequence; Training a pre-built synchronization network according to the training data set so that the synchronization network learns the potential connection between the 2D human pose sequence and the 3D human pose sequence; Constructing a 3D human posture reconstruction network and training it based on the 3D human posture sequence in the training data set, the 3D human posture reconstruction network comprising an encoder, a discretized latent space, and a decoder; A 3D human pose estimation network is constructed based on a sampler network and a discretized latent space and a decoder in a trained 3D human pose reconstruction network, the sampler network in the 3D human pose estimation network is trained according to a 2D human pose sequence in the training data set, and weak supervision is performed by calculating an auxiliary loss through the synchronization network during the training process, so that the 3D human pose estimation network can estimate the corresponding 3D human pose based on the input 2D human pose sequence; Three-dimensional human pose estimation is performed based on the pre-trained 2D human pose estimation network and 3D human pose estimation network.
2. The method according to claim 1, characterized in that During the training process of the 3D human posture reconstruction network, the encoder encodes the key point coordinates in the 3D human posture sequence into a 1024-dimensional feature vector through a linear layer, uses a temporal convolutional network to analyze the time dimension of the feature vector output by the linear layer, adds the position information of the time series through position encoding, and then converts it into an embedded vector through a 6-layer Transformer Block; The discretized latent space is designed as an [n, d] matrix containing n d-dimensional embedding vectors, and the embedding vector with the highest similarity to the embedding vector output by the encoder is searched in the discretized latent space to replace the embedding vector output by the encoder; The decoder converts the replaced embedding vector into a 1024-dimensional feature, uses a temporal convolutional network to analyze the time dimension of the feature, adds the position information of the time series through position encoding, and then converts the feature into the original data dimension through a 6-layer Transformer Block and a linear layer to complete the decoding.
3. The method according to claim 2, characterized in that According to the following formula, the total loss function of the 3D human posture reconstruction network is calculated: L=L motion +αL vq +βL reg Among them, α and β are coefficients; The action loss L in the total loss function motion The definition is as follows: L w is the WMPJPE loss, L t is the TC loss, L m is the MPJVE loss, is the weight coefficient; The WMPJPE loss is defined as follows: N is the number of key points of the human body, W is the weight of each joint, T is the number of frames in the input video sequence, and p i,j and gt i,j are the predicted value and true value of the 3D position corresponding to the ith joint in the jth frame, respectively. ||·||2 indicates L2 norm calculation; TC loss L t is defined as follows: p i,j and p i,j-1 are the predicted values of the 3D position corresponding to the i-th joint in the j-th frame and the j-1-th frame respectively; MPJVE Loss L m is defined as follows: The vector quantization loss L in the total loss function vq is defined as follows: L vq =||sg[z e (x)]-e k ||2 z e (x) is the continuous feature vector generated by the encoder, e k is the discrete code vector stored in the discretized latent space, and sg[·] represents the stop gradient operation, indicating that z will not be updated during back propagation. e (x); Embedding regression loss L in the total loss function reg is defined as follows: L reg =||z e (x)-sg[e k ]||2 Among them, z e (x) is the continuous feature vector generated by the encoder, e k is the discrete code vector stored in the discretized latent space, and sg[·] represents the stop gradient operation, indicating that e will not be updated during back propagation. k .
4. The method according to claim 1, characterized in that During the training process of the 3D human pose estimation network, only the sampler network is trained to encode the 2D human pose sequence into a feature vector consistent with the dimension of the embedding vector in the discretized latent space, so as to select the embedding vector with the smallest similarity distance in the discretized latent space for replacement, and then decode the replaced embedding vector through the decoder to obtain the 3D human pose.
5. The method according to claim 4, characterized in that The total loss of the 3D human pose estimation network is defined as follows: L=L reg +γL motion +δL sync Among them, γ and δ are coefficients, and the regression loss L reg is defined as follows: in, is the feature generated by the sampler network, e k is the discrete embedding vector stored in the discretized latent space, and sg[·] represents the stop gradient operation, indicating that e will not be updated during back propagation. k ; Action loss L motion The definition is as follows: Among them, L w is the WMPJPE loss, L t is the TC loss, L m is the MPJVE loss, is the weight coefficient; Synchronization loss L sync is defined as follows: in, They represent the features of 3D human posture data generated by 2D and prediction networks after being encoded by synchronous networks.
6. The method according to claim 1, characterized in that Get the training dataset, including: Preprocessing is performed based on the video image sequence to obtain a 2D human body posture sequence and a 3D human body posture sequence; Random misalignment is performed according to the 2D human body posture sequence and the 3D human body posture sequence to construct a plurality of training data pairs, wherein the training data pairs include positive sample pairs and negative sample pairs.
7. A video 3D human posture estimation device based on discretized latent space, characterized in that: include: An acquisition module is used to acquire a training data set, wherein the training data set includes a plurality of training data pairs, the plurality of training data pairs include positive sample pairs and negative sample pairs, and each of the training data pairs includes a 2D human body posture sequence and a 3D human body posture sequence; A first training module is used to train a pre-constructed synchronization network according to the training data set so that the synchronization network learns the potential connection between the 2D human body posture sequence and the 3D human body posture sequence; A second training module is used to construct a 3D human posture reconstruction network and train it based on the 3D human posture sequence in the training data set, wherein the 3D human posture reconstruction network includes an encoder, a discretized latent space, and a decoder; A third training module is used to construct a 3D human pose estimation network based on a sampler network and a discretized latent space and a decoder in a trained 3D human pose reconstruction network, train the sampler network in the 3D human pose estimation network according to the 2D human pose sequence in the training data set, and perform weak supervision by calculating the auxiliary loss through the synchronization network during the training process, so that the 3D human pose estimation network can estimate the corresponding 3D human pose based on the input 2D human pose sequence; The processing module is used to perform three-dimensional human posture estimation based on a pre-trained 2D human posture estimation network and a 3D human posture estimation network.
8. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for estimating three-dimensional human posture in a video based on a discretized latent space is implemented as claimed in any one of claims 1 to 6.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the video three-dimensional human posture estimation method based on discretized latent space as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN116543076A
Knowledge migration point cloud human body posture estimation model training and recognition method
CN117612200A
Three-dimensional modeling method and device based on monocular image and medium
CN118470208A
Three-dimensional human body posture estimation method and system based on video sequence spatio-temporal context
CN118823833A
Method for training a pose estimator
US20240212195A1