3D Human Joint Point Estimation Method, System and Device Based on Monocular Sequential Images

A two-stage method using spatial and temporal networks with a novel Transformer model addresses the inefficiencies of existing methods by prioritizing easier joints and reducing redundant features, enhancing the accuracy of three-dimensional human body joint estimation.

CN115223201BActive Publication Date: 2025-07-15ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210835636.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2025-07-15
Estimated Expiration
2042-07-15

AI Technical Summary

Technical Problem

The existing three-dimensional human joint node estimation method based on Transformer has large redundant information extraction and estimation errors when processing sequences, and cannot effectively utilize local information, resulting in insufficient accuracy of the model in human skeleton joint node applications.

Method used

The three-dimensional human joint node estimation method based on monocular sequence images is adopted. By constructing a spatial feature extraction network and a timing feature extraction network, combining targeted spatiotemporal Transformer network model, the human bone chain structure characteristics are used to estimate the joint nodes in a layered manner, and combining timing convolution and GeLU activation function to perform intermediate supervision and time smoothness optimization.

Benefits of technology

It effectively reduces the estimation error in the estimation process, improves the accuracy and generalization ability of the model, solves the problem of redundant information and estimation error transmission, and improves the accuracy of 3D human joint estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223201B_ABST
    Figure CN115223201B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, system and device for estimating three-dimensional human joint points based on monocular sequence images. The estimation method first obtains a two-dimensional human joint sequence in each frame of the monocular sequence images, then performs filtering processing on the two-dimensional human joint point sequence, adds position encoding to the two-dimensional human joint point sequence and inputs it into a newly constructed spatial feature extraction network to extract the spatial features of the human joint points in each frame of the monocular sequence images, thereby obtaining a three-dimensional human joint point pose feature sequence of n frames. Then, the three-dimensional human joint point pose feature sequence of n frames is input into a temporal feature extraction network to obtain the three-dimensional human joint point features of the intermediate frame. Finally, the three-dimensional human joint point features of the intermediate frame are input into a fully connected layer module two to obtain the three-dimensional human joint point coordinates of the intermediate frame. This estimation method can effectively reduce the estimation error in the process of joint point estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of joint point positioning, and particularly to a method, system and device for estimating three-dimensional human joint points based on monocular sequence images. Background Art

[0002] Human skeletal joint points play a crucial role in describing human postures and predicting human behaviors. The capture of human skeletal joint points is widely applied in fields such as video games, robot development, virtual reality, etc. In recent years, with the rapid development of artificial intelligence and image computing power, how to enable a machine to imitate the thinking characteristics of humans to capture and calculate joint points, and the calculation of joint points is more accurate than that of humans, has become an urgent problem to be solved in the current field of joint point positioning.

[0003] Due to its high efficiency, scalability and powerful modeling ability, Transformer has become the de facto model in natural language processing (NLP) and is now being introduced into computer vision tasks such as image classification, object detection and semantic segmentation. Thanks to the self-attention mechanism, Transformer can capture the internal correlations of long-time series inputs and is not restricted by their distances, and the global correlation features across long input sequences can be clearly captured. This makes it particularly suitable for architectures of sequence data problems and thus can naturally be extended to the estimation of three-dimensional human joint points of sequence data.

[0004] However, recent research has shown that in terms of visual tasks, in existing Transformer-based joint point estimation methods, there is a large amount of redundancy when processing sequences, and the memory requirements of the model are relatively high. In addition, a large amount of valuable information may be lost when extracting temporal feature information between frames, resulting in a large estimation error and the inability to make good use of local information, thus limiting the application of Transformer in human skeletal joint points. Summary of the Invention

[0005] Based on this, in view of the technical problem of relatively large estimation errors in the estimation of three-dimensional human joint points in the prior art, the present invention provides a method, system and device for estimating three-dimensional human joint points based on monocular sequence images.

[0006] The present invention discloses a method for estimating three-dimensional human joint points based on monocular sequence images, which includes the following steps:

[0007] S1: Collect multiple frames of monocular sequence images containing human joint actions, and obtain the two-dimensional human joint sequences in each frame of monocular sequence images.

[0008] S2: Perform filtering processing on the two-dimensional human joint point sequences, and then add position encoding to the two-dimensional human joint point sequences.

[0009] S3: Input the J two-dimensional human joint sequences after position encoding into a newly constructed spatial feature extraction network to extract the spatial features of the human joint points in each frame of the monocular sequence image, thereby obtaining an n-frame three-dimensional human joint point pose feature sequence. Among them, the construction method of the spatial feature extraction network includes the following steps:

[0010] S31: Divide each joint point of the human body into multiple joint sets according to the chain structure of the human skeleton.

[0011] S32: Assign multiple joint sets to multiple levels with different estimation difficulties according to the motion amplitude characteristics of each joint set.

[0012] S33: Divide multiple joint sets in each level into multiple channels representing different affiliated parts according to the affiliated characteristics of the chain structure, so that multiple joint sets are combined into a tree-like series structure. Among them, multiple levels correspond to the extension direction of the tree-like series structure in the order from easy to difficult.

[0013] S34: Design multiple groups of spatial feature extraction modules corresponding to multiple joint sets respectively, thereby constituting the spatial feature extraction network. Each group of spatial feature extraction modules is used to extract the joint point spatial feature vectors of the corresponding joint sets.

[0014] S4: Input the n-frame three-dimensional human joint point pose feature sequence into a temporal feature extraction network to obtain the three-dimensional human joint point features of the intermediate frame. Among them, the temporal feature extraction network includes multiple groups of temporal feature extraction modules. Each group of temporal feature extraction modules is used to extract the pose features of multiple consecutive human joint points, and then merge adjacent frames to reduce the frame sequence of multiple human joint point poses, and obtain the three-dimensional human joint point coordinates of the target frame through multiple groups of temporal feature extraction modules.

[0015] S5: Input the three-dimensional human joint point features of the intermediate frame into a fully connected layer module two with a dimension of T*J to obtain the three-dimensional human joint point coordinates of the intermediate frame.

[0016] As a further improvement of the present invention, in S32 and S33, a total of eight joint sets are provided. There are four levels arranged in ascending order of estimation difficulty: the first level, the second level, the third level, and the fourth level. A total of three channels representing different affiliated parts are provided: the first channel, the second channel, and the third channel. The first channel corresponds to the head, the second channel corresponds to the hand, and the third channel corresponds to the leg.

[0017] Among them, the first level is assigned one joint set, and this joint set includes a total of five joint points: the coccyx, the spine, the chest, the left hip bone, and the right hip bone.

[0018] The second level is allocated with three joint sets. The joint set located in the first channel includes the neck. The joint sets located in the second channel include the left shoulder and the right shoulder. The joint sets located in the third channel include the left knee and the right knee.

[0019] The third level is allocated with three joint sets. The joint set located in the first channel includes the head. The joint sets located in the second channel include the left elbow and the right elbow. The joint sets located in the third channel include the left ankle and the right ankle.

[0020] The fourth level is allocated with one joint set, and this joint set includes the left wrist and the right wrist.

[0021] As a further improvement of the present invention, in S3 and S4, the spatial feature extraction network and the temporal feature extraction network are connected in series, thereby constituting a targeted spatio-temporal Transformer network model. The targeted spatio-temporal Transformer network model is improved based on the classical Transformer network. The construction method of the targeted spatio-temporal Transformer network model includes the following steps:

[0022] (1) Obtain the standard Transformer network as the basic framework of the spatial feature extraction module and the temporal feature extraction module, use the GeLU function as the activation function of the spatial feature extraction module and the temporal feature extraction module respectively, and incorporate the random regularization function in the activation.

[0023] (2) Replace the fully connected layer in each group of temporal feature extraction modules with a strided convolutional unit. The strided convolutional unit is used to reduce the time dimension between layers.

[0024] (3) Adopt the residual structure two in each group of temporal feature extraction modules to realize the connection between units, and use the average pooling function as the dimensionality reduction function of the residual structure.

[0025] (4) Add a fully connected layer module one with a dimension of T*J at the output end of the spatial feature extraction network, and also add a fully connected layer module two at the output end of the temporal feature extraction network, thereby constructing the targeted spatio-temporal Transformer network model. The fully connected layer module one is used to obtain a three-dimensional human joint point sequence of n frames according to the three-dimensional human joint point pose feature sequence of n frames.

[0026] As a further improvement of the present invention, the expression formula of the activation function of the spatial feature extraction module and the temporal feature extraction module is:

[0027]

[0028] As a further improvement of the present invention, after constructing the targeted spatio-temporal Transformer network model, the targeted spatio-temporal Transformer network model is also trained, and the training process is as follows:

[0029] Obtain a standard monocular sequence image of multiple frames of known joint point coordinate real data, and mix the standard monocular sequence image with the corresponding monocular sequence image to be estimated to obtain a random set of monocular sequence images. Use the set of monocular sequence images as the sample data to form a data set for model training, and divide the data set into a training set and a validation set.

[0030] Complete the initialization of the targeted spatio-temporal Transformer network model, use the training set to train the targeted spatio-temporal Transformer network model, use the validation set to verify the training effect of the targeted spatio-temporal Transformer network model, and thus obtain the trained targeted spatio-temporal Transformer network model.

[0031] As a further improvement of the present invention, each spatial feature extraction module includes: a layer normalization unit one, a multi-head attention unit one, two fully connected layer units one, and a residual structure one.

[0032] Among them, the feature vector generated by each spatial feature extraction module generates a three-dimensional pose through the fully connected layer module one, and then calculates the intermediate supervision loss function L J For fast backpropagation, the intermediate supervision loss function L J Is set to optimize the average Euclidean distance between the joint points of each spatial feature extraction module and the corresponding joint points in the real data.

[0033] Take the average Euclidean distance between the three-dimensional human joint point sequence of n frames generated by the fully connected layer module one and the corresponding joint points in the n-frame corresponding real data as the sequence loss function L of the spatial feature extraction network K :

[0034]

[0035] In the formula, Represents the estimated three-dimensional joint point position of joint i at frame t. Represents the real three-dimensional joint point position of joint i at frame t.

[0036] The total loss L of the spatial feature extraction network S The expression formula of is:

[0037] L S =λ K L K +λ J L J

[0038] In the formula, λ K and λ J are weight factors corresponding to the intermediate supervision loss function and the sequence loss function respectively.

[0039] As a further improvement of the present invention, each group of temporal feature extraction modules includes: a layer normalization unit two, a multi-head attention unit two, two consecutive one-dimensional convolutional units, and a residual structure two.

[0040] Among them, the single-frame loss L T is used to minimize the distance between the three-dimensional joint coordinates X of the intermediate frame output by the temporal feature extraction network and the corresponding real three-dimensional human joint coordinates Y. The expression formula of L T is:

[0041]

[0042] As a further improvement of the present invention, the expression formula of the total loss L of the targeted spatio-temporal Transformer network model is:

[0043] L = λ S L S + λ T L T

[0044] In the formula, λ S and λ T are weight factors related to the spatial feature extraction network and the temporal feature extraction network respectively.

[0045] The present invention also discloses a three-dimensional human joint point estimation system based on monocular sequence images, which adopts any one of the above three-dimensional human joint point estimation methods based on monocular sequence images. The three-dimensional human joint point estimation system based on monocular sequence images includes: an image acquisition module, a preprocessing module, a spatial feature extraction network, a temporal feature extraction network, and a fully connected layer module two.

[0046] The image acquisition module is used to acquire multiple frames of monocular sequence images containing human joint actions, and obtain the two-dimensional human joint sequence in each frame of monocular sequence image.

[0047] The preprocessing module is used to filter the two-dimensional human joint point sequence, and then add position encoding to the two-dimensional human joint point sequence.

[0048] The spatial feature extraction network is used to extract the spatial features of the human joint points in each frame of monocular sequence image, and then obtain a three-dimensional human joint point pose feature sequence of n frames. The spatial feature extraction network includes multiple groups of spatial feature extraction modules. Each group of spatial feature extraction modules is used to extract the joint point spatial feature vectors of the corresponding joint set.

[0049] The temporal feature extraction network is used to obtain the 3D human joint point features of the middle frame based on the 3D human joint point pose feature sequence of n frames. The temporal feature extraction network includes multiple groups of temporal feature extraction modules. Each group of temporal feature extraction modules is used to extract the pose features of human joints in multiple consecutive frames, and then merge adjacent frames to reduce the frame sequence of the pose features of human joints in multiple frames. The 3D human joint point coordinates of the target frame are obtained through multiple groups of temporal feature extraction modules.

[0050] The fully connected layer module two is used to obtain the 3D human joint point coordinates of the middle frame based on the 3D human joint point features of the middle frame.

[0051] The present invention also discloses a 3D human joint point estimation device based on a monocular sequence image, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the 3D human joint point estimation method according to any one of the above are implemented.

[0052] Compared with the prior art, the technical solution disclosed by the present invention has the following beneficial effects:

[0053] 1. The estimation method uses a two-stage method to estimate the position of the 3D human joint points in the middle frame from the monocular sequence image. First, the newly constructed spatial feature extraction network is used to extract the spatial features of the human joint points in each frame of the monocular sequence image from the 2D human joint sequence, and then the 3D human joint point pose feature sequence is obtained. And the constructed spatial feature extraction network first determines the core five joint points according to the chain structure characteristics of the human skeleton, and then sequentially estimates the joint points close to the edge of the chain structure. Using the constraints between the joint points in the chain structure, from easy to difficult, layer by layer, it effectively improves the accuracy of the model, and to a certain extent alleviates the problem that the estimation error of one joint point in the overall estimation is transmitted to all joint points, and finally can effectively reduce the estimation error in the process of joint point estimation.

[0054] 2. This estimation method improves the network structure of PoseFormer and proposes a targeted spatio-temporal Transformer network model. First, it combines a temporal convolutional structure to process the temporal features between frames, replaces the fully connected layer in the Transformer with a strided convolution, gradually reduces the sequence length, effectively solves the redundancy problem of the temporal features of adjacent frames, and reduces the interference of invalid features. In addition, GeLU is also used as the activation function, and random regularization is incorporated into the activation function, effectively improving the generalization of the model. Finally, the improved Transformer balances the calculations in the MLP to build a deeper model, aggregates information in a global and local manner, improves the model capacity, and at the same time applies the idea of intermediate supervision to supervise the loss function of the sequence images spatially and temporally. The structure based on the spatial Transformer is more helpful for learning the spatial information feature extraction between single-frame human joint points, while the structure based on the temporal Transformer focuses on the temporal information feature extraction between frames, enhancing temporal smoothness.

[0055] 3. The beneficial effects of this estimation system and device are the same as those of the above estimation method and will not be elaborated here. Brief Description of the Drawings

[0056] Figure 1 Schematic diagrams of two ideas for applying the Transformer to human joint points in Embodiment 1 of the present invention;

[0057] Figure 2 Schematic diagram of the algorithm structure based on the pure Transformer module in Embodiment 1 of the present invention;

[0058] Figure 3 Flowchart of three-dimensional human joint point estimation based on monocular sequence images in Embodiment 1 of the present invention;

[0059] Figure 4 System block diagram when the targeted spatio-temporal Transformer network model executes the estimation method in Embodiment 1 of the present invention;

[0060] Figure 5 Comparison chart of the standard deviations of different joint points in a series of frames in Embodiment 1 of the present invention;

[0061] Figure 6 Schematic diagram of the division of the human joint structure in Embodiment 1 of the present invention;

[0062] Figure 7 Schematic diagram of the structure of the spatial feature extraction network in Embodiment 1 of the present invention;

[0063] Figure 8 Module schematic diagram of the spatial feature extraction module in Embodiment 1 of the present invention;

[0064] Figure 9 This is the schematic diagram of the timing feature extraction module in Embodiment 1 of the present invention;

[0065] Figure 10 This is the comparison chart of different activation functions in Embodiment 1 of the present invention. Detailed implementation manners

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0067] It should be noted that when a component is referred to as "installed on" another component, it can be directly on the other component or there may also be an intermediate component. When a component is considered to be "disposed on" another component, it can be directly disposed on the other component or there may be an intermediate component at the same time. When a component is considered to be "fixed to" another component, it can be directly fixed to the other component or there may be an intermediate component at the same time.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.

[0069] Embodiment 1

[0070] Due to its high efficiency, scalability, and powerful modeling capabilities, Transformer has become the de facto model for natural language processing (NLP) and is now being introduced into computer vision tasks such as image classification, object detection, and semantic segmentation. Thanks to the self-attention mechanism, Transformer can capture the internal correlations of long-time series inputs and is not restricted by their distances. The global correlation features across long input sequences can be clearly captured. This makes it particularly suitable for architectures of sequence data problems and thus can naturally be extended to the three-dimensional human joint point estimation of sequence data.

[0071] However, recent studies have shown that Transformers require specific designs to achieve comparable performance to their CNN counterparts for visual tasks. Specifically, they often require very large training datasets, or data augmentation and regularization if applied to smaller datasets. In addition, existing visual Transformers are mainly limited to image classification, object detection, and segmentation, but there is still little work on how to leverage the power of Transformers for 3D human joint estimation.

[0072] See also Figure 1 , some scholars have tried to apply Transformer directly to two-dimensional to three-dimensional human posture estimation. Figure 1 One approach in (a) is to treat the entire 2D pose of each frame in a given sequence as a token. Although this approach is useful to a certain extent, it ignores the spatial relationship, that is, the information between joints in a 2D pose of a frame. Figure 1 Another approach in (b) is to treat the coordinates of each joint of each frame of the 2D pose as tokens and provide input consisting of these joints from all frames of the sequence. However, when using sequences with a higher number of frames as input, the number of tokens will become larger and larger. For example, a common input for 3D human pose estimation is a 243-frame sequence with 17 joints per frame, and the number of tokens is 243×17=4131. Since the Transformer model directly calculates each token with another token, the memory requirement of the model reaches an unreasonable level.

[0073] See also Figure 2 To solve this problem, Zheng et al. proposed PoseFormer. PoseFormer directly uses two different Transformer modules to model the input two-dimensional joint point sequence from the spatial and temporal aspects. Specifically, the spatial Transformer module constructed by PoseFormer encodes the local relationship between the two-dimensional joint points in each frame, and the self-attention layer focuses on the spatial position information of the two-dimensional joint points and returns the potential feature representation. Next, the temporal Transformer module will analyze the global temporal dependencies between the spatial feature representations of each frame and finally generate an accurate three-dimensional pose estimate. PoseFormer can certainly achieve good results. It not only extracts features in the spatial and temporal dimensions, but also does so without generating huge token counts for long input sequences.

[0074] However, the extraction of temporal feature information between frames by PoseFormer actually contains a lot of redundancy in the pose estimation based on the joint sequence because adjacent frames are very similar. Therefore, adjacent frames should be gradually merged to reduce the sequence length until the three-dimensional joint representation of an intermediate frame. One method is to perform a pooling operation after the MLP (Multi-Layer Perception). However, this method may lose a large amount of valuable information and does not make good use of local information.

[0075] Based on this, this embodiment provides a method for estimating three-dimensional human joint points based on monocular sequence images, which will be introduced and verified below.

[0076] Please refer to Figure 3 and Figure 4 , this embodiment provides a method for estimating three-dimensional human joint points based on monocular sequence images, which includes steps S1 to S5.

[0077] S1: Collect multiple frames of monocular sequence images containing human joint actions, and obtain the two-dimensional human joint sequence in each frame of the monocular sequence image.

[0078] In this embodiment, CPN can be used as a two-dimensional joint detector to output the two-dimensional human pose sequence of the video sequence.

[0079] S2: Filter the two-dimensional human joint point sequence, and then add position encoding to the two-dimensional human joint point sequence.

[0080] S3: Input the J two-dimensional human joint sequences after position encoding into a spatial feature extraction network to extract the spatial features of the human joint points in each frame of the monocular sequence image, and then obtain a three-dimensional human joint point pose feature sequence of n frames. Among them, the construction method of the spatial feature extraction network includes the following steps:

[0081] S31: Divide each joint point of the human body into multiple joint sets according to the chain structure of the human skeleton.

[0082] S32: According to the motion amplitude characteristics of each joint set, allocate multiple joint sets to multiple levels with different estimation difficulties.

[0083] S33: According to the belonging characteristics of the chain structure, divide the multiple joint sets in each level into multiple channels representing different belonging parts, so that the multiple joint sets are combined into a tree-like series structure. Among them, the multiple levels correspond to the extension direction of the tree-like series structure in the order from easy to difficult.

[0084] S34: Design multiple sets of spatial feature extraction modules corresponding to multiple joint sets respectively, and then form a spatial feature extraction network. Each set of spatial feature extraction modules is used to extract the joint point spatial feature vectors of the corresponding joint set.

[0085] The joint point estimation method in this embodiment is based on the deep learning method of overall regression. Generally, it is difficult to process joint points with high complexity. The complexity of human joint points is closely related to their corresponding movement amplitudes, and can be judged by the coordinate standard deviation of each joint point in each direction during human movement for a period of time.

[0086] Although the range of motion of the human body during different actions is different, no matter what activities the human body is performing, along the human torso, the movement amplitude of the joint points will gradually increase in one or more directions. Under the same action, even in the low-dimensional human joint point trajectory that only analyzes the plane, different joint parts will have different movement amplitudes.

[0087] Please refer to Figure 5 , in this article, the true position coordinates of each joint of an actor in the Human3.6M dataset in the x, y, and z directions during different actions are selected, and their standard deviations are taken to analyze the movement amplitude of each joint point in this action. As Figure 4 shown, we found that whether it is a discussion action, a posed action, or a sitting down action, the movement amplitude of the joint points of the human wrist is always the largest among all joint points, and with the chain structure of the human skeleton, from the wrist to the elbow and then to the shoulder, the standard deviation always shows a downward trend. Therefore, it can be concluded that the movement amplitude of human joint points gradually increases with the chain structure of the skeleton. That is to say, the closer the joint point is to the wrist and ankle, the greater the movement amplitude and the greater the estimation difficulty. The joint points of the hip bone, spine, chest, etc. always have the smallest movement amplitude and the smallest estimation difficulty, and can be regarded as the root nodes of the human skeleton joint point structure and estimated first in the algorithm. After determining the positions of the joint points that are easy to estimate, and according to the constraints of the human skeleton joint point structure, it is easier to estimate the more difficult-to-estimate joint points close to the wrist and ankle.

[0088] Therefore, this embodiment proposes to divide all human joint points according to the standard deviation of their coordinates in all frames of a video.

[0089] Please refer to Figure 6 , this embodiment provides a targeted human joint structure as shown in Figure 6 , and the human joint points can be divided into four levels: the first level, the second level, the third level, and the fourth level. There can be eight joint sets. A total of three channels representing different affiliated parts are set: the first channel, the second channel, and the third channel. The first channel corresponds to the head, the second channel corresponds to the hand, and the third channel corresponds to the leg.

[0090] Among them, the first level contains 5 nodes, namely the coccyx, spine, chest, left hip bone, and right hip bone. The movement range of these 5 nodes is small and the estimation difficulty is low. In 3D human joint point estimation, they are regarded as core joint points and estimated first. The second level contains a total of five joint points, namely the left knee, right knee, left shoulder, right shoulder, and neck. The estimation difficulty of the joint points in this part is slightly increased and they are estimated after the first level in the algorithm. The third level contains a total of five joint points, namely the left ankle, right ankle, left elbow, right elbow, and head. The joint points in this part are close to the edge of the human joint point structure and the estimation difficulty is relatively large. They are estimated after the second level in the algorithm. The fourth level is the two joint points of the left wrist and right wrist with the highest movement complexity. These joint points are at the edge of the human joint point structure and have the largest movement range. They are estimated last in the algorithm. At the same time, since the arms, legs, and head belong to different chain structures, in order to distinguish them, in this embodiment, the joint points are re-divided according to these three chain structures on the basis of the previous division, that is, three channels.

[0091] Please refer to Figure 7 , the spatial feature extraction network in this embodiment includes 8 spatial feature extraction modules. Each spatial feature extraction module is further divided into 3 channels on the basis of the above four levels. In the spatial feature extraction network, the joint points with small movement range and low complexity are estimated first, and then the joint points adjacent to them, which are close to the edge of the chain structure and have high complexity, are estimated. Since there is a certain constraint effect between adjacent human joint points in the chain structure, after determining one joint point, it is easier to estimate the other joint point. Therefore, the model inputs the J feature vectors after adding position encoding into the first layer of the spatial feature extraction Transformer module to adjust the five joint points in the first layer with the lowest complexity, namely the hip bones, spine, chest, left and right hip bones. Then, according to the characteristics of the human chain structure, they are mapped into 3 d m -dimensional feature vectors and input into 3 channels respectively to adjust the joint points of the neck, hands, and legs. Then, according to the targeted human joint point structure, the estimation of the joint points is completed step by step from low to high complexity and from simple to difficult. For example, the second channel is used to optimize the joint points on the hand. First, the left and right shoulders with lower complexity are optimized, then the left and right elbows with higher complexity are optimized, and finally the left and right wrists with the highest complexity are optimized. Finally, the three channels are summarized to obtain a human pose in which all joint points are optimized.

[0092] Please refer to Figure 8, in this embodiment, each spatial feature extraction module includes: a layer normalization unit 1, a multi-head attention unit 1, two fully connected layer units 1, and a residual structure 1. The activation function can use the GeLU function to add a regularization function during activation to improve the generalization ability. The feature vectors generated by each group of spatial feature extraction modules pass through the fully connected layer module 1 to generate a three-dimensional pose, and then the intermediate supervision loss function L is calculated. J For fast backpropagation, the intermediate supervision loss function L J is set to optimize the average Euclidean distance between the joint points of each spatial feature extraction module and the corresponding joint points in the real data.

[0093] In this embodiment, first, according to the chain structure of the human skeleton, the core five joint points are determined, and then the joint points near the edge of the chain structure are estimated in turn. Using the constraints between the joint points in the chain structure, from easy to difficult, to a certain extent, it alleviates the problem that the estimation error of one joint point caused by the overall estimation is transmitted to all joint points.

[0094] In addition, we also note that the model directly supervised at the single target frame scale does not consider the temporal smoothness between frames, while the model supervised only at the full sequence scale cannot explicitly learn the specific representation of the target frame. To combine these two scale constraints into the framework, a complete-to-single scheme is proposed to further refine the intermediate prediction to produce more accurate estimates, rather than using a single component with a single output. We implement the full sequence scale supervision by imposing an additional temporal smoothness constraint during training. The three-dimensional human joint point pose feature sequence of n frames finally output by the spatial feature extraction network can be input into the fully connected layer module 1. The average Euclidean distance between the three-dimensional human joint point sequence of n frames generated by the fully connected layer module 1 and the corresponding joint points in the n-frame corresponding real data is used as the sequence loss function L of the spatial feature extraction network. K Perform intermediate supervision on the human joint points as a whole at the spatial level to improve the temporal consistency prediction of the single-frame sequence. Coupled with the intermediate supervision loss function L generated for each spatial extraction module before J , L K and the total loss L of the entire spatial feature extraction network S The expression formulas are as follows:

[0095]

[0096] L S = λ K L K + λ J L J

[0097] In the formula, Denote the estimated 3D joint position of joint i at frame t. Denote the true 3D joint position of joint i at frame t. λ K and λ J are the weight factors corresponding to the intermediate supervision loss function and the sequence loss function respectively.

[0098] S4: Input the 3D human joint pose feature sequence of n frames into a temporal feature extraction network to obtain the 3D human joint features of the intermediate frame. Among them, the temporal feature extraction network includes multiple groups of temporal feature extraction modules. Each group of temporal feature extraction modules is used to extract the pose features of multiple consecutive frames of the human joints, and then merge adjacent frames to reduce the frame sequence of the human joint poses of multiple frames. The 3D human joint coordinates of the target frame are obtained through multiple groups of temporal feature extraction modules.

[0099] In the aforementioned spatial feature extraction network, the spatial features of the human joints in each frame of the image are extracted. The temporal feature extraction network will be introduced below.

[0100] Please refer to Figure 9 , each group of temporal feature extraction modules includes: layer normalization unit two, multi-head attention unit two, two consecutive one-dimensional convolutional units, and residual structure two. Based on the method of temporal convolution to process sequences of different input lengths, we propose to replace the fully connected layer in the temporal feature extraction module with strided convolution to gradually reduce the sequence length. Among them, the self-attention unit two is used to extract global temporal features, and the strided convolution unit helps to extract the temporal features of adjacent frames. In this way, the time dimension is gradually reduced from one layer to another, and nearby poses are merged into a short sequence length representation. The temporal feature extraction network aggregates information in a global and local manner. More importantly, it reduces the redundancy of all frames, thereby improving the capacity of the model and enhancing the temporal smoothness. At the same time, in order to prevent the phenomenon of gradient disappearance or gradient explosion, we adopt residual structures in the multi-head attention unit two and the fully connected layer as the feed-forward network respectively, and use the average pooling function as the dimensionality reduction function of the residual structure to maximize the retention of the feature information of the residual structure.

[0101] Finally, after the 3D human joint pose feature sequence of n frames passes through the temporal feature extraction network, the 3D human joint features of the intermediate frame are obtained.

[0102] The time feature extraction network is a structure with gradually reduced dimensions layer by layer. Use past and future data to predict the 3D human poses of all frames in the input sequence as the output. In this embodiment, the single-frame loss L T is used to minimize the distance between the 3D joint coordinates X of the intermediate frame output by the temporal feature extraction network and the corresponding true 3D human joint coordinates Y. The expression formula of L T is:

[0103]

[0104] In this embodiment, the spatial feature extraction network and the temporal feature extraction network are connected in series to form a targeted spatio-temporal Transformer network model. The targeted spatio-temporal Transformer network model is improved based on the classical Transformer network. The construction method of the targeted spatio-temporal Transformer network model includes the following steps:

[0105] (1) Obtain the standard Transformer network as the basic framework of the spatial feature extraction module and the temporal feature extraction module, use the GeLU function as the activation function of the spatial feature extraction module and the temporal feature extraction module respectively, and incorporate the random regularization function in the activation.

[0106] (2) Replace the fully connected layer in each group of temporal feature extraction modules with a strided convolutional unit. The strided convolutional unit is used to reduce the time dimension between layers.

[0107] (3) Adopt the residual structure two in each group of temporal feature extraction modules to realize the connection between units, and use the average pooling function as the dimensionality reduction function of the residual structure.

[0108] (4) Add a fully connected layer module one with a dimension of T*J at the output end of the spatial feature extraction network, and also add a fully connected layer module two at the output end of the temporal feature extraction network, thereby constructing the targeted spatio-temporal Transformer network model. The fully connected layer module one is used to obtain a three-dimensional human joint point sequence of n frames according to the three-dimensional human joint point pose feature sequence of n frames.

[0109] After constructing the targeted spatio-temporal Transformer network model, the targeted spatio-temporal Transformer network model is also trained, and the training process is as follows:

[0110] Obtain the standard monocular sequence images of multiple frames of known joint point coordinate real data, and mix the standard monocular sequence images with the corresponding monocular sequence images to be estimated to obtain a random monocular sequence image set. Use the monocular sequence image set as the sample data to form the dataset for model training, and divide the dataset into a training set and a validation set.

[0111] Complete the initialization of the targeted spatio-temporal Transformer network model, use the training set to train the targeted spatio-temporal Transformer network model, use the validation set to verify the training effect of the targeted spatio-temporal Transformer network model, and thus obtain the trained targeted spatio-temporal Transformer network model.

[0112] S5: Input the 3D human joint point features of the middle frame into a fully connected layer module two with a dimension of T*J to obtain the 3D human joint point coordinates of the middle frame. The expression formula for the total loss L of the entire targeted spatio-temporal Transformer network model is:

[0113] L = λ S L S + λ T L T

[0114] In the formula, λ S and λ T are weight factors related to the spatial feature extraction network and the temporal feature extraction network respectively.

[0115] In recent years, with the continuous development of the network, the training of neural networks using the sigmoid activation function has been proven to be worse than that of the non-smooth and low-probability ReLU. The latter is generally faster in training speed and better in convergence effect than the sigmoid function. Based on the successful experience of ReLU, an optimized activation function called ELU allows non-linear functions like ReLU to output values less than 0, which can improve the training efficiency in some cases. In short, for neural networks, the selection of activation functions is very necessary to prevent neural networks from becoming linear deep networks.

[0116] Non-linear activation functions can fit the data well. To avoid overfitting, regularization needs to be added to improve its generalization ability. Therefore, network designers often face the problem of how to choose random regularization methods. For example, applying Dropout, and the regularization function is separate from the activation function. The random regularizer dropout creates a pseudo-ensemble by randomly changing some activation decisions by multiplying with zeros randomly. The non-linear activation function and dropout thus jointly determine the output of the neuron, but the randomness of the regularizer dropout is independent of the input and lacks flexibility.

[0117] Hendrycks D and Gimpel K proposed a new non - linear activation function, namely the Gaussian Error Linear Unit (GELU). It is related to stochastic regularizers because it is an optimization of the stochastic regularizer Dropout. It should be noted that both ReLU and Dropout output the result of a neuron. Among them, the former deterministically multiplies the input by 0 or 1 as the output, while the latter randomly multiplies by 0. And GELU also realizes this function by multiplying the input by 0 or 1, but whether the input is multiplied by 0 or 1 is randomly selected depending on the distribution of the input itself. In other words, whether it is 0 or 1 depends on the probability that the current input is greater than the rest of the inputs. This indicates that the neuron is more likely to be output. This special non - linear activation function outperforms the ReLU or ELU activation function in tasks in multiple fields.

[0118] Please refer to the figure. In this embodiment, the ReLU activation function, ELU activation function, and GELU activation function in 3D human pose estimation are also compared. The effects of the ReLU activation function (α = 1), ELU activation function (α = 1), and GELU activation function (μ = 0, σ = 1) are as shown in the figure. In this embodiment, we use an approximate GELU definition, that is

[0119]

[0120] To verify the 3D human joint point estimation method based on monocular sequence images proposed in this embodiment, the following performance verification tests are also carried out in this embodiment. The process of the performance verification test is as follows:

[0121] Experiment and Analysis

[0122] This embodiment explores the problem of estimating 3D human joint points from monocular sequence images using the recently popular Transformer model, and proposes a targeted spatio - temporal Transformer network. First, use CPN as a 2D joint detector to output the 2D human pose sequence of the video sequence, and then extract the feature information of the 2D human pose sequence from both spatial and temporal dimensions, and output the 3D human pose of the middle frame of the video sequence. To test the results of the new model. In this embodiment, the above - mentioned method is tested on the Human3.6M dataset, and the network is built using the Ubuntu system and the Pytorch framework, and the graphics card used is GTX1080Ti.

[0123] The experimental settings are as follows:

[0124] (1) Test the performance of the model under the standard protocol.

[0125] (2) Investigate the specific impacts of improvements in the spatial model, temporal model, and activation function on the experimental results.

[0126] (3) Conduct experiments on different combinations of model hyperparameters to explore the hyperparameter combination that can train the best results.

[0127] Standard protocol experiment

[0128] In this embodiment, the Human3.6M dataset is continuously introduced for model verification. According to the two-stage method from two-dimensional to three-dimensional, we use the CPN network as the two-dimensional joint detector, and then use the detected two-dimensional joint sequence as the input for training and testing. The experiment is based on Protocol 1, and S9 and S11 are used as the validation sets to calculate the error. The specific experimental results are shown in the table, and the last column provides the average value in all validation sets. For details, please refer to Table 1.

[0129] Table 1 Experimental results under standard protocol 1

[0130]

[0131] Table 1 (continued)

[0132]

[0133] In Table 1, the data unit is millimeters (mm). According to the test results of Protocol 1, the proposed targeted spatio-temporal Transformer network model in this embodiment has much better performance than the temporal convolutional network (4.6%). This clearly demonstrates the advantage of using the Transformer network to model human joint sequences in both time and space. From the data, it can be seen that the targeted spatio-temporal Transformer network model can more accurately predict difficult actions such as taking pictures, sitting down, walking the dog, and smoking. Different from other simple actions, the human body postures in these actions change faster, and some long-distance frames have strong correlations. In this case, global dependence plays an important role, and the self-attention mechanism in the Transformer network is particularly beneficial for extracting such features.

[0134] Under Protocol 1, the lowest average MPJPE of the targeted spatio-temporal Transformer network model proposed in this embodiment is 43.5 mm. Compared with the PoseFormer network proposed by Zheng et al., the improved targeted spatio-temporal Transformer network model in this embodiment reduces the MPJPE by about 1.58%. The reasons are as follows. First, the targeted spatio-temporal Transformer network model pays more attention to the research on the human chain structure, and the training of joint points for each frame is from easy to difficult, progressing layer by layer. Second, the targeted spatio-temporal Transformer network model uses a temporal convolutional module to replace the latter half of the MLP module, and gradually extracts temporal features in the way of dilated convolution, effectively improving the redundancy problem of features between adjacent frames. Finally, the targeted spatio-temporal Transformer network model uses GeLU as the activation function, incorporating random regularization into the activation function, effectively improving the generalization of the model.

[0135] Ablation experiment

[0136] To verify the contribution of individual components of the targeted spatio-temporal Transformer network model and the impact of hyperparameters on performance, we conducted extensive ablation experiments on the Human3.6M dataset under Protocol 1. This embodiment tested the impact of the improved structure of the model on the output results, and the specific situation is shown in Table 2.

[0137] Table 2 Network structure error analysis table

[0138]

[0139] As can be seen from Table 2, after changing the encoding layer of the original network to a hierarchical structure to extract spatial features, the error of the algorithm is reduced by about 0.3 mm; after changing the decoding layer of the original network to a convolutional structure to extract temporal features, the algorithm error is reduced by about 0.3 mm; after replacing the ReLU function with the GeLU function for the activation function, the algorithm error is reduced by about 0.2 mm.

[0140] Experiments show that the improved structures proposed in this paper are all practical and effective, and each part can improve the performance of the algorithm, giving positive feedback to the model. Combining these three improvement measures, the error is reduced by about 0.8 mm on the basis of the original network, effectively improving the performance of the model.

[0141] Parameter experiment

[0142] To verify the impact of hyperparameters on the performance of the targeted spatio-temporal Transformer network model, this embodiment also conducted hyperparameter experiments on the Human3.6M dataset under Protocol 1.

[0143] Table 3 Comparative study of different hyperparameter combinations

[0144]

[0145] As shown in Table 3, in this embodiment, various parameter combinations are explored to find the optimal network. c represents the dimensionality of the features embedded in the spatial Transformer, and L represents the number of layers used in the encoder of the Transformer model. In our targeted spatio-temporal Transformer model, the output of the spatial Transformer is flattened, and temporal position embeddings are added to form the input of the temporal Transformer encoder. Therefore, the dimensionality of the embedded features of the temporal Transformer encoder is c×j. The optimal parameters of our model are c = 32, L S = 4, L T = 4.

[0146] In summary, the three-dimensional human joint point estimation based on monocular sequence images provided in this embodiment has the following advantages:

[0147] This estimation method uses a two-stage method to estimate the positions of three-dimensional human joint points in the middle frame from monocular sequence images. First, a newly constructed spatial feature extraction network is used to extract the spatial features of the human joint points in each frame of the monocular sequence image from the two-dimensional human joint sequence, and then a three-dimensional human joint point pose feature sequence is obtained. And the constructed spatial feature extraction network first determines the core five joint points according to the chain structure characteristics of the human skeleton, and then sequentially estimates the joint points close to the edge of the chain structure. Using the constraints between the joint points in the chain structure, from easy to difficult, layer by layer, it effectively improves the accuracy of the model, and to a certain extent alleviates the problem that the estimation error of one joint point in the overall estimation is transmitted to all joint points, and finally effectively reduces the estimation error of the joint points.

[0148] This estimation method improves the network structure of PoseFormer and proposes a targeted spatio-temporal Transformer network model. First, it combines a temporal convolutional structure to process the temporal features between frames, replaces the fully connected layer in the Transformer with a strided convolution, gradually reduces the sequence length, effectively solves the redundancy problem of the temporal features of adjacent frames, and reduces the interference of invalid features. In addition, GeLU is also used as the activation function, and random regularization is incorporated into the activation function, effectively improving the generalization of the model. Finally, the improved Transformer balances the calculations in the MLP to build a deeper model, aggregates information in a global and local manner, improves the model capacity, and at the same time applies the idea of intermediate supervision to supervise the loss function of the sequence images in space and time respectively. The structure based on the spatial Transformer is more helpful for learning the extraction of spatial information features between single-frame human joint points, while the structure based on the temporal Transformer focuses on the extraction of temporal information features between frames, enhancing temporal smoothness.

[0149] Example 2

[0150] The present invention also discloses a three-dimensional human joint point estimation system based on monocular sequence images, which can adopt the three-dimensional human joint point estimation method based on monocular sequence images in Example 1. The three-dimensional human joint point estimation system based on monocular sequence images includes: an image acquisition module, a pre-processing module, a spatial feature extraction network, a temporal feature extraction network, and a second fully connected layer module. The image acquisition module is used to acquire multiple frames of monocular sequence images containing human joint actions and obtain the two-dimensional human joint sequence in each frame of monocular sequence image. The pre-processing module is used to filter the two-dimensional human joint point sequence and then add position encoding to the two-dimensional human joint point sequence.

[0151] The spatial feature extraction network is used to extract the spatial features of the human joint points in each frame of monocular sequence image, and then obtain a three-dimensional human joint point pose feature sequence of n frames. The spatial feature extraction network includes multiple groups of spatial feature extraction modules. Each group of spatial feature extraction modules is used to extract the joint point spatial feature vectors of the corresponding joint set.

[0152] The temporal feature extraction network is used to obtain the three-dimensional human joint point features of the intermediate frame according to the three-dimensional human joint point pose feature sequence of n frames. The temporal feature extraction network includes multiple groups of temporal feature extraction modules. Each group of temporal feature extraction modules is used to extract the pose features of multiple consecutive human joint points, and then merge adjacent frames to reduce the frame sequence of the pose of multiple human joint points. After passing through multiple groups of temporal feature extraction modules, the three-dimensional human joint point coordinates of the target frame are obtained.

[0153] The fully connected layer module two is used to obtain the three-dimensional human joint coordinates of the intermediate frame based on the three-dimensional human joint point features of the intermediate frame.

[0154] Embodiment 3

[0155] The present invention also discloses a three-dimensional human joint point estimation device based on monocular sequence images, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it can implement the steps of the three-dimensional human joint point estimation method in Embodiment 1.

[0156] The joint point estimation device may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers) that executes the program. The joint point estimation device in this embodiment at least includes, but is not limited to, a memory and a processor that can communicate with each other through a system bus.

[0157] In this embodiment, the memory (i.e., the readable storage medium) includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various data that have been output or will be output.

[0158] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data to implement the processing process of the joint point estimation method in the foregoing Embodiment 1, so as to accurately estimate the three-dimensional human joints in the monocular sequence images.

[0159] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0160] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation of the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A three-dimensional human joint point estimation method based on monocular sequence images, characterized in that, It includes the following steps: S1: Collect multiple frames of monocular sequence images containing human joint actions, and obtain the two-dimensional human joint sequence in each frame of the monocular sequence image; S2: Perform filtering processing on the two-dimensional human joint point sequence, and then add position encoding to the two-dimensional human joint point sequence; S3: Input the J two-dimensional human joint sequences after position encoding into a newly constructed spatial feature extraction network to extract the spatial features of the human joint points in each frame of the monocular sequence image, and then obtain a three-dimensional human joint point pose feature sequence of n frames; Among them, the construction method of the spatial feature extraction network includes the following steps: S31: Divide each joint point of the human body into multiple joint sets according to the chain structure of the human skeleton; S32: According to the motion amplitude characteristics of each joint set, allocate multiple joint sets to multiple levels with different estimation difficulties; S33: According to the subordination characteristics of the chain structure, divide the multiple joint sets in each level into multiple channels representing different subordination parts, so that multiple joint sets are combined into a tree-like series structure; among them, the multiple levels correspond to the extension direction of the tree-like series structure in the order from easy to difficult; S34: Design multiple groups of spatial feature extraction modules corresponding to multiple joint sets respectively, and then constitute the spatial feature extraction network; each group of spatial feature extraction modules is used to extract the spatial feature vectors of the joint points of the corresponding joint set; S4: Input the three-dimensional human joint point pose feature sequence of n frames into a temporal feature extraction network to obtain the three-dimensional human joint point features of the intermediate frame; among them, the temporal feature extraction network includes multiple groups of temporal feature extraction modules; each group of temporal feature extraction modules is used to extract the pose features of human joint points in multiple consecutive frames, and then merge adjacent frames to reduce the frame sequence of the pose of human joint points in multiple frames, and obtain the three-dimensional human joint point coordinates of the target frame through multiple groups of temporal feature extraction modules; S5: Input the three-dimensional human joint point features of the intermediate frame into a fully connected layer module two with a dimension of T*J to obtain the three-dimensional human joint point coordinates of the intermediate frame.

2. The three-dimensional human joint point estimation method based on monocular sequence images according to claim 1, wherein In S32 and S33, a total of eight joint sets are set; four levels are set in order of increasing estimation difficulty: the first level, the second level, the third level, and the fourth level; a total of three channels representing different subordination parts are set: the first channel, the second channel, and the third channel; the first channel corresponds to the head, the second channel corresponds to the hands, and the third channel corresponds to the legs; Among them, the first level is allocated one joint set, and this joint set includes five joint points: the coccyx, spine, chest, left hip bone, and right hip bone; The second level is allocated three joint sets. The joint set located in the first channel includes the neck; the joint set located in the second channel includes the left shoulder and the right shoulder; the joint set located in the third channel includes the left knee and the right knee; The third level is allocated three joint sets. The joint set located in the first channel includes the head; the joint set located in the second channel includes the left elbow and the right elbow; the joint set located in the third channel includes the left ankle and the right ankle; The fourth level is allocated one joint set, and this joint set includes the left wrist and the right wrist.

3. The three-dimensional human joint point estimation method based on monocular sequence images according to claim 1, wherein In S3 and S4, the spatial feature extraction network and the temporal feature extraction network are connected in series to form a targeted spatio-temporal Transformer network model; the targeted spatio-temporal Transformer network model is improved based on the classical Transformer network; the construction method of the targeted spatio-temporal Transformer network model includes the following steps: (1) Obtain the standard Transformer network as the basic framework of the spatial feature extraction module and the temporal feature extraction module, use the GeLU function as the activation function of the spatial feature extraction module and the temporal feature extraction module respectively, and incorporate the random regularization function into the activation; (2) Replace the fully connected layer in each group of temporal feature extraction modules with a strided convolution unit; the strided convolution unit is used to reduce the time dimension between layers; (3) Adopt the residual structure two in each group of temporal feature extraction modules to realize the connection between units, and use the average pooling function as the dimensionality reduction function of the residual structure; (4) Add a fully connected layer module one with a dimension of T*J at the output end of the spatial feature extraction network, and also add the fully connected layer module two at the output end of the temporal feature extraction network, thereby constructing a targeted spatio-temporal Transformer network model; the fully connected layer module one is used to obtain a three-dimensional human joint point sequence of n frames according to the three-dimensional human joint point pose feature sequence of n frames.

4. The three-dimensional human joint point estimation method based on monocular sequence images according to claim 3, wherein The expression formula of the activation function of the spatial feature extraction module and the temporal feature extraction module is:

5. The method for estimating three-dimensional human joint points based on monocular sequence images according to claim 3, wherein After constructing the targeted spatio-temporal Transformer network model, the targeted spatio-temporal Transformer network model is also trained, and the training process is as follows: Obtain the standard monocular sequence images of multi-frame known joint point coordinate real data, and mix the standard monocular sequence images with the corresponding monocular sequence images to be estimated to obtain a random monocular sequence image set; Use the monocular sequence image set as the sample data to form the data set for model training, and divide the data set into a training set and a validation set; Complete the initialization of the targeted spatio-temporal Transformer network model, use the training set to train the targeted spatio-temporal Transformer network model, use the validation set to verify the training effect of the targeted spatio-temporal Transformer network model, and thus obtain the trained targeted spatio-temporal Transformer network model.

6. The method for estimating three-dimensional human joint points based on monocular sequence images according to claim 5, wherein Each spatial feature extraction module includes: a layer normalization unit one, a multi-head attention unit one, two fully connected layer units one, and a residual structure one; Among them, the feature vectors generated by each spatial feature extraction module generate a three-dimensional pose through the first fully connected layer module, and then calculate the intermediate supervision loss function L J for fast backpropagation, the intermediate supervision loss function L J is set to the average Euclidean distance between the optimized joint points of each spatial feature extraction module and the corresponding joint points in the real data; Take the average Euclidean distance between the 3D human joint point sequences of n frames generated by the first fully connected layer module and the corresponding joint points in the corresponding real data of n frames as the sequence loss function L of the spatial feature extraction network K : In the formula, represents the estimated three-dimensional joint point position of joint i at frame t; represents the true three-dimensional joint point position of joint i at frame t; The total loss L of the spatial feature extraction network S The expression formula is as follows: L S = λ K L K + λ J L J where λ K and λ J are weight factors corresponding to the intermediate supervision loss function and the sequence loss function, respectively.

7. The three-dimensional human joint point estimation method based on monocular sequence images according to claim 6, wherein Each group of temporal feature extraction modules includes: a layer normalization unit two, a multi-head attention unit two, two consecutive one-dimensional convolution units, and the residual structure two; Among them, using the single-frame loss L T to minimize the distance between the three-dimensional joint coordinates X of the intermediate frame output by the temporal feature extraction network and the corresponding true three-dimensional human joint coordinates Y; L T The expression formula of is:

8. The method for estimating three-dimensional human joint points based on monocular sequence images according to claim 7, characterized in that, The expression formula of the total loss L of the targeted spatio-temporal Transformer network model is: L = λ S L S + λ T L T where λ S and λ T are weight factors related to the spatial feature extraction network and the temporal feature extraction network, respectively.

9. A three-dimensional human joint point estimation system based on monocular sequence images, characterized in that, It adopts the three-dimensional human joint point estimation method based on monocular sequence images as described in any one of claims 1 to 8; the three-dimensional human joint point estimation system based on monocular sequence images includes: An image acquisition module, which is used to acquire multiple frames of monocular sequence images containing human joint actions and obtain a two-dimensional human joint sequence in each frame of the monocular sequence image; A preprocessing module, which is used to perform filtering processing on the two-dimensional human joint point sequence and then add position encoding to the two-dimensional human joint point sequence; A spatial feature extraction network, which is used to extract the spatial features of human joint points in each frame of the monocular sequence image, and then obtain a three-dimensional human joint point pose feature sequence of n frames; the spatial feature extraction network includes multiple groups of spatial feature extraction modules; each group of spatial feature extraction modules is used to extract the joint point spatial feature vectors of the corresponding joint set; A temporal feature extraction network, which is used to obtain the three-dimensional human joint point features of the intermediate frame according to the three-dimensional human joint point pose feature sequence of n frames; the temporal feature extraction network includes multiple groups of temporal feature extraction modules; each group of temporal feature extraction modules is used to extract the pose features of human joint points in multiple consecutive frames, and then merge adjacent frames to reduce the frame sequence of the pose of human joint points in multiple frames, and obtain the three-dimensional human joint point coordinates of the target frame through multiple groups of temporal feature extraction modules; and A fully connected layer module two, which is used to obtain the three-dimensional human joint point coordinates of the intermediate frame according to the three-dimensional human joint point features of the intermediate frame.

10. A three-dimensional human joint point estimation device based on monocular sequence images, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the three-dimensional human joint point estimation method based on monocular sequence images according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Traffic police command gesture recognition method based on skeleton joint point sequence

    CN110837778A

  • Three-dimensional human body posture estimation method based on spatio-temporal context feature perception

    CN114241515A