Graphic converter decoder-based category-independent attitude estimation method
By using the graph transformer decoder and Transformer model in pose estimation, combined with the graph convolution network and self-attention mechanism, the problems of multi-category pose estimation and timing information modeling in the prior art are solved, and higher pose estimation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510298490.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Existing pose estimation methods are difficult to deal with multi-category pose estimation, and cannot fully model the geometric structure and timing information between nodes, resulting in poor accuracy and robustness in complex scenarios.
A class-independent pose estimation method based on the graph transformer decoder is adopted to extract spatial features through the graph convolution network, and time information is captured in combination with the Transformer model. The graph transformer decoder is used to integrate spatial and temporal information to accurately locate the location of the correlation node.
It significantly improves the accuracy and robustness of multi-category pose estimation, and can effectively handle pose estimation tasks of different object categories, especially in occlusion or complex backgrounds.
Smart Images

Figure CN120220232A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision estimation methods, and particularly relates to a category-agnostic pose estimation method based on a graph transformer decoder. Background Art
[0002] Human pose estimation, as one of the core tasks in the field of computer vision, aims to identify the precise positions of various joint points of objects or humans in images or videos, such as joints, facial feature points, object parts, etc. It has a wide range of applications, covering multiple fields such as human behavior analysis, intelligent monitoring, augmented reality, virtual reality, medical diagnosis, etc. In recent years, with the rapid development of convolutional neural networks and deep learning, pose estimation technology has made remarkable progress in multiple tasks.
[0003] Existing pose estimation methods can generally be divided into two categories: 2D pose estimation and 3D pose estimation. The goal of 2D pose estimation is to infer the positions of various joints of the human body through two-dimensional images, while 3D pose estimation infers the three-dimensional spatial structure of the human body through the depth information of images or videos. Currently, methods based on convolutional neural networks, especially those that extract features from images through deep convolutional networks, have become the mainstream methods. These methods usually rely on the training of image classification and regression problems, and predict the coordinates of joints through the layer-by-layer processing of images.
[0004] The prior art problems are firstly category-specific problems: Most traditional pose estimation models are designed for specific categories. This means that these models can only handle predefined object categories, such as humans, animals, or vehicles. For example, human pose estimation methods like OpenPose are designed assuming that the input image only contains humans, so their applications are limited. When encountering new object categories, such as new types of animals or objects with special shapes, these methods cannot generalize well, resulting in a significant decline in the performance of pose estimation. Specifically, traditional methods rely on specific annotation data for each category during training, and for unseen categories, the performance of existing models is usually poor. There is also the issue of being unable to handle multi-category pose estimation. To overcome the above problems, Category-Agnostic Pose Estimation (CAPE) emerged. The goal of CAPE is to use a single model for pose estimation of multiple categories, which can not only effectively handle objects of new categories but also requires less supporting data. However, most existing CAPE methods only predict the pose of the target object by matching joint points. They still regard each joint point as an independent individual, ignoring the geometric structure and internal relationships between different joint points. This processing method often leads to lower accuracy in complex pose estimation tasks, especially when facing occlusion or complex backgrounds, and the robustness of traditional methods is poor. Finally, existing CAPE methods usually use Convolutional Neural Networks (CNNs) to extract spatial features and complete pose estimation through simple joint point matching. However, most of these methods ignore the information in the time dimension. When processing video data or time series data, the model is difficult to effectively model the long-distance temporal dependencies between joint points, resulting in poor performance of the model in cases where long-term dependencies are required. Graph Convolutional Networks (GCNs) have shown great potential in processing graph-structured data in recent years. In tasks such as human behavior recognition, skeleton data can be naturally modeled as a graph structure, with each joint point as a node in the graph and the connection relationship between joint points as an edge in the graph. Although GCNs can handle this type of data to a certain extent, due to the fact that they process the structural information of the graph, traditional GCN models have the problem of over-smoothing. Especially when dealing with deep graph convolutional networks, the features of nodes may lose their uniqueness due to over-smoothing, resulting in a decline in recognition accuracy. Summary of the Invention
[0005] The objective of the present invention is to provide a category-agnostic pose estimation method based on a graph transformer decoder, which solves the problem in the prior art that the geometric structure and temporal information between joint points cannot be fully modeled.
[0006] The technical solution adopted by the present invention is a category-agnostic pose estimation method based on a graph transformer decoder, including the following steps: Step 1: Obtain an RGB video and extract image frames; Step 2: Use the OpenPose algorithm to extract skeletons from the image frames and output skeleton data; Step 3: Use the graph convolutional network (GCN) to extract the spatial features of the skeleton data for each frame of the image; Step 4: Use the Transformer model to extract the temporal features of the skeleton sequence in the skeleton data and output enhanced temporal feature vectors; Step 5: Use the graph transformer decoder, combined with the Transformer self-attention mechanism, to accurately locate the position of each joint point and obtain a prediction model; Step 6: Use a loss function to train the prediction model, optimize the neural network, obtain accurate joint point positions, and complete pose estimation independent of categories.
[0007] The features of the present invention also lie in: The specific process of Step 1 is as follows: Step 1.1: RGB video frame extraction; The total number of video frames is N, the video duration is T seconds, and the frame rate is f. Calculate the total number of video frames through the formula and extract each frame of the image from the RGB video to obtain a series of consecutive image frames ; Step 1.2: Image preprocessing; Crop each frame of the image , only retain the human activity area, remove background interference, standardize the size and color of each frame of the image, reduce the influence of light changes and noise on each frame of the image, and annotate each frame of the image through an image annotation tool to ensure that the positions of the human joint points in each frame of the image are accurately marked.
[0008] The specific process of Step 2 is as follows: Step 2.1: Joint point extraction; Use OpenPose to process each frame of the image , identify the human joint points, extract the positions of the human joint points from each frame of the image through a convolutional neural network, and output a set of two-dimensional coordinates , and the two-dimensional coordinates represent the positions of each joint point in the image; Step 2.2: Skeleton connection; Construct the human skeleton structure through the position information of the two-dimensional coordinates between the joint points to form a graph structure, where the edges in the graph structure are the connection relationships between the joint points, and then form a preliminary skeleton sequence; Step 2.3: Skeleton sequence completion; Standardize the preliminary skeleton sequences of different lengths to ensure that the skeleton sequence lengths of each video frame are the same. For the skeleton sequences with insufficient lengths, use the interpolation method for filling to obtain the completed skeleton sequences; Step 2.4: Output the skeleton data of the completed skeleton sequence in each frame of the image as the input for subsequent GCN and Transformer models.
[0009] The specific process of Step 3 is as follows: Step 3.1: Construct and represent the adjacency matrix; represent each frame of skeleton data as a graph, construct the adjacency matrix A, and the adjacency matrix A represents the connection relationship between joint points; the elements of the adjacency matrix A represent node and node whether there is an edge connection between them. If there is, it is 1; otherwise, it is 0. Step 3.2: Graph convolution operation; input the skeleton data into the graph convolution network, and extract the spatial features of each frame of the image through multiple layers of GCN. GCN aggregates the information of adjacent nodes to generate a new representation of each joint point. The output of each layer of graph convolution operation is , where is the node feature of the th layer. Step 3.3: Output the spatial feature representation of each frame of the image through the node features of each layer of graph convolution as the input for subsequent temporal feature extraction.
[0010] The specific process of Step 4 is as follows: Step 4.1: Sequence encoding; merge the spatial feature vector H extracted by GCN and the temporal information to obtain the feature vector Z containing spatial and temporal information. Step 4.2: Cosine position encoding; perform cosine encoding on each frame of the skeleton sequence to generate a vector representing the position of each frame in the time sequence, and obtain the sequential vector. Step 4.3: Add the sequential vector of each frame to the merged feature vector Z to generate the representation vector E with temporal information, and perform a linear transformation on the representation vector E to generate the query vector q, the key-value vector k, and the representative value vector v. Step 4.4: Perform multi-head self-attention operation through the query vector q, the key-value vector k, and the representative value vector v to capture the complex temporal dependence relationships in the skeleton sequence. Step 4.5: Output the enhanced temporal feature vector according to the temporal dependence relationships.
[0011] The specific process of Step 5 is as follows: Step 5.1: In each layer of the decoder, process the input joint point features through the graph convolution network. Through the GCN layer, aggregate the joint point features to accurately locate the joint points. Step 5.2: Perform adaptive interaction between the model at the joint points and the features of the query image. Step 5.3: Perform multi-layer iterative optimization on the joint positions through the graph transformer decoder, and output the optimized predicted joint positions. Step 5.4: Further optimize the predicted positions of each joint through the backpropagation algorithm, and output the accurate joint positions.
[0012] The loss function in Step 6 L is the heatmap loss function and the offset loss function weighted sum, that is: .
[0013] The specific process of training the prediction model using the loss function is as follows: Use the heatmap loss function to supervise the difference between the joint heatmap generated by the model and the real heatmap, so that the joint positions are as close as possible to the real positions; use the offset loss function to supervise the offset prediction of each joint by the network, so that the model further refines the joint positions.
[0014] The beneficial effects of the present invention are: The category-agnostic pose estimation method based on the graph transformer decoder provided by the present invention has been verified on the MP-100 benchmark dataset and significantly improved the joint location accuracy in the 1-shot and 5-shot settings. By using a graph convolutional network to extract spatial features and combining the self-attention mechanism of the Transformer, it can better model spatio-temporal relationships, thereby achieving accurate pose estimation. Specifically, the graph decoder based on GTD can effectively integrate the structural information from the support image and the query image, breaking the dependence on category restrictions in traditional methods and improving the adaptability and accuracy of multi-category pose estimation; by means of the self-attention mechanism of the Transformer, it overcomes the problem in traditional methods that when using convolutional operations to extract temporal features, long temporal dependencies are often unable to be processed. The self-attention mechanism can capture the long-range dependencies between the joints in the sequence, improving the understanding ability of the temporal changes of the joints in the skeleton sequence; the category-agnostic pose estimation CAPE method of the present invention has strong adaptability and can handle the pose estimation tasks of different object categories, such as human bodies, animals, vehicles, etc. By using a graph convolutional network and a Transformer to extract spatial features and temporal features, the model can adapt to the joint relationships of multiple categories and make effective inferences, avoiding the limitations of relying on single-category training in traditional methods; the graph convolutional network GCN and the self-attention mechanism provide strong robustness. Even when there are occlusions or missing joints in the image, the GTD method can still effectively recover and infer the missing joint positions, ensuring the stability and reliability of the model. Through the prior information of the graph structure, the model can infer the occluded or missing joints from adjacent joints and the image context, further enhancing the accuracy of joint location; it has significant advantages in terms of computational efficiency and training scalability. Due to the introduction of the joint design of graph convolution and the Transformer, the network can maintain a low computational cost while ensuring high accuracy; in addition, the method adopts end-to-end training. By optimizing the heatmap loss and the offset loss function, it improves the training efficiency and performance of the model in multi-category pose estimation tasks. Through multiple experiments, it is verified that the category-agnostic pose estimation method based on the graph transformer decoder can handle the pose estimation tasks of more categories with fewer computational resources compared with traditional pose estimation methods, and has good training scalability and computational efficiency. Brief Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of the category-agnostic pose estimation method based on the graph transformer decoder of the present invention. Detailed Embodiments
[0016] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0017] Example 1 The category-agnostic pose estimation method based on the graph transformer decoder proposed in this example is as Figure 1 shown and includes the following steps: Step 1: Obtain an RGB video and extract image frames; Step 2: Use the OpenPose algorithm to extract skeletons from the image frames and output skeleton data; Step 3: Use the graph convolutional network GCN to extract the spatial features of the skeleton data of each frame of image; Step 4: Use the Transformer model to extract the temporal features of the skeleton sequence in the skeleton data and output enhanced temporal feature vectors; Step 5: Use the graph transformer decoder, combined with the Transformer self-attention mechanism, to accurately locate the positions of each joint point and obtain a prediction model; Step 6: Use the loss function to train the prediction model, optimize the neural network, obtain accurate joint point positions, and complete category-agnostic pose estimation.
[0018] Example 2 The category-agnostic pose estimation method based on the graph transformer decoder proposed in this example is as Figure 1 shown and includes the following steps: Step 1: Obtain an RGB video and extract image frames; The specific process is as follows: Step 1.1: RGB video frame extraction; The total number of video frames is N, the video duration is T seconds, and the frame rate is f. Calculate the total number of video frames through the formula and extract each frame of image from the RGB video to obtain a series of consecutive image frames ; Step 1.2: Image preprocessing; Crop each frame of the image frames , only retain the human activity area, remove background interference, standardize the size and color of each frame of image, reduce the influence of illumination changes and noise on each frame of image, and annotate each frame of image through an image annotation tool to ensure that the positions of human joint points in each frame of image are accurately marked; Step 2: Use the OpenPose algorithm to extract skeletons from the image frames and output skeleton data; The specific process is as follows: Step 2.1: Joint point extraction; Use OpenPose to process each frame of image to identify human joint points, extract the positions of human joint points from each frame of image through a convolutional neural network, and output a set of two-dimensional coordinates , and the two-dimensional coordinates Indicates the position of each joint point in the image; Step 2.2, Skeleton connection; Through the two-dimensional coordinates of the joint points, construct the human skeleton structure to form a graph structure, where the edges in the graph structure are the connection relationships between the joint points, thereby forming a preliminary skeleton sequence; Step 2.3, Completing the skeleton sequence; Standardize the preliminary skeleton sequences of different lengths to ensure that the skeleton sequence lengths of each video frame are the same. For the skeleton sequences with insufficient lengths, use the interpolation method to fill them to obtain the completed skeleton sequence; Step 2.4, Output the skeleton data of the completed skeleton sequence in each frame of the image as the input for the subsequent GCN and Transformer models; Step 3, Use the graph convolutional network GCN to extract the spatial features of the skeleton data in each frame of the image; Step 4, Use the Transformer model to extract the temporal features of the skeleton sequence in the skeleton data and output the enhanced temporal feature vectors; Step 5, Use the graph transformer decoder, combined with the Transformer self-attention mechanism, to accurately locate the position of each joint point to obtain the prediction model; Step 6, Use the loss function to train the prediction model, optimize the neural network, obtain the accurate joint point positions, and complete the category-agnostic pose estimation.
[0019] Embodiment 3 The category-agnostic pose estimation method based on the graph transformer decoder proposed in this embodiment, as Figure 1 shown, includes the following steps: Step 1, Obtain the RGB video and extract the image frames; The specific process is as follows: Step 1.1, RGB video frame extraction; The total number of video frames is N, the video duration is T seconds, and the frame rate is f. Calculate the total number of video frames through the formula and extract each frame of image from the RGB video to obtain a series of consecutive image frames ; Step 1.2, Image preprocessing; Crop each frame of the image frame , only keep the human activity area, remove the background interference, standardize the size and color of each frame of the image, reduce the influence of light changes and noise on each frame of the image, and annotate each frame of the image through an image annotation tool to ensure that the positions of the human joint points in each frame of the image are accurately marked; Step 2, Use the OpenPose algorithm to extract the skeleton from the image frame and output the skeleton data; The specific process is as follows: Step 2.1, Joint point extraction; Use OpenPose to process each frame of image to identify human joint points, extract the positions of human joint points from each frame of image through a convolutional neural network, and output a set of two-dimensional coordinates , and the two-dimensional coordinates represent the position of each joint point in the image; Step 2.2, Skeleton connection; Construct the human skeleton structure through the position information of the two-dimensional coordinates between joint points to form a graph structure, where the edges in the graph structure are the connection relationships between joint points, and then form a preliminary skeleton sequence; Step 2.3, Completing the skeleton sequence; Standardize the preliminary skeleton sequences of different lengths to ensure that the skeleton sequence lengths of each video frame are the same. For the skeleton sequences with insufficient lengths, use the interpolation method to fill them to obtain the completed skeleton sequences; Step 2.4, Output the skeleton data of the completed skeleton sequence in each frame of image as the input for the subsequent GCN and Transformer models; Step 3, Use the graph convolutional network GCN to extract the spatial features of the skeleton data in each frame of image; The specific process is as follows: Step 3.1, Construct and represent the adjacency matrix; Represent each frame of skeleton data as a graph, construct the adjacency matrix A, and the adjacency matrix A represents the connection relationship between joint points; The elements of the adjacency matrix A represent whether there is an edge connection between node and node . If there is, it is 1, otherwise it is 0; Step 3.2, Graph convolution operation; Input the skeleton data into the graph convolutional network, and extract the spatial features of each frame of image through multiple layers of GCN. GCN aggregates the information of adjacent nodes to generate a new representation of each joint point. The output of each layer of graph convolution operation is , where is the node feature of the th layer; Step 3.3, Output the spatial feature representation of each frame of image through the node features of each layer of graph convolution as the input for subsequent time feature extraction; Step 4, Use the Transformer model to extract the time features of the skeleton sequence in the skeleton data and output the enhanced time series feature vectors; Step 5, Use the graph transformer decoder, combined with the Transformer self-attention mechanism, to accurately locate the position of each joint point to obtain the prediction model; Step 6, Use the loss function to train the prediction model, optimize the neural network, obtain the accurate joint point positions, and complete the pose estimation independent of categories.
[0020] Example 4 The category - agnostic pose estimation method based on the graph transformer decoder proposed in this example is as Figure 1 shown, and it includes the following steps: Step 1, obtain the RGB video and extract image frames; The specific process is as follows: Step 1.1, RGB video frame extraction; the total number of video frames is N, the video duration is T seconds, and the frame rate is f. Calculate the total number of video frames through the formula and extract each frame of image from the RGB video to obtain a series of consecutive image frames ; Step 1.2, image pre - processing; crop each frame of the image frame , only keep the human activity area, remove background interference, standardize the size and color of each frame of the image, reduce the influence of light changes and noise on each frame of the image, and annotate each frame of the image through an image annotation tool to ensure that the positions of human joint points in each frame of the image are accurately marked; Step 2, use the OpenPose algorithm to extract the skeleton of the image frame and output skeleton data; The specific process is as follows: Step 2.1, joint point extraction; use OpenPose to process each frame of the image to identify human joint points, extract the positions of human joint points from each frame of the image through a convolutional neural network, and output a set of two - dimensional coordinates , and the two - dimensional coordinates represent the positions of each joint point in the image; Step 2.2, skeleton connection; construct the human skeleton structure through the position information of the two - dimensional coordinates between joint points to form a graph structure, where the edges in the graph structure are the connection relationships between joint points, and then form a preliminary skeleton sequence; Step 2.3, skeleton sequence completion; standardize the preliminary skeleton sequences of different lengths to ensure that the skeleton sequence lengths of each video frame are the same. For the skeleton sequences with insufficient lengths, use the interpolation method for filling to obtain the completed skeleton sequences; Step 2.4, output the skeleton data of the completed skeleton sequences in each frame of the image as the input for the subsequent GCN and Transformer models; Step 3, use the graph convolutional network GCN to extract the spatial features of the skeleton data of each frame of the image; The specific process is as follows: Step 3.1. Construct an adjacency matrix and represent it; represent each frame of skeleton data as a graph, construct an adjacency matrix A, and the adjacency matrix A represents the connection relationship between joint points; the elements of the adjacency matrix A represent nodes and nodes to indicate whether there is an edge connection between them. If there is, it is 1; otherwise, it is 0; Step 3.2. Graph convolution operation; input the skeleton data into a graph convolutional network, extract the spatial features of each frame of image through multiple layers of GCN, and GCN aggregates the information of adjacent nodes to generate a new representation of each joint point. The output of each layer of graph convolution operation is where is the node feature of the th layer; Step 3.3. Output the spatial feature representation of each frame of image through the node features of each layer of graph convolution as the input for subsequent time feature extraction; Step 4. Use the Transformer model to extract time features from the skeleton sequence in the skeleton data and output an enhanced temporal feature vector; The specific process is as follows: Step 4.1. Sequence encoding; merge the spatial feature vector H extracted by GCN and the time information to obtain a feature vector Z containing spatial and time information; Step 4.2. Cosine position encoding; perform cosine encoding on each frame of the skeleton sequence to generate a vector representing the position of each frame in the time sequence, and obtain an order vector; Step 4.3. Add the order vector of each frame to the merged feature vector Z to generate a representation vector E with temporal information, and perform a linear transformation on the representation vector E to generate a query vector q, a key-value vector k, and a representative value vector v; Step 4.4. Perform multi-head self-attention operation through the query vector q, the key-value vector k, and the representative value vector v to capture the complex temporal dependence relationships in the skeleton sequence; Step 4.5. Output an enhanced temporal feature vector according to the temporal dependence relationships; Step 5. Use a graph transformer decoder, combined with the Transformer self-attention mechanism, to accurately locate the position of each joint point to obtain a prediction model; Step 6. Use a loss function to train the prediction model, optimize the neural network, obtain the accurate joint point positions, and complete class-agnostic pose estimation.
[0021] Example 5 The class-agnostic pose estimation method based on a graph transformer decoder proposed in this example, as Figure 1 shown, includes the following steps: Step 1: Obtain the RGB video and extract image frames; The specific process is as follows: Step 1.1: RGB video frame extraction; The total number of video frames is N, the video duration is T seconds, and the frame rate is f. Calculate the total number of video frames through the formula Extract each frame of image from the RGB video to obtain a series of consecutive image frames ; Step 1.2: Image preprocessing; Crop each frame of the image frame Only retain the human activity area, remove background interference, standardize the size and color of each frame of the image, reduce the influence of light changes and noise on each frame of the image, and annotate each frame of the image through an image annotation tool to ensure that the positions of human joint points in each frame of the image are accurately marked; Step 2: Use the OpenPose algorithm to extract the skeleton of the image frame and output skeleton data; The specific process is as follows: Step 2.1: Joint point extraction; Use OpenPose to process each frame of the image Identify the human joint points, extract the positions of human joint points from each frame of the image through a convolutional neural network, and output a set of two-dimensional coordinates , The two-dimensional coordinates Represent the position of each joint point in the image; Step 2.2: Skeleton connection; Construct the human skeleton structure through the position information of the two-dimensional coordinates Between joint points to form a graph structure, where the edges in the graph structure are the connection relationships between joint points, thereby forming a preliminary skeleton sequence; Step 2.3: Skeleton sequence completion; Standardize the preliminary skeleton sequences of different lengths to ensure that the skeleton sequence lengths of each video frame are the same. For the skeleton sequences with insufficient lengths, use the interpolation method for filling to obtain the completed skeleton sequence; Step 2.4: Output the skeleton data of the completed skeleton sequence in each frame of the image as the input for the subsequent GCN and Transformer models; Step 3: Use the graph convolutional network GCN to extract the spatial features of the skeleton data of each frame of the image; The specific process is as follows: Step 3.1: Construct and represent the adjacency matrix; Represent each frame of skeleton data as a graph, construct the adjacency matrix A, and the adjacency matrix A represents the connection relationship between joint points; The element Of the adjacency matrix A represents the node And the node Whether there is an edge connection between them. If there is, it is 1, otherwise it is 0; Step 3.2, Graph convolution operation; Input the skeleton data into the graph convolutional network, and extract the spatial features of each frame of image through multiple layers of GCN. GCN aggregates the adjacent node information to generate a new representation of each joint point. The output of each layer of graph convolution operation is , where is the node feature of the th layer; Step 3.3, Output the spatial feature representation of each frame of image through the node features of each layer of graph convolution, which is used as the input for subsequent time feature extraction; Step 4, Use the Transformer model to extract the time features of the skeleton sequence in the skeleton data, and output the enhanced temporal feature vector; The specific process is as follows: Step 4.1, Sequence encoding; Merge the spatial feature vector H extracted by GCN and the time information to obtain the feature vector Z containing spatial and time information; Step 4.2, Cosine position encoding; Perform cosine encoding on each frame of the skeleton sequence to generate a vector representing the position of each frame in the time sequence, and obtain the sequential vector; Step 4.3, Add the sequential vector of each frame to the merged feature vector Z to generate the representation vector E with temporal information, and perform a linear transformation on the representation vector E to generate the query vector q, the key-value vector k, and the representative value vector v; Step 4.4, Perform multi-head self-attention operation through the query vector q, the key-value vector k, and the representative value vector v to capture the complex temporal dependencies in the skeleton sequence; Step 4.5, Output the enhanced temporal feature vector according to the temporal dependencies; Step 5, Use the graph transformer decoder, combined with the Transformer self-attention mechanism, to accurately locate the position of each joint point to obtain the prediction model; The specific process is as follows: Step 5.1, In each layer of the decoder, process the input joint point features through the graph convolutional network. Through the GCN layer, aggregate the joint point features to accurately locate the joint points; Step 5.2, Perform adaptive interaction between the model among the joint points and the features of the query image; Step 5.3, Perform multi-layer iterative optimization on the joint point positions through the graph transformer decoder, and output the optimized predicted joint point positions, Step 5.4, Re-optimize the predicted positions of each joint point through the backpropagation algorithm, and output the accurate joint point positions; Step 6, Use the loss function to train the prediction model, optimize the neural network, obtain the accurate joint point positions, and complete the class-agnostic pose estimation; Loss function L is the heatmap loss function and the weighted sum of the offset loss function , that is: ; The specific process of training the prediction model using the loss function is as follows: Use the heatmap loss function to supervise the difference between the joint heatmap generated by the model and the true heatmap, so that the joint positions are as close as possible to the true positions; Use the offset loss function to supervise the offset prediction of each joint by the network, so that the model further refines the joint positions.
[0022] Example 6 Take 70% of the video data in a set of RGB videos containing different objects as the training set, and 30% of the video data as the test set. Each frame of image is processed by OpenPose to obtain the two-dimensional coordinates of 25 joints, forming a complete skeleton sequence; Step 1, preprocessing; First, extract each frame of image from the RGB video, and adjust the image size to 256x192 pixels. Crop each frame of image, only retaining the area related to the object's activity; Step 2, skeleton data extraction; Extract the two-dimensional coordinates of 25 joints in each frame of image through OpenPose to construct a skeleton graph. The skeleton data of each frame is arranged in chronological order to form a skeleton sequence; Step 3, GCN feature extraction; Use the graph convolutional network GCN to extract the spatial features of each frame of image and capture the spatial relationship between joints; Step 4, temporal information modeling; Combine the spatial features of each frame with the sequence vector through Transformer to capture the temporal dependence relationship in the skeleton sequence; Step 5, graph transformer decoder; In the graph transformer decoder, use graph convolutional operations to further strengthen the relationship between joints, thereby improving the joint localization accuracy; Step 6, training and optimization; Optimize the neural network through the cross-entropy loss and the offset loss function to finally obtain accurate joint position predictions; Through this series of steps, accurately predict the joint positions of objects in the query image and complete the pose estimation task that is independent of the category.
Claims
1. A class-independent pose estimation method based on a graph transformer decoder, characterized in that The following steps are involved: Step 1, obtain RGB video and extract image frames; Step 2: Use the OpenPose algorithm to extract the skeleton of the image frame and output the skeleton data; Step 3: Use graph convolutional network GCN to extract the spatial features of skeleton data of each frame of image; Step 4: Use the Transformer model to extract temporal features from the skeleton sequence in the skeleton data and output an enhanced temporal feature vector; Step 5: Use the graph transformer decoder and the Transformer self-attention mechanism to accurately locate the position of each joint point and obtain the prediction model; Step 6: Use the loss function to train the prediction model, optimize the neural network, obtain accurate joint point positions, and complete category-independent posture estimation.
2. The class-independent pose estimation method based on a graph transformer decoder according to claim 1, characterized in that: The specific process of step 1 is as follows: Step 1.1, RGB video frame extraction; the total number of video frames is N, the video length is T seconds, the frame rate is f, and the formula Calculate the total number of video frames, extract each frame from the RGB video, and get a series of continuous image frames ; Step 1.2: Image preprocessing: Each frame of the image is cropped to retain only the area where the human body is active, remove background interference, standardize the size and color of each frame of the image, reduce the impact of lighting changes and noise on each frame of the image, and annotate each frame of the image through image annotation tools to ensure that the position of the human body joints in each frame of the image is accurately marked.
3. The class-independent pose estimation method based on a graph transformer decoder according to claim 1, characterized in that: The specific process of step 2 is: Step 2.1, joint point extraction; use OpenPose for each frame image Processing is performed to identify the human joints, and the positions of the human joints are extracted from each frame of the image through a convolutional neural network, outputting a set of two-dimensional coordinates , two-dimensional coordinates Indicates the position of each joint point in the image; Step 2.2, skeleton connection; Through the two-dimensional coordinates between joint points The position information of the human body is used to construct a human skeleton structure, forming a graph structure, in which the edges in the graph structure are the connection relationships between the joints, thereby forming a preliminary skeleton sequence; Step 2.3, skeleton sequence completion: standardize the preliminary skeleton sequences of different lengths to ensure that the skeleton sequence length of each video frame is consistent. For skeleton sequences of insufficient length, use interpolation to fill them and obtain the completed skeleton sequence; Step 2.4: Output the skeleton data of the completed skeleton sequence in each frame image as the input of the subsequent GCN and Transformer models.
4. The class-independent pose estimation method based on a graph transformer decoder according to claim 1, characterized in that: The specific process of step 3 is as follows: Step 3.1, construct an adjacency matrix and represent it; represent each frame of skeleton data as a graph, construct an adjacency matrix A, and the adjacency matrix A represents the connection relationship between joint points; the elements of the adjacency matrix A are Representation Node and nodes Is there an edge connection between them? If yes, it is 1, otherwise it is 0; Step 3.2, graph convolution operation: Input the skeleton data into the graph convolution network, extract the spatial features of each frame image through multi-layer GCN, GCN aggregates the adjacent node information, and generates a new representation of each joint point. The output of each layer of graph convolution operation is ,in For the Node characteristics of the layer; In step 3.3, the spatial feature representation of each frame image is output through the node features of each layer of graph convolution, which serves as the input for subsequent temporal feature extraction.
5. The class-independent pose estimation method based on a graph transformer decoder according to claim 1, characterized in that: The specific process of step 4 is as follows: Step 4.1, sequence encoding: merge the spatial feature vector H extracted by GCN with the temporal information to obtain a feature vector Z containing spatial and temporal information; Step 4.2, cosine position encoding: cosine encoding is performed on each frame of the skeleton sequence to generate a vector representing the position of each frame in the time series, and a sequence vector is obtained; Step 4.3, add the sequence vector of each frame to the merged feature vector Z to generate a representation vector E with time series information, perform a linear transformation on the representation vector E to generate a query vector q, a key value vector k and a representative value vector v; Step 4.4: Perform multi-head self-attention operation on query vector q, key value vector k and representative value vector v to capture the complex temporal dependencies in the skeleton sequence. Step 4.5: Output the enhanced timing feature vector according to the timing dependency.
6. The class-independent pose estimation method based on a graph transformer decoder according to claim 1, characterized in that: The specific process of step 5 is as follows: Step 5.1: In each layer of the decoder, the input joint point features are processed by the graph convolutional network, and the joint point features are aggregated through the GCN layer to accurately locate the joint points; Step 5.2, adaptively interact the model with the features of the query image between the joint points; Step 5.3: Perform multi-layer iterative optimization on the joint point positions through the graph transformer decoder, and output the optimized predicted joint point positions. Step 5.4: Optimize the predicted position of each joint point again through the back propagation algorithm and output the precise joint point position.
7. The class-independent pose estimation method based on a graph transformer decoder according to claim 1, characterized in that: The loss function described in step 6 L is the heat map loss function And the offset loss function The weighted sum of is: .
8. The class-independent pose estimation method based on a graph transformer decoder according to claim 7, characterized in that: The specific process of using the loss function to train the prediction model is as follows: using the heat map loss function The difference between the joint point heat map generated by the supervised model and the real heat map is used to make the joint point position as close to the real position as possible; Using the offset loss function The supervised network predicts the offset of each joint point, allowing the model to further refine the joint point positions.