A Pose Estimation Method Based on Self-Supervised Learning

Through the pose estimation method based on self-supervised learning, the self-supervised training visual backbone model and partial segmentation network are used to solve the problems of insufficient labeled data and insufficient feature fine-graining in the pose estimation and partial overall relationship discovery tasks in the prior art, and a more robust and interpretable computer vision model is achieved.

CN115661246BActive Publication Date: 2025-06-24SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211312697.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-06-24
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Existing computer vision models are difficult to generalize in strong adversarial scenarios, especially in pose estimation and partial overall relationship discovery tasks. The lack of sufficient labeled data and fine-grained features leads to insufficient robustness and interpretation.

Method used

Using a pose estimation method based on self-supervised learning, the visual backbone model is pre-trained through the disclosed image data set, and self-supervised training is performed in combination with part of the overall relationship constraints to obtain a partial segmentation network and a key point estimator, reducing the complexity of data labeling and extracting fine-grained features.

Benefits of technology

It realizes effective pose estimation and partial segmentation under the condition of a small number of data samples, which improves the robustness and interpretability of the model and reduces the complexity of data labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661246B_ABST
    Figure CN115661246B_ABST
Patent Text Reader

Abstract

The present invention discloses a pose estimation method based on self-supervised learning. First, a visual backbone model is pre-trained by a self-supervised learning algorithm based on a contrastive method; then a part segmentation network is obtained through self-supervised training based on part-whole relationship constraints; then a key point estimator is obtained through regression learning training; after that, the target image is sequentially passed through the visual backbone model, the part segmentation network, and the key point estimator to obtain a key point map and a calibrated perspective feature map, and then combined with a depth map to extract the calibrated perspective feature and depth value of the key point. According to the depth value and the key point coordinates, the three-dimensional coordinates of the key point in the camera coordinate system are obtained, and then a similarity transformation is performed between the camera coordinate system and the world coordinate system to obtain the pose estimation result. The present invention can extract image features suitable for fine-grained downstream tasks, and can directly provide key points and calibrated perspective features, effectively reducing the complexity and workload of data annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a pose estimation method based on self-supervised learning. Background Art

[0002] Pose estimation and part-whole relationship discovery are both long-standing challenges in computer vision and important processes for artificial intelligence to understand the real 3D world. The traditional computer vision field mainly focuses on visual understanding on 2D images, such as tasks like image recognition, object detection, semantic segmentation, etc. With the development of fields such as autonomous driving and robotics, the understanding of the real 3D world by artificial intelligence has gradually received attention. Some researchers also focus on generating RGB-D images with depth information or point cloud information through sensors capable of acquiring the real 3D world, such as depth cameras, LiDAR, etc., for further use in artificial intelligence's understanding of the real 3D world. However, it has been found that humans can often obtain an accurate cognitive understanding of the real 3D world only through 2D images and their 3D priors about the real world. Different from most artificial intelligence methods, this ability of humans is highly generalizable. That is to say, even if humans have not seen a certain category of objects, they can still extract a 3D understanding of the target object through 2D images. This understanding can be interpreted as a bottom-up process when humans perceive the world, by comparing the parts in the target object with the parts in known objects, thus forming a cognitive understanding of the target object in a combinatorial way. This idea has inspired a class of methods in computer vision, called combinatorial methods. Combinatorial methods mostly rely on the features of parts of the image (at the pixel level or block level), and by introducing combinatorial models to model the relationships between pixels in the image, thus forming an abstract concept understanding of the target or an understanding of the part-whole relationship.

[0003] Traditional machine learning is often limited by the form of its input data. For example, traditional computer vision methods need to use manually designed feature extractors to convert image data into the input of the machine learning subsystem; while deep learning is a representation learning method based on multi-level representations, by combining simple but non-linear modules, converting features at a certain level into higher-order and more abstract features. From this perspective, deep learning methods are also an implicit combinatorial method, by learning different levels of features for further use in downstream tasks.

[0004] Although computer vision benefits from deep learning, it is still limited by the security and robustness that need to be considered in real-world deployments. Research has found that in strongly adversarial scenarios such as partial occlusion, computer vision models may not be well generalized, leading to possible fatal consequences. The current vision models have the following defects: (1) The annotation of the relationship between the part and the whole of the target or pose estimation requires a more complex annotation process compared to conventional vision tasks such as classification. For example, after introducing the 3D CAD model of the target, manual adjustment of the 3D CAD model is required to align the target in the image. For specific sensitive targets, it is even more difficult to obtain a CAD model for annotation. Therefore, there is a lack of sufficient relevant annotation data, and there is a problem of insufficient dataset annotation. (2) The current deep learning backbone models of computer vision are all pre-trained network models based on image labels as supervision signals. As a coarse-grained supervision signal, the backbone network model obtained by corresponding pre-training is difficult to perform some more fine-grained tasks downstream, such as the discovery of the relationship between the part and the whole of the target and pose estimation, which both require features with high fine-grainedness and distinctiveness.

[0005] Therefore, the present invention hopes to construct a robust and interpretable computer vision model to cope with these strongly adversarial scenarios; hopes to guide the model to discover the relationship between the part and the whole, so that it can obtain visual understanding similar to human cognition of things, and obtain a more robust model intuitively; hopes to complete further image understanding tasks, such as pose estimation, through the learning-based discovery of the relationship between the part and the whole of the target. Summary of the Invention

[0006] The present invention provides a self-supervised learning-based pose estimation method, which can extract image pixel-level features applicable to fine-grained downstream tasks such as pose estimation and part segmentation, and can reflect its interpretability through the part segmentation results. At the same time, it can directly provide key points and calibration perspective features for the pose estimation task, reduce the complexity and workload of data annotation, obtain effective pose estimation, and better complete the image understanding task.

[0007] The technical solution of the present invention is as follows:

[0008] A self-supervised learning-based pose estimation method includes the following steps:

[0009] S1. Using a publicly available image dataset, pre-train a vision backbone model based on a self-supervised learning algorithm of a contrastive method, and the vision backbone model outputs image features;

[0010] S2. Using the image features, self-supervised training based on the part-whole relationship constraint to obtain a part segmentation network, and the part segmentation network outputs a part response map;

[0011] S3. Using the images with key points marked and their corresponding calibrated perspective features as learning targets, taking the feature points of some response maps as inputs, and then training a network through regression learning to obtain a key point estimator. The key point estimator outputs the key point map and the calibrated perspective feature map corresponding to the image.

[0012] S4. Input the target image into the trained visual backbone model to obtain the image characteristics of the target image. Then, input the image characteristics of the target image into the trained partial segmentation network to obtain the partial response map of the target image. After that, input the partial response map of the target image into the trained key point estimator to obtain the key point map and the calibrated perspective feature map of the target image.

[0013] S5. Obtain the depth map of the target image, use the non-maximum suppression algorithm to filter out multiple key points from the key point map of the target image, extract the coordinates of multiple key points, and then use the key point coordinates to extract the calibrated perspective features q i and depth values d i ;

[0014] S6. Combine the depth value d i and the key point coordinates to obtain the three-dimensional coordinates p i of multiple key points in the camera coordinate system. Represent the transformation relationship between the camera coordinate system and the world coordinate system as a similarity transformation, which is parameterized by a scalar s ∈ R + , a rotation matrix R ∈ SO(3), and a translation t, and obtain it by minimizing the following objective function:

[0015]

[0016] where w i ∈ [0, 1] represents the confidence score, and N1 represents the number of key points;

[0017] s ★ , R ★ , t ★ are the optimal parameterizations obtained after minimizing the objective function, and s ★ , R ★ , t ★ are the pose estimation results of the target image.

[0018] The present invention forms training samples using a publicly available large-scale image dataset, and then pre-trains a visual backbone model based on a self-supervised learning algorithm of a contrastive method. The visual backbone model mainly provides image features for a key point estimator and a part segmentation network of downstream tasks. Among them, the part segmentation network performs further self-supervised learning training on an unannotated dataset through part-whole relationship constraints, and finally obtains a part-whole relationship discovery model that can output part segmentation, and its interpretability is reflected through the part segmentation results. The key point estimator is trained through regression learning based on the above-mentioned trained visual backbone model and part segmentation network. The key point estimator can directly provide key points and calibration perspective features for the pose estimation task, reducing the complexity and workload of data annotation. After obtaining the visual backbone model, the part segmentation network, and the key point estimator, the target image is predicted. First, the target image is sequentially passed through the visual backbone model, the part segmentation network, and the key point estimator to obtain a key point map and a calibration perspective feature map, and then combined with the depth map of the image itself, the calibration perspective features and depth values corresponding to the positions of multiple key points on the calibration perspective feature map and the depth map are extracted. According to the depth values and the key point coordinates, the three-dimensional coordinates of multiple key points in the camera coordinate system are obtained, and then a similarity transformation between the camera coordinate system and the world coordinate system is performed to obtain the pose estimation result of the target image.

[0019] Further, the image dataset used in step S1 includes ImageNet-1K or ImageNet-21K.

[0020] Further, the specific process of pre-training the visual backbone model based on the self-supervised learning algorithm of the contrastive method in step S1 is as follows:

[0021] Introduce a proxy task at the pixel level. The proxy task involves two parts, one is a pixel propagation module, and the other is an asymmetric structure design. One branch of the structure design generates a normal feature map, and the other branch combines the pixel propagation module. The asymmetric structure design only requires the consistency of positive sample pairs and does not require careful debugging of negative sample pairs.

[0022] For each pixel feature, the vector after its smooth transformation is calculated through the pixel propagation module. This vector is obtained by propagating all pixel features on the same image Ω to the current pixel feature, as shown in the following formula:

[0023] y i =Σ j∈Ω s(x i ,x j )·g(x j )

[0024] In the formula, x iis the i-th pixel feature, x j is the j-th pixel feature, i is the vector after the i-th pixel feature is smoothed;

[0025] where s(·,·) is a similarity function, defined as follows:

[0026] s(x i , x j ) = (max(cos(x i , x j ), 0)) γ

[0027] where γ is a sharpness index that controls the similarity function and is default set to 2;

[0028] g(·) is a transformation function, instantiated through several linear layers containing batch normalization and rectified linear unit functions;

[0029] In the asymmetric structure design, there are two different encoders: one is the propagation encoder loaded with the pixel propagation module for post-processing to generate smoothed features, and the other is the momentum encoder without the pixel propagation module; both enhanced perspectives are fed into the two encoders, and the features generated by different encoders are encouraged to be consistent:

[0030]

[0031] where, represents the pixel propagation loss, and i and j are two positive pixel pairs based on the threshold assignment rule under the enhanced perspective; x i ’ is the i-th pixel feature enhanced by the momentum encoder, x′ j is the j-th pixel feature enhanced by the momentum encoder, y j is the vector after the j-th pixel feature is smoothed; this loss is calculated on average for each image of all positive sample pairs, and then averaged for each batch of data to represent learning.

[0032] Furthermore, the specific process of obtaining the part segmentation network through self-supervised training based on the part-whole relationship constraint in step S2 is as follows:

[0033] Adopt the self-supervised constraints of geometric concentration loss, equivalence loss, semantic consistency loss, and foreground-background discrimination loss for self-supervised learning training, and finally obtain a part-whole relationship discovery model that can output part segmentation, that is, the part segmentation network.

[0034] Furthermore, the definition process of the geometric concentration loss is as follows:

[0035] Pixels of the same target part are more spatially concentrated on the same image and form a connected component without occlusion or multiple instances. Based on this, geometric concentration is an important property for part segmentation. Therefore, a loss term is used to encourage the spatial concentration of the same part;

[0036] For the part center of a certain part k on axis u There is:

[0037]

[0038] For the part center of a certain part k on axis v There is:

[0039]

[0040] Where is a normalization term used to transform the part response map into a spatial probability distribution function. Then, the geometric concentration loss is defined as:

[0041]

[0042] Moreover, this loss is differentiable. This loss function encourages each part to form geometric concentration and attempts to minimize the variance of the spatial probability distribution function R(k, u, v) / z k of.

[0043] Furthermore, the definition process of the equivalence loss is as follows:

[0044] For each training image, use a random spatial transformation T s (·) and appearance variation T a (·). For the input image and the transformed image, obtain the corresponding part response maps Z and Z' respectively. According to these two part response maps, calculate the part centers and respectively. Then, the equivalence loss is defined as:

[0045]

[0046] Where D KL (·) is the KL divergence distance, is the balance coefficient;

[0047] The first term in the above formula corresponds to the equivalence constraint of part segmentation, and the second term in the above formula corresponds to the equivalence constraint of part centers.

[0048] Furthermore, the definition process of the semantic consistency loss is as follows:

[0049] The intermediate layer information of the neural network has the semantic information of the target and parts. Therefore, a loss function that constrains semantic consistency is used to utilize the hidden information contained in the features of the neural network pre-trained on ImageNet, and find representative feature clusters from the given pre-trained classification features to correspond to different part segmentations.

[0050] Formally, given C-dimensional classification features It is hoped to find K representative partial feature vectors d k ∈R D , k ∈ {1, 2, …, K}. At the same time, it is hoped to learn the partial segmentation results and the corresponding dictionary of partial feature vectors, so that the classification features are close to d k Then there is the following semantic consistency loss:

[0051]

[0052] where V(u, v) is the feature vector at spatial position (u, v). Through the constraint of the semantic consistency loss, the partial basis vectors w k composing the semantic dictionary {w k} can be learned, which ensures cross-instance semantic consistency, thus ensuring that the same partial response corresponds to similar semantic features in the pre-trained classification feature space;

[0053] When training the semantic consistency loss, there may be a situation where different partial bases correspond to similar feature vectors. Therefore, an additional orthogonality constraint on the partial basis vectors w k is introduced to distinguish different basis vectors from each other. Let represent the normalized partial basis vectors for each row Formally, the orthogonality constraint is regarded as a loss function acting on :

[0054]

[0055] where is the F-norm, and II K is the identity matrix of size K × K; through this constraint, the mutual correlation of different basis vectors is minimized to obtain more accurate partial basis vectors, thereby obtaining better partial segmentation results.

[0056] Furthermore, the definition process of the foreground-background discrimination loss is as follows:

[0057] Using the saliency detection model pre-trained on other training sets to generate a saliency map, and using the saliency map, the background loss function is:

[0058]

[0059] where D ∈ [0, 1] H×W is the saliency map, H represents the number of rows of the matrix, W represents the number of columns of the matrix, D(u, v) is the saliency value of the saliency map at the spatial position (u, v), and R(0, u, v) is the segmentation result of the background.

[0060] Furthermore, by using multiple loss functions to train the partial segmentation network and the semantic partial basis, the obtained objective function is a linear combination of multiple loss functions:

[0061]

[0062] In the formula, λ con , λ eqv , λ sc , λ bg are the balance coefficients of the corresponding loss functions respectively.

[0063] Furthermore, the specific process of training the key point estimator through regression learning in step S3 is as follows:

[0064] Using the partial response map Z(k) H×W obtained by the segmentation network, where k takes 1, 2,..., K. For each partial response map, a series of feature points are extracted using the non-maximum suppression method, and this series of feature points are used as the input of the key point estimator. The key point estimator is a multi-layer perceptron, and a heat map is also obtained as the output. The heat map is processed using non-maximum suppression to obtain a series of key points

[0065] Denote the normalized labeled key point as kp i =(a i , b i ), a i ∈ [0, 1], b i ∈ [0, 1], and the estimated key point is Then there is a regression loss:

[0066]

[0067] The beneficial effects of the present invention are as follows:

[0068] 1) For the current target part-whole relationship discovery algorithm, a pre-trained model based on supervised learning is usually used to obtain image features, and the features extracted by this supervised learning are usually coarse-grained supervision signals based on categories, which are not sufficient to meet the needs of the target part-whole relationship discovery algorithm. However, the present invention uses a self-supervised learning algorithm based on a contrastive method to pre-train a visual backbone model, which can extract image pixel-level features suitable for fine-grained downstream tasks such as pose estimation and partial segmentation, and can meet the needs of the target part-whole relationship discovery algorithm.

[0069] 2) For the current target pose estimation algorithm, a complex manual labeling process is usually required. However, the present invention introduces a self-supervised visual backbone model and a partial segmentation network, which can fine-tune and train a key point estimator with a small amount of data sample labeling, and use the key point estimator to directly provide key points and calibration view features for the pose estimation task, which can effectively reduce the workload of manual labeling and the complexity of data labeling, obtain effective pose estimation, and better complete the image understanding task. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 The present invention is a flowchart of a posture estimation method based on self-supervised learning. DETAILED DESCRIPTION

[0071] The drawings are only for illustrative purposes and cannot be construed as limiting the present invention. To better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged, or reduced, and do not represent the size of the actual product. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the drawings. The positional relationships described in the drawings are only for illustrative purposes and cannot be construed as limiting the present invention.

[0072] Embodiment 1:

[0073] Existing part-whole relationship discovery algorithm research can be divided into three categories: capsule network-based methods, combinational model-based methods, and part-based methods. They all use different methods based on image features to discover part-whole relationship concepts. The method used in this invention is a self-supervised method based on pixel-level features generated by a self-supervised method, which is different from the previous part-whole relationship discovery method based on learning with some kind of supervisory signal.

[0074] Part-based methods are often used in fine-grained object recognition. In fine-grained object recognition, since objects of the same category often have a common appearance and only differ in local locations, the paradigm of locating the parts of the object and extracting the landmark information of the parts often plays an important role in the task of fine-grained object recognition.

[0075] Self-supervised learning is a category of algorithms relative to supervised learning. Self-supervised learning does not require data to have labeled information. Instead, it optimizes predefined proxy tasks on a large amount of unlabeled data and uses the information of the data itself as a supervisory signal to learn meaningful representations for downstream tasks. Since it does not require labeled data, self-supervised learning can use more data for training, which is also an advantage of self-supervised learning over supervised learning. Self-supervised learning methods can be divided into two categories according to the form of their proxy tasks:

[0076] (1) Contrastive methods: These methods obtain positive or negative samples by enhancing the data itself or random sampling; then, through a loss function, they minimize the similarity distance between positive samples and maximize the similarity distance between negative samples. For example, methods such as MoCo (Momentum Contrastive) in the field of computer vision obtain positive sample pairs by data augmentation of images, randomly sample other images in the dataset as negative sample pairs, and learn semantic representations for use in downstream tasks such as image classification, object detection, and semantic segmentation.

[0077] (2) Generative methods: Utilize distribution information such as the context of the data itself to generate a distribution that completes the proxy task, thereby achieving the purpose of extracting information from unlabeled data. Common proxy tasks include: instance discrimination, clustering discrimination, image reconstruction, cloze test, etc. For example, the classic model BERT (Bidirectional Encoder Representation from Transformers) in the field of natural language processing randomly masks words in a sentence and uses the cloze test as a proxy task to enable the model to learn the context information between words.

[0078] Based on this, as Figure 1 shown, the present invention proposes a self-supervised learning pose estimation method based on a contrastive method. The process of discovering the part-whole relationship is not completely consistent with the existing three methods, and its interpretability can be reflected through the partial segmentation results.

[0079] The specific process is as follows:

[0080] S1. Use a publicly available image dataset and pre-train a visual backbone model based on a self-supervised learning algorithm of the contrastive method. The visual backbone model outputs image features.

[0081] S2. Use the image features and self-supervised train a partial segmentation network based on part-whole relationship constraints. The partial segmentation network outputs partial response maps.

[0082] S3. Use the image with labeled key points and its corresponding calibrated perspective features as the learning target, take the feature points of the partial response map as the input, and then train a network through regression learning as a key point estimator. The key point estimator outputs the key point map and calibrated perspective feature map corresponding to the image.

[0083] S4. Input the target image into the trained visual backbone model to obtain the image characteristics of the target image. Then, input the image characteristics of the target image into the trained part segmentation network to obtain the part response map of the target image. After that, input the part response map of the target image into the trained key point estimator to obtain the key point map and the calibrated perspective feature map of the target image;

[0084] S5. Obtain the depth map of the target image, and use the non-maximum suppression algorithm to screen out multiple key points from the key point map of the target image, extract the coordinates of multiple key points, and then use the key point coordinates to extract the calibrated perspective feature q i and the depth value d i ;

[0085] S6. Combine the depth value d i and the key point coordinates to obtain the three-dimensional coordinates p i of multiple key points in the camera coordinate system. Represent the transformation relationship between the camera coordinate system and the world coordinate system as a similarity transformation, which is parameterized by a scalar s ∈ R + , a rotation matrix R ∈ SO(3), and a translation t, and is obtained by minimizing the following objective function:

[0086]

[0087] where w i ∈ [0,1] represents the confidence score, and N1 represents the number of key points;

[0088] s ★ , R ★ , t ★ are the optimal parameterizations obtained after minimizing the objective function, and s ★ , R ★ , t ★ are the pose estimation results of the target image.

[0089] The present invention forms training samples using publicly available large-scale publicly available image datasets, and then pre-trains a visual backbone model based on a self-supervised learning algorithm of a contrastive method. The visual backbone model mainly provides image features for a key point estimator and a part segmentation network for downstream tasks; among them, the part segmentation network is further self-supervised learning trained on an unlabeled dataset through part-whole relationship constraints, and finally a part-whole relationship discovery model that can output part segmentation is obtained, and its interpretability is reflected through the part segmentation results; while the key point estimator is obtained through regression learning based on the above-mentioned trained visual backbone model and part segmentation network. The key point estimator can directly provide key points and calibrated perspective features for the pose estimation task, reducing the data annotation complexity and workload. After obtaining the visual backbone model, the part segmentation network and the key point estimator, the target image is predicted. First, the target image is sequentially passed through the visual backbone model, the part segmentation network and the key point estimator to obtain a key point map and a calibrated perspective feature map, and then combined with the depth map of the image itself, the calibrated perspective features and depth values corresponding to the positions of multiple key points on the calibrated perspective feature map and the depth map are extracted. According to the depth values and the key point coordinates, the three-dimensional coordinates of multiple key points in the camera coordinate system are obtained, and then a similarity transformation between the camera coordinate system and the world coordinate system is performed to obtain the pose estimation result of the target image.

[0090] In step S1 of this embodiment, publicly available large-scale image datasets such as ImageNet-1K, ImageNet-21K, etc. are used as the training set, and a visual backbone model is pre-trained based on a self-supervised learning algorithm of a contrastive method. The specific process is as follows:

[0091] Introduce a pixel-level proxy task - pixel-to-propagation for propagation, which can simultaneously extract the spatial sensitivity and spatial smoothness of the representation during the self-supervised representation learning process; this proxy task mainly involves two parts, one is a pixel propagation module, and the other is an asymmetric structure design. One branch of the structure design generates a normal feature map, and the other branch combines the pixel propagation module; the asymmetric structure design only requires the consistency of positive sample pairs and does not require careful debugging of negative sample pairs;

[0092] For each pixel feature, the smoothly transformed vector is calculated through the pixel propagation module, and this vector is obtained by propagating all pixel features on the same image Ω to the current pixel feature, as shown in the following formula:

[0093] y i =Σ j∈Ω s(x i ,xj)·g(x j )

[0094] Wherein, x i is the i-th pixel feature, and x j is the j-th pixel feature, i is the vector after the i-th pixel feature is smoothed;

[0095] Where s(·,·) is a similarity function, defined as follows:

[0096] s(x i , x j ) = (max(cos(x i , x j ), 0)) γ

[0097] Where γ is a sharpness index that controls the similarity function and is default set to 2;

[0098] g(·) is a transformation function instantiated through several linear layers containing batch normalization and rectified linear units;

[0099] In the asymmetric structure design, there are two different encoders: one is a propagation encoder loaded with a pixel propagation module for post-processing to generate smoothed features, and the other is a momentum encoder without a pixel propagation module; the two enhanced views are fed into both encoders, and the features generated by different encoders are encouraged to be consistent:

[0100]

[0101] Wherein, represents the pixel propagation loss, and i and j are two pairs of positive pixels based on the threshold assignment rule under the enhanced view; x i ’ is the i-th pixel feature enhanced by the momentum encoder, and x′ j is the j-th pixel feature enhanced by the momentum encoder, y j is the vector after the j-th pixel feature is smoothed; this loss is calculated on average for each image of all positive sample pairs, and then averaged in each batch of data to represent learning.

[0102] In step S2 of this embodiment, the specific process of obtaining the part segmentation network through self-supervised training based on the part-whole relationship constraint is as follows:

[0103] Self-supervised constraints of geometric concentration loss, equivalence loss, semantic consistency loss, and foreground-background discrimination loss are adopted for self-supervised learning training, and finally a part-whole relationship discovery model capable of outputting part segmentation, that is, a part segmentation network, is obtained.

[0104] The definition process of the geometric concentration loss is as follows:

[0105] Generally speaking, the pixels of the same target part will be more spatially concentrated on the same image and form a connected component without occlusion or multiple instances; based on this, geometric concentration is an important property for forming part segmentation; therefore, a loss term is used to encourage the spatial concentration of the same part;

[0106] For the partial center of a certain part k on axis u There is:

[0107]

[0108] For the partial center of a certain part k on axis v There is:

[0109]

[0110] Where is a normalization term used to convert the partial response map into a spatial probability distribution function. After that, the geometric concentration loss is defined as:

[0111]

[0112] Moreover, this loss is differentiable. This loss function encourages each part to form geometric concentration and attempts to minimize the variance of the spatial probability distribution function R(k, u, v) / z k of.

[0113] The definition process of the equivalence loss is as follows:

[0114] The part-whole relationship that the present invention hopes to obtain is robust to the appearance and pose changes of the target. Therefore, for each training image, a random spatial transformation T s (·) and appearance variation T a (·) are used. For the input image and the transformed image, the corresponding partial response maps Z and Z' are obtained respectively. According to these two partial response maps, the partial centers and are calculated respectively. After that, the equivalence loss can be defined as:

[0115]

[0116] Where D KL (·) is the KL divergence distance, is the balance coefficient;

[0117] The first term in the above formula corresponds to the equivalence constraint of part segmentation, and the second term in the above formula corresponds to the equivalence constraint of partial centers.

[0118] The definition process of the semantic consistency loss is as follows:

[0119] Although the equivalence loss has made some of the segmentation results robust to some appearance and pose changes, these synthetic transformations still cannot fully guarantee the consistency between different instances; for example, the appearance and pose changes between images often cannot be modeled by artificial transformations; to encourage the semantic consistency between different target instances, this needs to be explicitly reflected in the loss function;

[0120] The intermediate layer information of the neural network has the semantic information of the object and its parts. Therefore, a loss function that constrains semantic consistency can be used to utilize the hidden information contained in the features of the neural network pre-trained on ImageNet. Representative feature clusters can be found from the given pre-trained classification features and made to correspond to different part segmentations;

[0121] Formally, given the C-dimensional classification feature it is desired to find K representative partial feature vectors d k ∈R D , k ∈ {1, 2, …, K}. At the same time, it is desired to learn the part segmentation results and the corresponding dictionary of partial feature vectors such that the classification feature is close to d k . Then there is the following semantic consistency loss:

[0122]

[0123] where V(u, v) is the feature vector at the spatial position (u, v). Through the constraint of the semantic consistency loss, the partial basis vectors w k shared by different target instances can be learned to form the semantic dictionary {w k}, which guarantees the cross-instance semantic consistency, thus ensuring that the same part response corresponds to similar semantic features in the pre-trained classification feature space;

[0124] When training the semantic consistency loss, there is a possibility that different partial bases correspond to similar feature vectors, especially when K is large or the rank of the subspace is smaller than K. Similar partial bases may lead to noise in the part segmentation results. For example, multiple parts may actually correspond to the same part block; therefore, an additional orthogonality constraint on the partial basis vectors w k is introduced to distinguish different basis vectors from each other. Let represent the normalized partial basis vectors for each row Formally, the orthogonality constraint is taken as a loss function acting on :

[0125]

[0126] where is the F-norm, ∏ K is the identity matrix of size K×K; with this constraint, the mutual correlation of different basis vectors is minimized, and more accurate partial basis vectors are obtained, thus obtaining a better partial segmentation result.

[0127] The definition process of the foreground-background discrimination loss is as follows:

[0128] In addition to the above losses to extract the part-whole relationship of the target, an additional loss function is needed to enable the model to distinguish the target whole and the background part in the image; for this purpose, a saliency detection model pre-trained on other training sets is used to generate a saliency map, and using the saliency map, the background loss function can be obtained as:

[0129]

[0130] where D ∈ [0, 1] H×W is the saliency map, H represents the number of rows of the matrix, W represents the number of columns of the matrix, D(u, v) is the saliency value of the saliency map at the spatial position (u, v), and R(0, u, v) is the segmentation result of the background.

[0131] In summary, using multiple loss functions to train the partial segmentation network and the semantic partial basis, the obtained objective function is a linear combination of multiple loss functions:

[0132]

[0133] In the formula, λ con , λ eqv , λ sc , λ bg are the balance coefficients of the corresponding loss functions respectively.

[0134] In step S3 of this embodiment, the specific process of training the key point estimator through regression learning is as follows:

[0135] Using the partial response map Z(k) obtained by the segmentation network H×W , where k takes 1, 2,..., K. For each partial response map, a series of feature points are extracted using the non-maximum suppression method, and this series of feature points are used as the input of the key point estimator. The key point estimator is a multi-layer perceptron, and the output also obtains a heat map. The heat map is processed using non-maximum suppression to obtain a series of key points

[0136] Denote the normalized labeled key point as kp i =(a i , b i ), a i ∈[0, 1], b i∈[0,1], the estimated key points are Then there is a regression loss:

[0137]

[0138] Generally speaking, the data required for pose estimation is a quadruple containing the target image, the key points on the image, the calibrated perspective features corresponding to the key points, and the depth map. The calibrated perspective features corresponding to the key points are the 3D coordinate points corresponding to the 2D key points on the image in the 3D calibration coordinate space. The depth map is a grayscale image with the same size as the image, and the gray value corresponds to the depth. Using the previously pre-trained visual backbone model and part of the segmentation network, a partial segmentation result is obtained. Further, using a small number of target images with labeled key points and their corresponding calibrated perspective features as learning targets, with the specific values of the partial segmentation result as input, and then through regression learning training, a network is obtained as a key point estimator. Then, through the key point estimator obtained by fine-tuning on few samples above, the data acquisition and annotation process can be simplified. The target image and the corresponding depth map in the quadruple can be directly acquired through the sensor, while the key points on the image and the calibrated perspective features corresponding to the key points can be generated by the key point estimator obtained by fine-tuning on few samples, thus effectively reducing the data annotation complexity and workload.

[0139] When performing pose estimation, this embodiment refers to the classic work on pose estimation, "StarMap for Category-Agnostic Keypoint and Viewpoint Estimation" published in ECCV2018. This work predicts three components for each input image: the key point map (StarMap), the calibrated perspective features, and the depth map, where StarMap is a single-channel heat map, and its local maximum encodes the position of the corresponding point in the image. Compared with using StarMap in this work to obtain category-agnostic key points, the present invention uses the output of the above key point estimator as StarMap and its corresponding calibrated perspective features, and then the target pose can be estimated by further combining the depth map.

[0140] Given the coordinates of the key points in the image, the corresponding calibrated perspective features, and the depth map, the perspective estimation result (pose estimation result) of the input image compared with the calibrated perspective can be output through an optimization method.

[0141] Denote p i =(u i –c x ,v i –c y ,d i ) as the 3D coordinates of the key points before normalization, where (cx , c y ) is the image center; denote q i as the corresponding part under the calibration perspective. Denote the value of each key point on the heat map as w i ∈ [0, 1], representing a confidence score. It is desired to solve the similarity transformation parameterized by a scalar s ∈ R + , a rotation matrix R ∈ SO(3), and a translation t between the camera coordinate system and the world coordinate system, that is, it can be obtained by minimizing the following objective function:

[0142]

[0143] where w i represents the confidence score, and N1 represents the number of key points;

[0144] s ★ , R ★ , t ★ is the optimal parameterization obtained after minimizing the objective function, and s ★ , R ★ , t ★ are the pose estimation results of the target image.

[0145] There is an explicit solution to the above equation, that is:

[0146]

[0147] where UΣV T = M is the singular value decomposition, is the mean of p i , q i .

[0148] The present invention pre-trains a visual backbone model using a self-supervised learning algorithm based on a contrastive method, which can extract image pixel-level features applicable to fine-grained downstream tasks such as pose estimation and part segmentation, and can meet the requirements of the target overall relationship discovery algorithm. The present invention introduces a self-supervised visual backbone model and a part segmentation network, which can, under the condition of a small number of data sample annotations, fine-tune and train to obtain a key point estimator, and use the key point estimator to directly provide key points and calibration perspective features for the pose estimation task, which can effectively reduce the manual annotation workload and data annotation complexity, obtain effective pose estimation, and better complete the image understanding task.

[0149] Example 2:

[0150] The following is a specific example for explaining the pose estimation method based on self-supervised learning in the above Example 1.

[0151] 1. Training of the visual backbone model based on self-supervised learning of pixel-level proxy tasks:

[0152] Feature pre-training is carried out using the widely used ImageNet-1K dataset, which contains approximately 1.28 million training images. ResNet-50

[30] is used as the backbone network. Two branches use different encoders. One uses a conventional backbone network and a conventional projection head, and the other uses a momentum network and a projection head obtained by using the moving average parameter update method with a conventional backbone network. The Pixel Propagation Module (PPM) is applied to the conventional branch. Conventional data augmentation strategies are adopted, that is, two slices obtained by independent sampling on the same image are rescaled to a size of 224×224, and random horizontal flipping, color distortion, Gaussian blur, and overexposure are performed. The loss calculation for slice pairs without overlap is skipped, that is, only a small part of all slices is calculated.

[0153] 400 epochs are used as the training length. During training, the LARS optimizer with a base learning rate of 1.0 and cosine annealing learning rate scheduling is used, and the learning rate is linearly scaled with respect to the batch size by lr = lr base ×#bs / 256. The weight decay is set to 1e-5. The total batch size is set to 1024 and is distributed to 8 V100 GPUs for optimization. For the momentum encoder, the momentum value gradually increases from 0.99 to 1. Synchronized batch normalization is also used during training.

[0154] 2. Training of the partial segmentation network:

[0155] Multiple loss functions are used to train the partial segmentation network and the semantic partial basis, including the geometric concentration loss equivalence loss and semantic consistency loss as well as foreground-background discrimination loss The final objective function is a linear combination of the above loss functions:

[0156]

[0157] For spatial transformation, random rotation, translation, scaling, and thin plate spline interpolation are adopted; for color transformation, random perturbations of brightness, contrast, saturation, and hue are adopted. Then, through a deep learning optimizer, different learning rates are used for the partial segmentation network and the visual backbone model (the learning rate of the partial segmentation network is greater than that of the visual backbone model), and fine-tuning is performed simultaneously on the self-supervised objective partial-global relationship.

[0158] 3. Training of pose estimation:

[0159] Annotations of 2D keypoints here, along with their corresponding depths and 3D localizations in the calibrated view, are required to train the hybrid representation. Such training data is available and open to the public. The 2D keypoint annotations for each image can be directly recovered and are widely available. Given an interactive 3D user interface such as MeshLab, it is not difficult to annotate 3D keypoints of CAD models. The calibrated view of a CAD model can be defined as the front view where the largest dimension of the target 3D bounding box is scaled to [-0.5, 0.5]. Note that it is only necessary to annotate some 3D CAD models in each category. Because the degree of variation in keypoint configurations is much less than that of image appearances. Given a set of images and a small series of CAD models corresponding to the category, human annotators will select the CAD model that is closest to the content of the picture and perform similar operations on Pascal3D+ and ObjectNet3D. By dragging the selected CAD model to align with the image appearance, a rough view can be obtained. In summary, all the annotations for training the hybrid representation are relatively easy to obtain. Assuming that the StarMap method is transferable for both depth estimation and estimation of calibrated view features, after obtaining the relevant annotations on the public dataset, it is possible to use the model trained on the public dataset to fine-tune and obtain an estimation model for targets unknown to other CAD models.

[0160] The part segmentation network obtained through self-supervised learning can obtain the part-whole relationship of the target. The part-whole relationship is reflected in the form of part segmentation. The part centers of each part segmentation are extracted and aggregated to obtain StarMap, thus eliminating the need to annotate keypoints on targets unknown to other CAD models.

[0161] The pose estimation network requires calibrated view features and depth maps. The calibrated view features provide the 3D localization of keypoints in the calibrated view. In the implementation, three channels are used to represent the calibrated view features, that is, the part centers obtained in the part segmentation network are used as keypoints, and the values in the three channels correspond to the 3D positions of the corresponding pixels in the calibrated coordinate system. When considering the keypoint configuration space in the calibrated space, it is easy to find features that are invariant to target poses and image appearances (scaling, translation, rotation, illumination), small changes in target shapes (e.g., the left front wheels of different vehicles will always be in the front left of the vehicle), and small changes in target categories (the front wheels of different categories will always be in the front bottom position). Although the calibrated view features only provide 3D localization, this can still be utilized to classify keypoints through the nearest neighbor association of category-level keypoint templates.

[0162] The training process of a conventional pose estimation network is as follows, regarded as the pre-training process of the pose estimation network: all three output components of the model are subject to supervised learning. The training is specifically completed through supervised heatmap regression. For example, the L2 distance between them and the ground truth is minimized on the five-channel heatmap of the output. It should be noted that for the calibrated perspective features and depth maps, only the outputs at the spike positions are concerned, and the outputs at non-peak positions are ignored, without forcing them to zero. At this time, the network output and the ground truth can be multiplied by a mask matrix, and then the standard L2 loss is used for training.

[0163] In subsequent applications, for the pre-trained pose estimation network, replace StarMap with the part segmentation centers obtained by the part-whole relationship discovery algorithm as key points, and introduce the information extracted by self-supervised learning to achieve the perspective estimation result (pose estimation result) on the target object without key point annotation.

[0164] Embodiment 3:

[0165] The present invention also provides a self-supervised learning pose estimation system based on a contrastive method for implementing a self-supervised learning pose estimation method based on a contrastive method in Embodiment 1 above.

[0166] The system includes a visual backbone model unit, a part segmentation network unit, a key point estimator unit, and a pose estimator unit that are communicatively connected to a controller;

[0167] The visual backbone model unit uses a publicly available image dataset to pre-train a visual backbone model based on a self-supervised learning algorithm based on a contrastive method, and outputs image features through the visual backbone model;

[0168] The part segmentation network unit uses the image features to self-train a part segmentation network based on part-whole relationship constraints, and outputs a part response map through the part segmentation network;

[0169] The key point estimator unit takes the image with key points marked and its corresponding calibrated perspective features as the learning target, takes the feature points of the part response map as the input, and then trains a network through regression learning as the key point estimator, and outputs the key point map and the calibrated perspective feature map corresponding to the image through the key point estimator;

[0170] The target image to be evaluated in the controller is sequentially processed through the visual backbone model unit, the part segmentation network unit, and the key point estimator unit to obtain the key point map and the calibrated perspective feature map of the target image. Then the controller directly obtains the depth map of the target image through the sensor, and the controller inputs the key point map, the calibrated perspective feature map, and the depth map of the target image into the pose estimator unit;

[0171] The pose estimation unit filters out multiple key points from the key point map of the target image through the non-maximum suppression algorithm, extracts the coordinates of multiple key points, and then uses the key point coordinates to extract the calibrated perspective features q at the corresponding positions of multiple key points on the calibrated perspective feature map and the depth map i and the depth value d i ; then, combining the depth value d i and the key point coordinates, the three-dimensional coordinates p of multiple key points in the camera coordinate system are obtained i , and then the transformation relationship between the camera coordinate system and the world coordinate system is represented as a similarity transformation, which is parameterized by a scalar s ∈ R + , a rotation matrix R ∈ SO(3), and a translation t, and is obtained by minimizing the following objective function:

[0172]

[0173] where w i ∈ [0,1], representing the confidence score, and N1 represents the number of key points;

[0174] s ★ , R ★ , t ★ is the optimal parameterization obtained after minimizing the objective function, and s ★ , R ★ , t ★ are the pose estimation results of the target image. Finally, the pose estimation unit outputs the pose estimation results and feeds them back to the controller;

[0175] The controller displays the results on the display screen

[0176] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention

Claims

1. A pose estimation method based on self-supervised learning, characterized in that, It includes the following steps: S1. Using a publicly available image dataset, a visual backbone model is pre-trained based on a self-supervised learning algorithm of a contrastive method, and the visual backbone model outputs image features; S2. Using the image features, a part segmentation network is self-supervised trained based on part-whole relationship constraints, and the part segmentation network outputs part response maps; S3. Taking the image marked with key points and its corresponding calibrated perspective features as the learning target, using the feature points of the part response map as the input, and then training a network through regression learning as a key point estimator, and the key point estimator outputs the key point map and the calibrated perspective feature map corresponding to the image; S4. Inputting the target image into the trained visual backbone model to obtain the image characteristics of the target image, then inputting the image characteristics of the target image into the trained part segmentation network to obtain the part response map of the target image, and then inputting the part response map of the target image into the trained key point estimator to obtain the key point map and the calibrated perspective feature map of the target image; S5. Obtain the depth map of the target image, and use the non-maximum suppression algorithm to screen out multiple key points from the key point map of the target image, extract multiple key point coordinates, and then use the key point coordinates to extract the calibration perspective features q corresponding to the positions of multiple key points on the calibration perspective feature map and the depth map i and the depth value d i ; S6. Combine with the depth value d i and the key point coordinates to obtain the three-dimensional coordinates p of multiple key points in the camera coordinate system i . Represent the transformation relationship between the camera coordinate system and the world coordinate system as a similarity transformation, which is parameterized by a scalar s ∈ R + , a rotation matrix R ∈ SO(3), and a translation t, and is obtained by minimizing the following objective function: where w i ∈ [0, 1], representing the trust score, and N1 represents the number of key points; s ★ ,R ★ ,t ★ is the optimal parameterization representation obtained after minimizing the objective function, s ★ ,R ★ ,t ★ is the pose estimation result of the target image.

2. The pose estimation method based on self-supervised learning according to claim 1, characterized in that, The image dataset used in step S1 includes ImageNet-1K or ImageNet-21K.

3. The pose estimation method based on self-supervised learning according to claim 1, wherein The specific process of pre-training the visual backbone model based on the self-supervised learning algorithm of the contrastive method in step S1 is as follows: Introduce a proxy task at the pixel level. The proxy task involves two parts. One is a pixel propagation module, and the other is an asymmetric structure design. One branch of the structure design generates a normal feature map, and the other branch combines the pixel propagation module. The asymmetric structure design only requires the consistency of positive sample pairs and does not require careful debugging of negative sample pairs; For each pixel feature, its smoothly transformed vector is calculated through the pixel propagation module. This vector is obtained by propagating all pixel features on the same image Ω to the current pixel feature, as shown in the following formula: y i = Σ j∈Ω s(x i , x j )·g(x j ) where x i is the i-th pixel feature, and x j is the j-th pixel feature, and y i is the vector obtained by smoothing transformation of the i-th pixel feature; where s(·,·) is a similarity function, defined as follows: s(x i ,x j ) = (max(cos(x i ,x j )), 0)) γ where γ is a sharpness index that controls the similarity function and is default set to 2; g(·) is a transformation function, instantiated through several linear layers containing batch normalization and rectified linear units; In the asymmetric structure design, there are two different encoders: one is a propagation encoder loaded with a pixel propagation module for post-processing to generate smooth features, and the other is a momentum encoder without a pixel propagation module; both enhanced perspectives are fed into the two encoders, and the features generated by different encoders are encouraged to be consistent; Among them, represents the pixel propagation loss, where i and j are two pairs of positive pixels based on the threshold assignment rule in the enhanced perspective; x i ’ is the i-th pixel feature enhanced by the momentum encoder, and x′ j is the j-th pixel feature enhanced by the momentum encoder, and y j is the vector after the smooth transformation of the j-th pixel feature; this loss is averaged for each image of all positive sample pairs and then averaged for each batch of data to represent learning.

4. The pose estimation method based on self-supervised learning according to claim 1, characterized in that The specific process of self-supervised training to obtain the part segmentation network based on part-whole relationship constraints in step S2 is as follows: Adopt self-supervised constraints of geometric concentration loss, equivalence loss, semantic consistency loss, and foreground-background discrimination loss for self-supervised learning training, and finally obtain a part-whole relationship discovery model that can output part segmentation, that is, the part segmentation network.

5. A method for pose estimation based on self-supervised learning according to claim 4, characterized in that The definition process of the geometric concentration loss is as follows: Pixels of the same target part are more spatially concentrated on the same image and form a connected component without occlusion or multiple instances. Based on this, geometric concentration is an important property for forming part segmentation. Therefore, a loss term is used to encourage the concentration of the same part in spatial distribution; For the partial center of a certain part k on the axis u There is: For the partial center of a certain part k on the axis v There is: where is a normalization term used to transform the partial response map into a spatial probability distribution function. After that, the geometric concentration loss is defined as: Moreover, this loss is differentiable. This loss function encourages each part to form a geometric concentration and attempts to minimize the variance of the spatial probability distribution function R(k, u, v) / z k of.

6. The pose estimation method based on self-supervised learning according to claim 5, wherein The definition process of the equivalence loss is as follows: For each training image, use a random spatial transformation T s (·) and appearance variation T a (·). For the input image and the transformed image, obtain the corresponding partial response maps Z and Z', respectively. According to these two partial response maps, calculate the partial centers and respectively. After that, the equivalence loss is defined as: Among which D KL (·) is the KL divergence distance, is the balance coefficient; The first term in the above formula corresponds to the equivalence constraint of part segmentation, and the second term corresponds to the equivalence constraint of part centers.

7. The method for pose estimation based on self-supervised learning according to claim 6, characterized in that The definition process of the semantic consistency loss is as follows: The intermediate layer information of the neural network has the semantic information of the target and parts. Therefore, a loss function that constrains semantic consistency is used to utilize the hidden information contained in the features of the ImageNet pre-trained neural network, and find representative feature clusters from the given pre-trained classification features to correspond to different part segmentations; Formally, given a C-dimensional categorical feature it is desired to find K representative partial feature vectors d k ∈R D , k ∈ {1, 2, …, K}, and at the same time, it is desired to learn partial segmentation results and the corresponding partial feature vector dictionaries such that the categorical feature is close to d k Then there is the following semantic consistency loss: Where V(u, v) is the feature vector at the spatial location (u, v). Through the constraint of the semantic consistency loss, the partial basis vectors w shared by different target instances can be learned. k The semantic dictionary {w k} composed of them ensures cross-instance semantic consistency, thus ensuring that the same partial responses will correspond to similar semantic features in the pre-trained classification feature space. When training the semantic consistency loss, there is a possibility that different partial bases correspond to similar feature vectors. Therefore, an additional orthogonal constraint on the partial basis vectors w k is introduced to distinguish different basis vectors. Let represent the partial basis vectors with each row being normalized Formally, the orthogonal constraint is regarded as a loss function acting on : where is the F-norm, is the identity matrix of size K×K; Through this constraint, the cross-correlation of different basis vectors is minimized to obtain more accurate part basis vectors, thereby obtaining better part segmentation results.

8. The pose estimation method based on self-supervised learning according to claim 4, wherein The definition process of the foreground-background discrimination loss is as follows: A saliency detection model pre-trained on other training sets is used to generate a saliency map. Using the saliency map, the background loss function is: where D ∈ [0, 1] H×W is the saliency map, H represents the number of rows of the matrix, W represents the number of columns of the matrix, D(u, v) is the saliency value of the saliency map at the spatial position (u, v), and R(0, u, v) is the segmentation result of the background.

9. The method for pose estimation based on self-supervised learning according to claim 8, characterized in that, Training the part segmentation network and the semantic part basis using multiple loss functions, the objective function obtained is a linear combination of multiple loss functions: where λ con , λ eqv , λ sc , λ bg are the equilibrium coefficients of the corresponding loss functions, respectively.

10. A method for pose estimation based on self-supervised learning according to claim 1, characterized in that, The specific process of training the key point estimator through regression learning in step S3 is as follows: Partial response map Z(k) obtained using the segmentation network H×W , where k takes values 1, 2, …, K. For each partial response map, a series of feature points are extracted using non-maximum suppression. This series of feature points is used as the input to the key point estimator, which is a multi-layer perceptron. The output also obtains a heat map, and a series of key points are obtained by processing the heat map using non-maximum suppression Denote the normalized labeled key points as kp i =(a i , b i ), a i ∈[0, 1], b i ∈[0, 1], and the estimated key point is Then there is a regression loss:

Citation Information

Patent Citations

  • Color image hand posture estimation method for shielding condition

    CN111027407A

  • Image key point detection method based on feature pyramid network

    CN111126412A