Human pose estimation domain generalization method based on contour information

CN118968546BActive Publication Date: 2026-09-29UESTC (SHENZHEN) ADVANCED RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410970722.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-09-29
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

然而,基于轮廓的方法在处理复杂场景时可能会遇到挑战,主要原因是传统的轮廓检测方法对人体的轮廓检测率较低

Benefits of technology

[0029]1)本发明通过轮廓信息构建了不同域数据的统一表征,基于现有轮廓检测和实例分割技术对人体姿态估计进行优化,以过滤背景噪声并提升人体部分的轮廓检测精度,显著提高了对人体轮廓提取的质量,为关键点推断提供了更准确的轮廓信息;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968546B_ABST
    Figure CN118968546B_ABST
Patent Text Reader

Abstract

The application discloses a human posture estimation domain generalization method based on contour information, first, three sub-models are set and pre-trained by using corresponding training sample sets respectively, including an instance segmentation model, a contour detection model and an initial human posture estimation model, then a human posture estimation model based on contour information is constructed based on the three pre-trained sub-models, two contour detection models are used to extract human contours from an input image and a pedestrian image obtained by segmentation respectively and a final human contour is obtained by fusion, an amplitude processing module is added after each layer of a Transform encoder in a visual extraction encoder in the initial human posture estimation model to enhance the features, and after the human posture estimation model based on contour information is trained, the model is used to estimate a human posture image of the input image. The application improves the performance of human posture estimation domain generalization by introducing contour information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human pose estimation technology, and more specifically, relates to a generalization method for human pose estimation based on contour information. Background Technology

[0002] Currently, deep neural networks (DNNs) demonstrate superior performance in human pose estimation. However, DNNs often only fit data from the training set. When applied to different scenarios, DNNs may encounter insufficient generalization ability. This is because deep learning models are typically trained on large amounts of data, learning patterns and features from that data. If there are significant differences between the training data and the actual application scenario, the model may not adapt well to the new scenario, leading to performance degradation. To improve the generalization ability of deep neural networks in human pose estimation, researchers have adopted various strategies, which can be categorized as follows: contour representation-based methods, data generation-based methods, and optimization-based methods.

[0003] While existing methods have achieved significant success, the domain generalization problem in human pose estimation remains far from solved. In contour-based methods, edges and contours typically refer to regions where color or brightness changes significantly. These regions are often the boundaries between objects. Detecting these edges and contours yields important information such as shape and structure in the image, which can be further used for tasks like object recognition and pose estimation. Other work has demonstrated that human contours play a crucial role in the generalization of human pose estimation. However, contour-based methods may encounter challenges when handling complex scenes, primarily due to the low detection rate of traditional contour detection methods for human contours. Low-quality contour extraction significantly impacts the domain generalization performance of human pose estimation. Furthermore, existing pre-trained models are often trained on natural images, which differ significantly from human contour maps, making it difficult for them to generalize effectively when processing contour maps. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a generalization method for human pose estimation domain based on contour information. By introducing contour information to construct a unified representation between different data, and optimizing the initial human pose estimation model based on frequency domain amplitude information, the model's ability to extract contour features is improved, thereby improving the performance of human pose estimation domain generalization.

[0005] To achieve the above-mentioned objectives, the human pose estimation domain generalization method based on contour information of the present invention includes the following steps:

[0006] S1: Construct three sub-models according to actual needs and pre-train them using corresponding training sample sets. The three sub-models are as follows:

[0007] An instance segmentation model is used to remove the background from the image input to the model and extract the human image.

[0008] A contour detection model is used to detect contour edges in images input to the model.

[0009] An initial human pose estimation model is used to estimate the human pose in the image input to the model. The initial human pose estimation model includes a visual extraction encoder and a pose estimation decoder. The visual extraction encoder includes L stacked Transformer coding blocks, which are used to extract visual features from the input image layer by layer. The pose estimation decoder is used to generate a human pose image based on the visual features.

[0010] S2: Construct a human pose estimation model based on contour information, including an instance segmentation model, a first contour detection model, a second contour detection model, a human contour fusion module, a visual extraction encoder, and a pose estimation decoder, wherein:

[0011] The parameters of the instance segmentation model are fixed as those obtained during pre-training in step S1, and are used to extract the human image x from the input image x. c Then the human body image x c Send to the second contour detection model;

[0012] The parameters of the first contour detection model are fixed as the parameters Θ obtained in the pre-training in step S1, which are used to extract the human contour image F(x;Θ) from the input image x and send it to the feature fusion module;

[0013] The parameters Θ of the second contour detection model c These are trainable parameters, initially set to the parameters Θ obtained during pre-training in step S1, used to train the human image x. c Extracting the human body contour image F(x) c ;Θ c And send it to the feature fusion module;

[0014] The human body contour fusion module is used to process human body contour images F(x; Θ) and F(x... c ;Θ c The images are then fused to obtain the fused human body contour image y. c And send it to the visual extraction encoder, the fusion formula is as follows:

[0015] y c =F(x;Θ)+Z(F(x) c ;Θ c );Θz )

[0016] Wherein, Z(·;Θ z ) indicates that the parameter is Θ z Zero convolutional layers;

[0017] The visual extraction encoder is used to process the input image x and the human body contour image y respectively. c Visual feature extraction is performed to obtain input visual features and human contour visual features, which are then output to the pose estimation decoder. The visual extraction encoder includes L stacked Transformer coding blocks and an amplitude processing module. The parameters of the Transformer coding blocks are fixed to the parameters pre-trained in step S1. The i-th amplitude processing module is used to perform amplitude processing on the two features output by the i-th Transformer coding block, i = 1, 2, ..., L. The specific processing method is as follows:

[0018] Let the features of the input image extracted from the i-th layer Transformer coded block be . Human body contour image features are Where D i H represents the number of channels of the feature extracted from the i-th Transformer coding block. i ×W i This represents the size of the feature extracted by the i-th Transformer encoding block; for the input image feature X i and human body contour image features X i,c Perform 3D Fourier transforms on each feature, then combine the real and imaginary features obtained from the 3D Fourier transforms to obtain the corresponding frequency domain features. and

[0019]

[0020] Then each frequency domain feature is enhanced using the following formula:

[0021]

[0022] in, This indicates element-wise multiplication, σ represents the sigmoid activation function, LN represents layer normalization, and Conv... 3×3 This indicates a convolution kernel with a parameter kernel of 3x3;

[0023] Then the enhanced feature F i ′ and F i, ′ c The input image features X are converted into spatial domain representation by performing 3D inverse Fourier transform. i ′ and pedestrian silhouette features X i ′,c The amplitude processing modules from layer 1 to layer i-1 output the two obtained features to the next layer Transformer encoding block, and the amplitude processing module of layer L outputs the two obtained features as input visual features and human contour visual features to the pose estimation decoder.

[0024] The parameters of the pose estimation decoder are fixed to the parameters obtained in the pre-training in step S1, and are used to generate human pose images based on the received input visual features and human contour visual features respectively.

[0025] S3: Obtain the training sample set for human pose estimation according to actual needs, then train the human pose estimation model based on contour information constructed in step S2, update the parameters of the second contour detection model and L amplitude processing modules, and obtain the trained human pose estimation model based on contour information.

[0026] S4: Input the image of the human pose to be generated into the human pose estimation model based on contour information trained in step S3 to obtain the corresponding human pose image.

[0027] This invention presents a contour-based generalization method for human pose estimation. First, three sub-models are set up and pre-trained using corresponding training sample sets, including an instance segmentation model, a contour detection model, and an initial human pose estimation model. Then, a contour-based human pose estimation model is constructed based on the three pre-trained sub-models. Two contour detection models are used to extract human contours from the input image and the segmented pedestrian image, respectively, and then fused to obtain the final human contour. In the initial human pose estimation model, an amplitude processing module is added after each Transformer encoder layer in the visual extraction encoder to enhance the features. After training the contour-based human pose estimation model, the human pose image of the input image is estimated using this model.

[0028] The present invention has the following beneficial effects:

[0029] 1) This invention constructs a unified representation of data from different domains through contour information, optimizes human pose estimation based on existing contour detection and instance segmentation technologies, filters background noise and improves the contour detection accuracy of human body parts, significantly improves the quality of human contour extraction, and provides more accurate contour information for key point inference.

[0030] 2) This invention designs an amplitude processing module to adapt to the perception ability of the pre-trained initial human pose estimation model on contour images, thereby improving the model's adaptability to images from different domains.

[0031] 3) This invention not only improves the generalization performance of the model in the domain of unknown targets, but also enhances the interpretability of the model by modeling and perceiving human body contour features. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating a specific implementation of the human pose estimation domain generalization method based on contour information of the present invention.

[0033] Figure 2 This is a structural diagram of the human pose estimation model based on contour information of the present invention;

[0034] Figure 3 This is a flowchart of the amplitude processing module in this invention. Detailed Implementation

[0035] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.

[0036] Example

[0037] Figure 1 This is a flowchart illustrating a specific implementation of the human pose estimation domain generalization method based on contour information according to the present invention. Figure 1 As shown, the specific steps of the human pose estimation domain generalization method based on contour information of the present invention include:

[0038] S101: Sub-model pre-training:

[0039] To better construct a human pose estimation model based on contour information, this invention first constructs three sub-models according to actual needs and pre-trains them using corresponding training sample sets. The three sub-models are as follows:

[0040] An instance segmentation model is used to remove the background from the input image and extract the human body image. Using an instance segmentation model can obtain a clean human body image after background removal, and the accuracy of the human body image can be ensured by training the model on large-scale data. In this embodiment, the ConvNext model is used for instance segmentation.

[0041] A contour detection model is used to detect human contours in an input image to obtain a human contour image. In this embodiment, the contour detection model adopts the RCF model.

[0042] The initial human pose estimation model is used to estimate the human pose in the input image. This model includes a visual extraction encoder and a pose estimation decoder. The visual extraction encoder consists of L stacked Transformer encoding blocks, used to extract visual features from the input image layer by layer. The pose estimation decoder is used to generate a human pose image based on these visual features. In this embodiment, the initial human pose estimation model uses a DinoV2 or MAE pre-trained backbone as the encoder, while the decoder uses a simple convolutional network.

[0043] S102: Constructing a human pose estimation model based on contour information:

[0044] To improve the performance of generalization in the human pose estimation domain, this invention optimizes existing edge detection models using human body masks, thereby filtering background noise and improving the accuracy of human body contour detection. Figure 2 This is a structural diagram of the human pose estimation model based on contour information of this invention. (See diagram below.) Figure 2 As shown, the human pose estimation model based on contour information of this invention includes an instance segmentation model, a first contour detection model, a second contour detection model, a human contour fusion module, a visual extraction encoder, and a pose estimation decoder. The following is a detailed description of each component model.

[0045] The parameters of the instance segmentation model are fixed as those obtained through pre-training in step S101, and are used to extract the human image x from the input image x. c Then the human body image x c Send to the second contour detection model.

[0046] The parameters of the first contour detection model are fixed as the parameters Θ obtained in the pre-training in step S101, which are used to extract the human contour image F(x;Θ) from the input image x and send it to the feature fusion module.

[0047] The parameters Θ of the second contour detection model c These are trainable parameters, initially set to the parameters Θ obtained from pre-training in step S101, used to train the human image x. c Extracting the human body contour image F(x) c ;Θ c And send it to the feature fusion module.

[0048] The human body contour fusion module is used to process human body contour images F(x; Θ) and F(x... c ;Θ c The images are then fused to obtain the fused human body contour image y. c And send it to the visual extraction encoder, the fusion formula is as follows:

[0049] y c =F(x;Θ)+Z(F(x) c ;Θ c );Θ z )

[0050] Wherein, Z(·;Θ z ) indicates that the parameter is Θ z The zero convolutional layer (i.e., a 1x1 convolutional layer with weights and biases initialized to zero).

[0051] The visual extraction encoder is used to process the input image x and the human body contour image y respectively. c Visual feature extraction is performed to obtain input visual features and human contour visual features, which are then output to the pose estimation decoder. The visual extraction encoder includes L stacked Transformer coding blocks and an amplitude processing module. The parameters of the Transformer coding blocks are fixed to the parameters pre-trained in step S101. The i-th amplitude processing module is used to perform amplitude processing on the two features output by the i-th Transformer coding block, i = 1, 2, ..., L. Figure 3 This is a flowchart of the amplitude processing module in this invention. For example... Figure 3 As shown, the specific processing method of the amplitude processing module in this invention is as follows:

[0052] Let the features of the input image extracted from the i-th layer Transformer coded block be . Human body contour image features are Where D i H represents the number of channels of the feature extracted from the i-th Transformer coding block. i ×W i This represents the size of the feature extracted by the i-th Transformer encoding block. For the input image feature X... i and human body contour image features X i,c Perform 3D Fourier Transform (3D FFT) on each part, and then merge the real and imaginary features obtained from the 3D Fourier Transform to obtain the corresponding frequency domain features. and The expressions for the 3D Fourier transform can be represented as follows:

[0053]

[0054] Each frequency domain feature is then enhanced using the following formula to improve the pre-trained model's ability to perceive contour information:

[0055]

[0056]

[0057] in, This indicates element-wise multiplication, σ represents the sigmoid activation function, LN represents layer normalization, and Conv... 3×3 This indicates a convolution kernel with a parameter kernel of 3x3.

[0058] Then the enhanced feature F i ′ and F i, ′ c The input image features X are converted into spatial domain representation by performing 3D Inverse Fourier Transform (3D IFFT) on each of them. i ′ and pedestrian silhouette features X i ′ ,c The amplitude processing modules from layer 1 to layer i-1 output the two obtained features to the next Transformer encoding block. The amplitude processing module at layer L outputs the two obtained features as input visual features and human contour visual features to the pose estimation decoder. The expressions for the 3D inverse Fourier transform can be represented as follows:

[0059]

[0060] As can be seen from the above description, in the pedestrian pose estimation model based on contour information constructed in this invention, an amplitude processing module is used to scale the feature maps of different frequency components, and to learn the feature representation in the frequency domain to adjust the visual extraction encoder's perception ability of the contour image, so that the obtained features are more suitable for the generalization requirements of the human pose estimation domain.

[0061] The parameters of the pose estimation decoder are fixed to the parameters obtained in step S1 during pre-training, and are used to generate human pose images based on the received visual features as input and human contour visual features, respectively.

[0062] S103: Training a human pose estimation model based on contour information:

[0063] According to actual needs, obtain the training sample set for human pose estimation, then train the human pose estimation model based on contour information constructed in step S2, update the parameters of the second contour detection model and amplitude processing module, and obtain the trained human pose estimation model based on contour information.

[0064] The setting of the loss function is crucial for model training. In this embodiment, the formula for calculating the loss function in the training process of the human pose estimation model based on contour information is as follows:

[0065]

[0066] Where N represents the number of training samples in the training batch, K represents the number of human pose key points, and y nk y′ is the label value of the k-th human pose keypoint in the n-th training sample of the current training batch. nk and y′ n ′ k These represent the predicted values ​​of key points of the k-th human posture in the n-th training sample of the current training batch, based on the input visual features and the visual features of the human body contour.

[0067] S104: Human pose estimation:

[0068] The image for which human pose needs to be generated is input into the human pose estimation model based on contour information trained in step S103 to obtain the corresponding human pose image.

[0069] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention.

[0070] Example

[0071] In this embodiment, the experimental conditions are set as follows: System: Ubuntu 20.04, Software: Python 3.9, Processor: Intel(R) Xeon(R) CPU E5-2678 v3@2.50GHz×2, Memory: 256GB, Graphics Processor: NVIDIA A100 x4.

[0072] This embodiment was trained on the CrowdPose and MSCOCO datasets, and then tested in 19 data domains on the Human-Art dataset. Using Dino-v2 as the initial human pose estimation model, the present invention was compared with the commonly used fine-tuning method, Adapter. Table 1 compares the domain generalization performance of the present invention and the compared methods in Embodiment 1:

[0073]

[0074]

[0075] Table 1

[0076] As shown in Table 1, the present invention improves the average generalization ability by 4.4% and 6.4% compared to state-of-the-art methods, with improvements ranging from 2.9% to 12.2% across different target domains. Furthermore, the improvement becomes more pronounced with increasing model parameters, indicating that the amplitude processing module proposed in this invention effectively preserves the generalization ability of the pre-trained human pose estimation model while enhancing the perception of human contours. These data effectively demonstrate that the human pose estimation model of this invention possesses good domain generalization ability across different scenarios.

[0077] To further verify the effectiveness of this invention under different pre-trained human estimation models, MAE was used as the original model, and this invention was compared with the current state-of-the-art fine-tuning method Adapter. Table 2 shows the performance comparison of this invention and the comparison method on the Human-Art dataset in Example 1:

[0078]

[0079]

[0080] Table 2

[0081] As shown in Table 2, the average generalization ability of this invention still reaches the state-of-the-art level under different pre-trained models. The above data effectively demonstrates that this invention has good domain generalization ability for different pre-trained models.

[0082] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.

Claims

1. A generalization method for human pose estimation based on contour information, characterized in that, Includes the following steps: S1: Construct three sub-models according to actual needs and pre-train them using corresponding training sample sets. The three sub-models are as follows: An instance segmentation model is used to remove the background from the image input to the model and extract the human image. A contour detection model is used to detect contour edges in images input to the model. An initial human pose estimation model is used to estimate the human pose in the image input to the model. The initial human pose estimation model includes a visual extraction encoder and a pose estimation decoder. The visual extraction encoder includes L stacked Transformer coding blocks, which are used to extract visual features from the input image layer by layer. The pose estimation decoder is used to generate a human pose image based on the visual features. S2: Construct a human pose estimation model based on contour information, including an instance segmentation model, a first contour detection model, a second contour detection model, a human contour fusion module, a visual extraction encoder, and a pose estimation decoder, wherein: The parameters of the instance segmentation model are fixed as those obtained during pre-training in step S1, and are used to extract the human image x from the input image x. c Then the human body image x c Send to the second contour detection model; The parameters of the first contour detection model are fixed as the parameters Θ obtained in the pre-training in step S1, which are used to extract the human contour image F(x;Θ) from the input image x and send it to the feature fusion module; The parameters Θ of the second contour detection model c These are trainable parameters, initially set to the parameters Θ obtained during pre-training in step S1, used to train the human image x. c Extracting the human body contour image F(x) c ;Θ c And send it to the feature fusion module; The human body contour fusion module is used to process human body contour images F(x; Θ) and F(x... c ;Θ c The images are then fused to obtain the fused human body contour image y. c And send it to the visual extraction encoder, the fusion formula is as follows: y c =F(x;Θ)+Z(F(x c ;I c );I z ) Wherein, Z(·;Θ z ) indicates that the parameter is Θ z Zero convolutional layers; The visual extraction encoder is used to process the input image x and the human body contour image y respectively. c Visual feature extraction is performed to obtain input visual features and human contour visual features, which are then output to the pose estimation decoder. The visual extraction encoder includes L stacked Transformer coding blocks and an amplitude processing module. The parameters of the Transformer coding blocks are fixed to the parameters obtained in step S1 during pre-training. The i-th amplitude processing module is used to perform amplitude processing on the two features output by the i-th Transformer coding block, i = 1, 2, ..., L. The specific processing method is as follows: Let the features of the input image extracted from the i-th layer Transformer coded block be . Human body contour image features are Where D i H represents the number of channels of the feature extracted from the i-th Transformer coding block. i ×W i This represents the size of the feature extracted by the i-th Transformer encoding block; for the input image feature X i and human body contour image features X i,c Perform 3D Fourier transforms on each feature, then combine the real and imaginary features obtained from the 3D Fourier transforms to obtain the corresponding frequency domain features. and Then each frequency domain feature is enhanced using the following formula: in, This indicates element-wise multiplication, σ represents the sigmoid activation function, LN represents layer normalization, and Conv... 3×3 This indicates a convolution kernel with a parameter kernel of 3x3; Then the enhanced feature F i ′ and F′ i,c The input image features X′ are converted into spatial domain representation by performing 3D inverse Fourier transform. i and pedestrian silhouette features X′ i,c The amplitude processing modules from layer 1 to layer i-1 output the two obtained features to the next layer Transformer encoding block, and the amplitude processing module of layer L outputs the two obtained features as input visual features and human contour visual features to the pose estimation decoder. The parameters of the pose estimation decoder are fixed to the parameters obtained by pre-training in step S1, and are used to generate human pose images based on the received input visual features and human contour visual features respectively. S3: Obtain the training sample set for human pose estimation according to actual needs, then train the human pose estimation model based on contour information constructed in step S2, update the parameters of the second contour detection model and L amplitude processing modules, and obtain the trained human pose estimation model based on contour information. S4: Input the image of the human pose to be generated into the human pose estimation model based on contour information trained in step S3 to obtain the corresponding human pose image.

2. The human pose estimation domain generalization method according to claim 1, characterized in that, The formula for calculating the loss function in the training process of the human pose estimation model based on contour information in step S3 is as follows: Where N represents the number of training samples in the training batch, K represents the number of human pose key points, and y nk y′ is the label value of the k-th human pose keypoint in the n-th training sample of the current training batch. nk and y′ n ′ k These represent the predicted values ​​of key points of the k-th human posture in the n-th training sample of the current training batch, based on the input visual features and the visual features of the human body contour.

Citation Information

Patent Citations

  • Multi-task deep learning model for improving human body analysis effect

    CN111709289A

  • Attitude estimation method and device, electronic equipment and storage medium

    CN113822102A