A 2D Human Pose Estimation Method Based on the Coupling of Local Features and Global Representations

By combining convolutional neural network and Transformer in human pose estimation, local features and global representations are extracted, and multiple feature coupling and communication are performed, the problem of failure to fully utilize coupled features in the prior art is solved, and a higher accuracy and stable human pose estimation is achieved.

CN115661858BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211249602.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-12
Publication Date
2025-05-27
Estimated Expiration
2042-10-12

AI Technical Summary

Technical Problem

The prior art fails to fully explore the role of coupling local features and global representations in human posture estimation, resulting in insufficient accuracy when dealing with complex postures and texture changes.

Method used

Convolutional neural network is used to extract local features, and global representation is captured through Transformer, combining feature coupling modules and local-global information exchange modules, feature coupling and communication are performed multiple times, and finally returning to the human joint node position.

Benefits of technology

By coupling local features and global representations, the network can more accurately capture the details and overall structure of human poses, improving the accuracy and generalization ability of human pose estimation, especially when dealing with complex poses and texture changes, it performs more stably.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661858B_ABST
    Figure CN115661858B_ABST
Patent Text Reader

Abstract

The present invention claims protection for a 2D human pose estimation method based on the coupling of local features and global representations, comprising the following steps: obtaining a dataset, selecting and partitioning the training set and validation set required for the human pose estimation task; obtaining the specific position of the human body, segmenting the sample image, and simultaneously performing data augmentation on the segmented image; inputting the processed data into a convolutional neural network designed based on the open-source deep learning framework Pytorch; the output heatmap can represent the position of the human body joints; calculating the loss between the output heatmap of the network model and the corresponding annotated heatmap, training and optimizing the detection model; in step five, using the optimized model parameters to detect the position of the human body joints in the real-scene image to obtain the corresponding human bone framework. The present invention improves the feature extraction ability of the network based on the local features extracted by the coupled convolutional neural network and the global representations captured by the Transformer, and realizes an end-to-end human pose estimation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to convolutional neural networks, Transformers, and human pose estimation technologies, and particularly to a method for human pose estimation based on coupling local features and global representations. Background Art

[0002] Human pose estimation is one of the important research topics in computer vision and also one of the very challenging tasks in computer vision. Its aim is to enable a machine to recognize the spatial positions of human joint points in an input image and generate a human bone framework representing the pose. With the development of artificial intelligence theory, deep learning-based human pose estimation algorithms have gradually replaced traditional algorithms and become a popular research topic in the computer field, opening the curtain on the era of artificial intelligence. The application fields include intelligent sports, intelligent monitoring, security, and other fields.

[0003] Currently, neural network-based human pose estimation methods are divided into two categories: top-down and bottom-up. The former first determines the position of the human body and then determines the positions of human joint points within the position. The latter first calculates the positions of all joint points in the image and then groups them using a human model fitting or related algorithm to form an independent human bone framework. Top-down algorithms include: Pose Machines, HRNet, TransPose, and bottom-up algorithms include: OpenPose, etc. Among the above methods, only TransPose uses both convolutional neural networks and Transformers, but it does not deeply explore the role of coupling local features and global representations in the field of human pose estimation.

[0004] CN114565938A, a three-dimensional human pose estimation method integrating local and global features, includes: proposing an observation mode I applicable to a 3D human pose estimation model; proposing an observation mode II applicable to a 3D human pose estimation model; adjusting the trained image-2D pose network; a fully connected module capturing the global features of the human pose; a grouped connection module capturing the local features of the human pose; fusing the local features and global features and regressing them to a 3D pose, and finally obtaining a 3D human pose estimation model. The beneficial effect of the present invention is: a three-dimensional human pose estimation method integrating local and global features is proposed. By separately using a fully connected module and a grouped connection module to capture the global features and local features of the human pose, and then fusing the extracted features, a 3D human pose estimation model with better generalization performance is obtained, and the pose estimation performance index with high local feature similarity is significantly improved.

[0005] Difference: CN114565938A uses a convolutional neural network to extract local features and global features. This paper uses a convolutional neural network to extract local features and a Transformer to capture global features. After the global feature extraction module and local feature extraction module of CN114565938A extract features, there is only one fusion and then it regresses to the positions of human joints. There is only one fusion in the whole process. The method adopted by the author of this paper is: after the convolutional neural network and the Transformer extract features, they are coupled, then features are extracted based on the coupled technology, and then continue to be coupled. After the above process is repeated multiple times, the positions of human joints are regressed. There are multiple couplings in the whole process.

[0006] The advantages of this paper are as follows: Through the feature coupling module, it is possible to enrich the global information of local features and the local information of global representations; convolutional neural networks are good at extracting detailed information such as textures, and the texture information around symmetric joints is more similar; Transformers are good at capturing long- and short-distance dependence relationships to form the spatial relative relationships between joints; after coupling the two types of features using the feature coupling module, the network can not only extract rich detailed information but also capture more accurate long- and short-distance relationships; when the texture information around symmetric joints in the image of a human body joint varies greatly due to clothing changes, the network can restore these difficulties through the extracted spatial relative relationships. At the same time, due to the very flexible human body postures, the dataset cannot cover all human body postures. When the network encounters uncommon postures such as acrobatics, the network is relatively poor at capturing spatial relative relationships, but the model can make up for it through the rich detailed features extracted. Summary of the Invention

[0007] The present invention aims to solve the above problems of the prior art. A 2D human pose estimation method based on the coupling of local features and global representations is proposed. The technical solution of the present invention is as follows:

[0008] A 2D human pose estimation method based on the coupling of local features and global representations, which includes the following steps:

[0009] 1). Obtain a dataset, select and divide the training set and validation set required for the human pose estimation task;

[0010] 2). Obtain the specific positions of human joints, segment the sample images in the training set images and validation set images respectively, and at the same time perform data augmentation on the segmented images;

[0011] 3) Input the data processed in step 2) into a convolutional neural network designed based on the Pytorch open-source deep learning framework; the convolutional neural network designed based on the Pytorch open-source deep learning framework is a serial structure. The picture first extracts features through the basic feature network and then sends them to the local-global feature coupling module for feature coupling. The output of the previous feature coupling module is the input of the next feature coupling module. The output of each feature coupling module will also be sent to the head network to participate in intermediate supervision calculation. What the network finally outputs is a heat map, and each heat map represents a joint point;

[0012] 4) Calculate the loss between the output heat map of the network model and the corresponding annotated heat map. The corresponding annotated heat map is converted from the joint point positions of the original picture, and train and optimize the detection model;

[0013] 5) Use the optimized model to detect the positions of human joint points in real-scene images and obtain the corresponding human bone framework.

[0014] Further, the training set in step 1) uses the coco2017 train dataset, which contains 118,287 pictures and 149,813 visible human instances; the validation set uses the coco2017 val dataset, which contains 5,000 pictures and 6,352 visible human instances.

[0015] Further, step 2) adopts a top-down scheme. The top-down scheme means: first detect the position of the human body in the picture, and then detect the position of the human joint points in the human body position. Therefore, when processing the dataset, first perform segmentation processing according to the position of the human body in the picture, so that each segmentation result only contains one human instance.

[0016] Further, the basic feature extraction network is resnet and hrnet.

[0017] Further, the specific feature coupling of the local-global feature coupling module includes:

[0018] The local features and global representations obtained in the previous stage are used to obtain a similarity matrix through dot product operation. After the similarity matrix is calculated by softmax, it is dot product calculated with the global representation. The result of the previous calculation is channel concatenated with the local features and dimension adjustment is performed using convolution. The result obtained is the coupled local feature; the formula description of the calculation process is as follows:

[0019]

[0020]

[0021] where local mand global m represent the coupled local features and global representations respectively. Conv represents the convolution operation. local and global represent the local features and global representations before coupling respectively. and represent the dimensions of the local and global vectors. * represents the dot product operation; + represents the feature map channel concatenation.

[0022] Furthermore, the local-global feature coupling module further includes a local-global feature communication module, specifically:

[0023] The image or feature map is divided into non-overlapping blocks. After the blocks are further subdivided and encoded into different PatchTokens, the Patch Tokens represent the non-overlapping blocks formed after splitting the picture. All Patch Tokens belonging to the same block are downsampled to obtain local-global communication tokens, and the local-global communication tokens represent the high-level semantic information of this block. After equally shuffling the local-global communication tokens, each new local-global communication token contains the high-level semantic information of all blocks, and then the new local-global communication tokens are concatenated with the Patch Tokens of the original divided blocks in channels. After the above calculations, each block not only contains the semantic information representing itself, but also contains a small amount of high-level semantic information of other blocks.

[0024] Furthermore, in step 4), calculate the loss between the output heat map of the network model and the corresponding labeled heat map. The corresponding labeled heat map is converted from the joint positions of the original image, and train and optimize the detection model, specifically including:

[0025]

[0026] where loss represents the loss, num joints represents the number of human joint points labeled in the dataset, height and weight represent the height and width of the heat map respectively, output and target represent the heat map calculated by the model and the heat map representing the actual positions of human joint points respectively.

[0027] The cross-entropy function is used to calculate the loss between the heat map output by the model and the heat map converted from the true label. The size of the heat map is 56*56; since each joint is associated with a heat map, when calculating the loss, first calculate the loss between each joint, and then calculate the average value; calculating the loss between each joint point is to calculate the cross-entropy sum of 3136 pairs of elements corresponding to each pair of heat maps.

[0028] The advantages and beneficial effects of the present invention are as follows:

[0029] Based on a convolutional neural network and a Transformer, the present invention first constructs a local-global feature coupling module, which deeply couples the local features extracted by a convolutional neural network and the global representations extracted by a Transformer using the mechanisms of attention calculation and residual; secondly, a local-global information exchange module is designed to reduce the loss of model accuracy caused by the difference in the data source range between local features and global representations during the coupling process.

[0030] Couple local features and global representations by means of attention calculation, so that the local features contain global information and the global features contain local information, which can enhance the feature extraction ability of the model. At the same time, the point set calculation in the attention calculation process can be used to judge the correlation degree of vectors. The similarity of vectors with key points, especially human symmetric key points such as the wrists of the left and right hands, is relatively high. Therefore, the method of coupling local features and global representations by means of attention calculation can continuously refine and determine the positions of human joints during the calculation process. Description of the Drawings

[0031] Figure 1 is the overall framework diagram of the network model of the preferred embodiment provided by the present invention;

[0032] Figure 2 is the diagram of the local-global feature coupling module;

[0033] Figure 3 is the local-global information exchange module. Detailed Implementation Manner

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0035] The technical solution of the present invention to solve the above technical problems is:

[0036] The process of the method is as Figure 1 , after the image undergoes shallow feature extraction through the basic network, the features are sent to the local-global feature coupling module. The output of the previous module is the input of the next module, and the output of each module will be sent to the head network to calculate the heat map representing the joint positions of the final output of the model. Specifically, it is divided into the following steps:

[0037] 1), Obtain the dataset, select and divide the training set and validation set required for the human pose estimation task;

[0038] 2), Since the proposed solution by the author is of the top-down type, this solution requires knowing the specific position of the human body first. Therefore, the sample images in the training set images and validation set images are respectively segmented, and at the same time, data augmentation is performed on the segmented images;

[0039] 3), Input the data processed in step 2 into a convolutional neural network designed based on the open-source deep learning framework Pytorch; the output heat map can represent the positions of human joints;

[0040] 4) Calculate the loss between the output heatmap of the network model and the corresponding labeled heatmap (converted from the joint positions of the original image), and train and optimize the detection model;

[0041] 5) Use the optimized model parameters to detect the positions of human joints in real-scene images and obtain the corresponding human skeletal frameworks.

[0042] For the described 2D human pose estimation method based on the coupling of local features and global representations, in step 1), the training set used in this training is the coco2017 train dataset, which contains 118,287 images and 149,813 visible human instances. The validation set uses the coco2017 val dataset, which contains 5,000 images and 6,352 visible human instances.

[0043] For the described 2D human pose estimation method based on the coupling of local features and global representations, in step 2), this scheme adopts a top-down approach. Therefore, when processing the dataset, the images are first segmented according to the positions of the humans in the images, so that each segmentation result only contains one human instance.

[0044] For the described 2D human pose estimation method based on the coupling of local features and global representations, in step 3), the network is as Figure 1 shown and is described in detail as follows:

[0045] The overall structure of the network is a serial structure. The image first passes through the basic feature network to extract features and then is sent to the local-global feature coupling module for feature coupling. The output of the previous feature coupling module is the input of the next feature coupling module. The output of each feature coupling module will also be sent to the head network to participate in the intermediate supervision calculation. What the network finally outputs is the heatmap. Each heatmap represents a joint point.

[0046] Basic feature extraction network: The basic feature extraction network adopted in this network is resnet and hrnet.

[0047] Local-global feature coupling module: Its structure is as Figure 2 shown: The local features and global representations obtained in the previous stage are used to obtain a similarity matrix through dot product operation. In order to enrich the global information of the local features, the similarity matrix is calculated through softmax and then dot product calculation is performed with the global representation. In order to ensure the properties of the local features, the result of the previous calculation is concatenated with the local features in the channel dimension and convolution is used for dimension adjustment. The result obtained is the coupled local feature; the process of enriching the locality of the global representation is similar. The formula description of the calculation process is as follows:

[0048]

[0049]

[0050] where local m and global m represent the coupled local features and global representations respectively, Conv represents the convolution operation, local and global represent the local features and global representations before coupling respectively, and represent the dimensions of the local and global vectors, * represents the dot product operation; + represents the concatenation of feature map channels.

[0051] Local-Global Feature Communication Module: The structural diagram is as shown in Figure 2 and is described in detail as follows: In the local-global feature coupling module, the attention calculation range is restricted to non-overlapping blocks, but the calculation process of the convolutional neural network is continuous and overlapping. Therefore, the calculation source of the global representation is smaller than that of the local features during the coupling calculation process, that is, there is a difference in the data source ranges of the local features and the global representation, which will affect the accuracy of the model. Therefore, the author proposes a local-global feature communication module to reduce the impact caused by the difference in the data source ranges. The process is as follows: The image or feature map is divided into non-overlapping blocks, and after being further divided within the block, it is encoded into different Patch Tokens. All Patch Tokens belonging to the same block are downsampled to obtain local-global communication tokens (LGC Tokens in the figure), and the local-global communication tokens represent the high-level semantic information of this block. After equally shuffling the local-global communication tokens, each new local-global communication token contains the high-level semantic information of all blocks, and then the new local-global communication tokens are concatenated with the Patch Tokens of the original divided blocks in channels; after the above calculations, each block not only contains the semantic information representing itself, but also contains a small amount of high-level semantic information of other blocks. Through the above process, the difference in the data source ranges of the local features and the global representation can be reduced.

[0052] The described 2D human pose estimation method based on the coupling of local features and global representations, wherein step 4) includes:

[0053] Calculate the loss between the heatmap output by the model and the heatmap converted from the ground truth label using the cross-entropy function. The size of the heatmap is 56*56. Since each joint is associated with a heatmap, when calculating the loss, first calculate the loss between each joint, and then calculate the average value. Calculating the loss between each joint point is to calculate the cross-entropy sum of 3136 (56*56 = 3136) pairs of elements corresponding to each pair of heatmaps.

[0054] The described 2D human pose estimation method based on the coupling of local features and global representations, wherein step 5) includes:

[0055] Using the optimized model, select the validation set images to test the detection performance of the trained model, that is, through forward propagation, calculate the positions of human joint points, and finally restore the human joint point skeletal frameworks of all people in the picture.

[0056] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0057] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0058] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0059] The above embodiments should be understood as only illustrative of the present invention and not limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A 2D human pose estimation method based on the coupling of local features and global representations, characterized in that, it includes the following steps: 1). Obtain a dataset, and select and divide the training set and validation set required for the human pose estimation task; 2). Obtain the specific positions of human joint points, segment the sample images in the training set images and validation set images respectively, and at the same time perform data augmentation on the segmented images; 3). Input the data processed in step 2) into a convolutional neural network designed based on the Pytorch open-source deep learning framework; the convolutional neural network designed based on the Pytorch open-source deep learning framework is a serial structure. The picture first extracts features through the basic feature network and then sends them to the local-global feature coupling module for feature coupling. The output of the previous feature coupling module is the input of the next feature coupling module. The output of each feature coupling module will also be sent to the head network to participate in the intermediate supervision calculation. What the network finally outputs is the heat map, and each heat map represents a joint point; 4). Calculate the loss between the output heat map of the network model and the corresponding annotated heat map. The corresponding annotated heat map is converted from the joint point positions of the original image, and train and optimize the detection model; 5). Use the optimized detection model to detect the positions of human joint points in real-scene images and obtain the corresponding human bone framework.

2. The 2D human pose estimation method based on the coupling of local features and global representations according to claim 1, characterized in that, the training set in step 1) uses the coco2017 train dataset, which contains 118,287 pictures and 149,813 visible human instances; the validation set uses the coco2017 val dataset, which contains 5,000 pictures and 6,352 visible human instances.

3. The 2D human pose estimation method based on the coupling of local features and global representations according to claim 1, characterized in that, the step 2) adopts a top-down scheme. The top-down scheme means: first detect the position of the human body in the picture, and then detect the position of the human joint points in the human body position. Therefore, when processing the dataset, first perform segmentation processing according to the position of the human body in the picture, so that each segmentation result only contains one human instance.

4. The 2D human pose estimation method based on the coupling of local features and global representations according to claim 1, characterized in that, the basic feature extraction network is resnet and hrnet.

5. The 2D human pose estimation method based on the coupling of local features and global representations according to claim 4, characterized in that, the specific feature coupling of the local-global feature coupling module includes: The local features and global representations obtained in the previous stage are used to obtain a similarity matrix through dot product operation. After the similarity matrix is calculated by softmax, it is used to perform dot product calculation with the global representation. The result of the previous step is concatenated with the local features in the channel dimension and convolutional operation is used for dimension adjustment. The obtained result is the coupled local feature. The formula description of the calculation process is as follows: where local m and global m represent the coupled local features and global representations respectively, Conv represents the convolution operation, local and global represent the local features and global representations before coupling, and represent the dimensions of the local and global vectors, * represents the dot product operation; + represents the concatenation of feature map channels.

6. A 2D human pose estimation method based on the coupling of local features and global representations according to claim 5, wherein, the local-global feature coupling module further includes a local-global feature communication module, specifically: The image or feature map is divided into non-overlapping blocks, which are further subdivided within the blocks and encoded into different Patch Tokens. A Patch Token is a non-overlapping block formed after slicing the picture. All Patch Tokens belonging to the same block are downsampled to obtain local-global communication tokens, and the local-global communication tokens represent the high-level semantic information of this block; after equally shuffling the local-global communication tokens, each new local-global communication token contains the high-level semantic information of all blocks, and then the new local-global communication tokens are concatenated with the Patch Tokens of the original divided blocks in the channel dimension; after the above calculations, each block contains not only the semantic information representing itself, but also a small amount of high-level semantic information of other blocks.

7. A 2D human pose estimation method based on the coupling of local features and global representations according to claim 5, wherein, in step 4), calculating the loss between the output heat map of the network model and the corresponding annotated heat map, and the corresponding annotated heat map is converted from the joint positions of the original image, and training and optimizing the detection model, specifically including: where loss represents loss, num joints represents the number of human joint points labeled in the dataset, height and weight respectively represent the height and width of the heatmap, output and target respectively represent the heatmap calculated by the model and the heatmap representing the actual positions of human joint points; Calculating the loss between the heat map output by the model and the heat map converted from the true label using the cross-entropy function, and the size of the heat map is 56*56; since each joint is associated with a heat map, when calculating the loss, first calculate the loss between each joint, and then calculate the average value; calculating the loss between each pair of joint points is to calculate the cross-entropy sum of 3136 pairs of elements corresponding to each pair of heat maps.

Citation Information

Patent Citations

  • Three-dimensional human body posture estimation method fusing local and global features

    CN114565938A