Human body posture estimation method and device in dense posture scene, equipment and medium
Through the combination of deep learning backbone network and human key point decoder, the interbody occlusion problem in human posture estimation in dense posture scenarios is solved, and the estimation accuracy and feature extraction ability are improved.
Patent Information
- Application Number
- CN202510209260.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
In dense posture scenarios, existing multi-person human posture estimation methods are difficult to effectively deal with the problem of mutual occlusion between human bodies, resulting in difficulty in matching key points and low estimation accuracy.
The deep learning backbone network is adopted, including the image blocking module, the position embedding module and the encoder. The deep features of the human body's posture are extracted through the multi-head attention module and the optimized feedforward network module, and the key point decoder is used to generate the key point heat map, and the model is trained through the loss function.
The accuracy of human posture estimation in dense posture scenarios is improved, the interference of interbody occlusion on key point detection is reduced, and the ability to extract local and global features is enhanced.
Smart Images

Figure CN120220180A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision human pose estimation, and particularly to a human pose estimation method, device, equipment and medium in a dense pose scenario. Background Art
[0002] Human pose estimation refers to estimating the positions of human joints and their associations in an image or video, which is a popular research topic in the field of computer vision. It has wide applications in scenarios such as action analysis, human-computer interaction, healthcare, virtual reality, motion analysis, and computer-generated images. According to the spatial dimension of the output result, the monocular human pose estimation task can be divided into two categories: two-dimensional pose estimation and three-dimensional pose estimation. Two-dimensional pose estimation focuses on locating the two-dimensional coordinates of human key points (body joints) in an image or video frame, while three-dimensional pose estimation aims to predict depth information for a more accurate spatial representation.
[0003] In recent years, with the rapid development of deep learning and convolutional networks, significant progress and excellent performance have been achieved in using deep learning techniques for single-person human pose estimation tasks. However, in two-dimensional multi-person human pose estimation in a dense pose scenario, it is more challenging due to problems such as mutual occlusion between humans, truncation of humans in the image, and changes in human appearance.
[0004] Existing multi-person human pose estimation methods mainly include two directions: top-down and bottom-up. The former uses a human object detector to first obtain the prediction boxes of multiple humans in a dense pose scenario, and then performs single-person key point detection on each prediction box separately; the latter obtains the key point heat map through methods such as Gaussian filtering, gets all the human key point prediction results, and then re-classifies and combines these key points to obtain the complete human pose estimation result. There are also some single-stage multi-person human pose estimation methods that directly regress the final result from the overall feature vector. The above methods will all face the problem that it is difficult to match adjacent key points to the corresponding human instances due to severe mutual occlusion between the human instances to be detected in human pose estimation in a dense pose scenario. Summary of the Invention
[0005] To solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a human pose estimation method, device, equipment and medium in a dense pose scenario.
[0006] The first technical solution adopted by the present invention is:
[0007] A human pose estimation method in a dense pose scenario, comprising the following steps:
[0008] Obtain a dense pose scene dataset, perform data preprocessing and data augmentation to obtain human instance images;
[0009] Construct a deep learning backbone network, which includes an image chunking module, a position embedding module, and an encoder;
[0010] Input the human instance images into the deep learning backbone network initialized with pre-trained weights for feature extraction to obtain a feature set of the human instance images;
[0011] Input the feature set output by the deep learning backbone network into a human keypoint decoder for decoding to generate a predicted set of human keypoint heatmaps;
[0012] Calculate the loss function value, train the deep learning backbone network to obtain a final human pose estimation model for human pose estimation.
[0013] Furthermore, the working mode of the deep learning backbone network is as follows:
[0014] Input the human instance images into the deep learning backbone network. The image chunking module divides the input human instance images into a series of small blocks of a fixed size. Each small block is linearly expanded and represented as a token of a fixed length to obtain an image chunk sequence token;
[0015] Input the image chunk sequence token into the position embedding module to add corresponding sequence position encoding information to each token;
[0016] Input the image chunk sequence after position encoding into the encoder for image chunk relationship modeling and feature extraction.
[0017] Furthermore, the encoder includes a multi-head attention module and an optimized feed-forward network module; among them, the multi-head attention module models the tokens based on the feature similarity of the dense pose scene images to be estimated; the optimized feed-forward network module models the dense pose scene tasks to enhance the extraction ability of global and local features;
[0018] The feed-forward network module includes a local and global information enhanced feed-forward network unit and an expert feed-forward network unit.
[0019] Furthermore, the global information enhanced feed-forward network unit contains 2 branches. Branch 1 maintains a linear transformation, and branch 2 introduces convolution and channel attention. The specific calculation process expression is as follows:
[0020] Branch 1:
[0021] X out1 = Linear2(GELU(Linear1(Z in1 )))
[0022] Wherein, X in1 represents the sequence of image block tokens of the input branch 1, with dimensions B×tokens×dims; Linear1 represents the first linear connection layer in branch 1; GELU represents the GELU activation function; Linear2 represents the second linear connection layer in branch 1; X out1 represents the output of branch 1;
[0023] Branch 2:
[0024]
[0025] X cbh1 = Hardswish(BN(Conv 1×1 (X B×C×H×W )))
[0026] X dwcbh = Hardswish(BN(DWConv 3×3 (X cbh1 )))
[0027] X cbh2 = Hardswish(BN(Conv 1×1 (X dwcbh )))
[0028] X se = Sigmoid(Linear4(ReLU(Linear3(AvgPool(X cbh2 )))))
[0029] Wherein, X in2 represents the input of branch 2, X out2 represents the output of branch 2; reshape1 represents adjusting the dimension size of X in1 from B×tokens×dims to B×C×H×W; CBH1 represents the first CBH module, X cbh1 is the output after passing through the first CBH module; DWCBH represents the DWCBH module, X dwcbh is the output after passing through the DWCBH module; CBH2 represents the second CBH module, X cbh2 is the output after passing through the second CBH module; SE-Layer represents the channel attention layer, se is the output after passing through the channel attention layer; reshape2 represents adjusting the dimension size of X se from B×C×H×W to B×tokens×dims; Hardswish represents the Hard-Swish activation function; BN represents the batch normalization layer; Conv1×1 Denotes a 1×1 convolutional layer; DWConv 3×3 Denotes a 3×3 depthwise separable convolutional layer; Sigmoid denotes the Sigmoid activation function; ReLU denotes the ReLU activation function; Linar4 and Linear3 denote 2 linear connection layers; AvgPool denotes the global average pooling layer;
[0030] Fusion output:
[0031] X out = X out1 + X out2
[0032] In the formula, X out Is the output image patch sequence token.
[0033] Furthermore, the expert feed-forward network unit includes a main path and 2 branches. After passing through the main path, it is further divided into 2 branches; Branch 1 maintains a linear transformation, and Branch 2 introduces convolution and channel attention. Finally, the two branches are concatenated and output. The specific calculation process expression is as follows:
[0034] Main path:
[0035] X1 = GELU(Linear1(X in ))
[0036] In the formula, X in Is the input image patch sequence token; Linear1 represents the linear connection layer in the main path; GELU represents the GELU activation function;
[0037] Branch 1:
[0038] X out3 = Linear2(X1)
[0039] In the formula, X out3 Is the output of Branch 1; Linear2 represents the linear connection layer in Branch 1;
[0040] Branch 2:
[0041]
[0042] X cbh2 = Hardswish(BN(Conv 1×1 (X dwcbh )))
[0043] X se = Sigmoid(Linear2(ReLU(Linear1(AvgPool(X cbh2 )))))
[0044] Fusion output:
[0045] X out = concat(X out3 , X out4 )
[0046] Wherein, X out4 represents the output of branch 2, and X out is the sequence of tokens of the output image patches.
[0047] Furthermore, the obtaining of the dense pose scene dataset, followed by data preprocessing and data augmentation to obtain human instance pictures, includes:
[0048] Adjust the true human detection boxes marked on the input dense pose scene pictures, expand the height or width of the human detection boxes to a fixed aspect ratio, crop multiple human instance pictures from the pictures according to the marked true human detection boxes, and adjust the human instance pictures to a fixed size;
[0049] Perform data augmentation processing and Gaussian blur processing on the human instance pictures.
[0050] Furthermore, the human key point decoder includes 2 deconvolution modules and a 1×1 convolution module; the 2 deconvolution modules are used to gradually restore a higher-resolution feature map from a low-resolution feature map for subsequent generation of more detailed key point heat maps, and each deconvolution module can double the height and width of the input feature map; the 1×1 convolution module is used to adjust the number of output channels to the number of human key points marked in the dense pose dataset.
[0051] The second technical solution adopted by the present invention is:
[0052] A human pose estimation device in a dense pose scene, comprising:
[0053] A data acquisition module, configured to acquire a dense pose scene dataset, and perform data preprocessing and data augmentation to obtain human instance pictures;
[0054] A model construction module, configured to construct a deep learning backbone network, and the deep learning backbone network includes an image chunking module, a position embedding module, and an encoder;
[0055] A feature extraction module, configured to input the human instance pictures into the deep learning backbone network initialized with pre-trained weights for feature extraction to obtain a feature set of the human instance pictures;
[0056] A key point extraction module, configured to input the feature set output by the deep learning backbone network into the human key point decoder for decoding to generate a predicted set of human key point heat maps;
[0057] A model training module, configured to calculate a loss function value, train a deep learning backbone network, and obtain a final human pose estimation model.
[0058] The third technical solution adopted by the present invention is:
[0059] An electronic device, comprising a processor and a memory, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned human pose estimation method in a dense pose scenario.
[0060] The fourth technical solution adopted by the present invention is:
[0061] A computer-readable storage medium, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned human pose estimation method in a dense pose scenario.
[0062] The fifth technical solution adopted by the present invention is:
[0063] A computer program product or a computer program, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method.
[0064] The beneficial effects of the present invention are: Based on the parallel idea and the mixture-of-experts structure idea, the present invention enhances the extraction of local features and global features beneficial to key point detection by introducing a convolutional layer and a channel attention mechanism. In addition, by adopting the mixture-of-experts structure and adjusting the positions of the convolutional layer and the channel attention, while enhancing the model's ability to extract local features and global features beneficial to key point detection, the number of parameters of the model is not increased too much. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention, and those skilled in the art can also obtain other drawings according to these drawings without creative efforts.
[0066] Figure 1 It is a human pose estimation method in a dense pose scenario in an embodiment of the present invention;
[0067] Figure 2 It is a human key point annotation map of a dense pose data set in an embodiment of the present invention;
[0068] Figure 3 It is a deep learning backbone network model diagram in an embodiment of the present invention;
[0069] Figure 4 It is a network structure diagram of a local and global information enhanced feedforward network unit in an embodiment of the present invention;
[0070] Figure 5 It is a network structure diagram of an expert feedforward network unit in an embodiment of the present invention. Detailed implementation manners
[0071] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0072] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0073] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the sequence of the indicated technical features.
[0074] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.
[0075] In view of the existing technical problems, the present invention proposes a method for multi-person human pose estimation in a dense pose scenario based on a top-down paradigm. The method first divides the input image into a series of small blocks of a fixed size through an image chunking module, and then captures deeper relationships and features between the image chunks through an encoder. The tokens are modeled based on the feature similarity of the dense pose scenario image to be estimated through a multi-head self-attention module, and the dense pose scenario task is modeled through an optimized feed-forward network module. The first two modules are simply concatenated alternately multiple times to form an encoder, which can better extract the deep individual features of human pose estimation. Then, the deep individual features are input into a decoder module to obtain a prediction result containing heatmaps of multiple human key points. Then, the value of the loss function is calculated based on the prediction result, and the parameters of the model are learned and adjusted through the value of the loss function. After multiple rounds of iteration, finally, a multi-person human pose estimation model in a dense pose scenario is obtained. Through the combined action of the above several modules, the present invention can fully decouple the multi-person images in the dense pose scenario into the feature expressions of different individual bodies, reduce the interference of mutual occlusion and overlap between different humans on the detection of human key points, more accurately estimate the poses of each human body, and improve the accuracy of human pose estimation of the system in the dense pose scenario. The multi-person human pose estimation model can be applied to scenarios such as video surveillance, sports competitions, human-computer interaction, medical rehabilitation, and autonomous driving, but is not limited thereto.
[0076] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0077] Embodiment 1
[0078] As Figure 1 shown, this embodiment provides a method for human pose estimation in a dense pose scenario, which improves the feature extraction ability of the multi-person human pose estimation model for human key points in a dense pose scenario, and further improves the detection effect of human key points. The method includes the following steps:
[0079] S1. Obtain a dense pose scene dataset, perform data preprocessing and data augmentation to obtain human instance images.
[0080] In some embodiments, the dense pose scene dataset is the Crowdpose dataset. The Crowdpose dataset proposes a definition of crowding degree for scenes with severe crowd occlusion. The dataset includes 20,000 original images containing human poses, as Figure 2 shown. Figure 2 Figure 7 is a schematic diagram of key point annotation adopted by the Crowdpose dataset, in which 14 key points are annotated for each human body, including the head, left and right shoulders, left and right elbows, left and right wrists, left and right hip joints, left and right knees, left and right ankles, and the body center point.
[0081] Divide the Crowdpose dataset into a training set and a validation set, and then perform preprocessing and data augmentation on the training set. As an implementation method, the data preprocessing and data augmentation steps include:
[0082] S11. Adjust the real human detection box annotated in the input dense pose scene image, expand the height or width of the human detection box to a fixed aspect ratio: height: width = 4:3, and then crop multiple human instance images from the image according to the annotated real human detection box, and adjust them to a fixed size, 256×192 or 384×288;
[0083] S12. Randomly rotate the human instance image, and the rotation angle range is ([ - 45°, 45°]);
[0084] S13. Randomly scale the human instance image according to the center point, and the scaling scale range is ([0.65, 1.35]);
[0085] S14. Randomly horizontally flip the human instance image, and the random factor is 0.5;
[0086] S15. Randomly perform half-body data augmentation, and the random factor is 0.7;
[0087] S16. Perform Gaussian blur and median filter blur on the human instance image.
[0088] Among them, the validation set only needs to perform the data preprocessing of step S11.
[0089] S2. Construct a deep learning backbone network, and the deep learning backbone network includes an image chunking module, a position embedding module, and an encoder.
[0090] Refer to Figure 3, the backbone network consists of an Image Patch Embedding module, a Position Embedding module, and an Encoder. Among them, the Encoder is composed of multiple identical EncoderLayer modules connected in series. The EncoderLayer module includes a Layer norm module, a Residual connection module, a Multi-Head Attention module, and an optimized Feed-Forward Network module (FFN′).
[0091] The specific data processing process is as follows:
[0092] The human body instance image passes through the Image Patch Embedding module, which is responsible for dividing the input image into a series of small blocks of a fixed size. Each small block is linearly expanded and represented as a token of a fixed length, and then an output sequence is formed. Each sequence item corresponds to the linear vectorized representation of an image block. For example: if the input image size is H×W×3 and each small block size is P×P, then the image is divided into small blocks. After each small block is linearly expanded, it becomes a token with Len = P×P×3. N small blocks form a sequence, and the size of the sequence can be represented by N×Len.
[0093] The output sequence then passes through the Position Embedding module, which adds the corresponding sequence position encoding information to each token, and then enters the Encoder for image block relationship modeling and feature extraction.
[0094] The sequence input to the Encoder module is divided into two paths. The first path passes through the 1st Layer norm module, which can ensure the gradient fluidity in the deep learning process, and then passes through the Multi-Head Attention module, which models the tokens based on the feature similarity of the human body instance image. The first path and the second path are connected and output through the 1st Residual connection module; then it continues to be divided into two paths. The first path passes through the 2nd Layer norm module and then enters the optimized Feed-Forward Network module, which can model for the dense pose scene task after optimization and better extract global and local features, and then is connected and output with the second path through the 2nd Residual connection module. The data process of each EncoderLayer module is the same as the above description.
[0095] Specifically, referring to Figure 4 , Figure 4It is the network structure diagram of the local and global information enhanced feed-forward network unit. This unit is an optimized feed-forward network module. Based on the idea of parallelism, depthwise separable convolution and the lightweight attention module SE-Layer (Squeeze-and-Excitation Layer) are introduced into the feed-forward network module. The convolutional layer and SE-Layer can enhance the global feature and local feature extraction capabilities of the feed-forward network module for the human pose estimation task in the dense pose scenario, making the global features and local features of the key points richer.
[0096] The local and global information enhanced feed-forward network unit consists of 2 parallel modules. Branch 1 includes 2 linear layers (Linear). The first linear layer performs feature expansion, and the default feature expansion multiple is set to 4. Then it is processed by the GELU non-linear activation function, and finally the second linear layer performs feature recovery to restore to the feature dimension number when inputting the first linear layer.
[0097] Branch 2 is successively the CBH module, DWCBH module, CBH module, and SE-Layer module. The first CBH module refers to a 1×1 convolutional layer, a batch normalization layer, and a Hard-Swish activation function, which expands the number of channels through 1×1 convolution, and the default expansion multiple is set to 4. The DWCBH module refers to a 3×3 depthwise separable convolutional layer with a stride of 1 and a padding of 1, a batch normalization layer, and a Hard-Swish activation function. The second CBH module refers to a 1×1 convolutional layer, a batch normalization layer, and a Hard-Swish activation function, which reduces the number of channels to 1 / 4 of the number of channels of the tensor input to this module. The SE-Layer module consists of a global average pooling layer and 2 linear layers.
[0098] The specific calculation process expression of the local and global information enhanced feed-forward network unit is as follows:
[0099] Branch 1:
[0100] X out1 = Linear2(GELU(Linear1(X in1 )))
[0101] In the formula, X in1 represents the image patch sequence token input to Branch 1, with a dimension of B×tokens×dims; Linear1 represents the first linear connection layer in Branch 1; GELU represents the GELU activation function; Linear2 represents the second linear connection layer in Branch 1; X out1 represents the output of Branch 1.
[0102] Branch 2:
[0103]
[0104] X cbh1 = Hardswish(BN(Conv 1×1 (X B×C×H×W )))
[0105] X dwcbh = Hardswish(BN(DWConv 3×3 (X cbh1 )))
[0106] X cbh2 = Hardswish(BN(Conv 1×1 (X dwcbh )))
[0107] X se = Sigmoid(Linear4(ReLU(Linear3(AvgPool(X cbh2 )))))
[0108] In the formula, X in2 represents the input of branch 2, and X out2 represents the output of branch 2; reshape1 means to adjust the dimension size of X in1 from B×tokens×dims to B×C×H×W; CBH1 represents the 1st CBH module, and X cbh1 is the output after passing through the 1st CBH module; DWCBH represents the DWCBH module, and X dwcbh is the output after passing through the DWCBH module; CBH2 represents the 2nd CBH module, and X cbh2 is the output after passing through the 2nd CBH module; SE-Layer represents the channel attention layer, and X se is the output after passing through the channel attention layer; reshape2 means to adjust the dimension size of X se from B×C×H×W to B×tokens×dims; Hardswish represents the Hard-Swish activation function; BN represents the batch normalization layer; Conv 1×1 represents the 1×1 convolutional layer; DWConv 3×3 represents the 3×3 depthwise separable convolutional layer; Sigmoid represents the Sigmoid activation function; ReLU represents the ReLU activation function; Linear4 and Linear3 represent 2 linear connection layers; AvgPool represents the global average pooling layer.
[0109] Fusion output:
[0110] X out = X out1 + X out2
[0111] In the formula, X out is the output image block sequence token.
[0112] Another optimized feed-forward network module, referring to Figure 5 , Figure 5 is the network structure diagram of the expert feed-forward network unit. Based on the idea of mixture-of-experts (MoE), this unit introduces depthwise separable convolution and the lightweight attention module SE-Layer (Squeeze-and-Excitation Layer) into the feed-forward network module. The convolutional layer and SE-Layer can enhance the global and local feature extraction capabilities of the feed-forward network module for the human pose estimation task in dense pose scenarios, making the global and local features of key points richer.
[0113] On the basis of the local and global information enhanced feed-forward network unit's first path, the expert feed-forward network unit retains the first linear layer of the first path as the main path, and splits the second linear layer into two paths, namely the shared expert module and the information enhanced expert module. The shared expert is actually also a linear layer, but the output dimension becomes the specified output dimension × 0.75, while the output dimension of the information enhanced expert module is the specified output dimension × 0.25. Then the outputs of the two paths are concatenated into the specified output dimension size.
[0114] Specifically, the local information expert module, that is, the second branch, is successively the DWCBH module, the CBH module, and the SE-Layer module. The DWCBH module refers to a 3×3 depthwise separable convolutional layer with a stride of 1 and a padding of 1, a batch normalization layer, and a Hard-Swish activation function. The CBH module refers to a 1×1 convolutional layer, a batch normalization layer, and a Hard-Swish activation function, which reduces the number of channels by 1×1 convolution to the specified output dimension × 0.25. The SE-Layer module consists of a global average pooling layer and two linear layers.
[0115] The specific calculation process expression of the expert feed-forward network unit is as follows:
[0116] Main path:
[0117] X1 = GELU(Linear1(X in ))
[0118] In the formula, X in is the input image block sequence token; Linear1 represents the linear connection layer in the main path; GELU represents the GELU activation function.
[0119] Branch 1:
[0120] X out3 = Linear2(X1)
[0121] In the formula, X out3 is the output of branch 1; Linear2 represents the linear connection layer in branch 1.
[0122] Branch 2:
[0123]
[0124] X cbh2 = Hardswish(BN(Conv 1×1 (X dwcbh )))
[0125] X se = Sigmoid(Linear4(ReLU(Linear3(AvgPool(X cbh2 )))))
[0126] Fusion output:
[0127] X out = concat(X out3 , X out4 )
[0128] In the formula, X out4 represents the output of branch 2, and X out is the sequence token of the output image patch.
[0129] S3. Input the human instance picture into the deep learning backbone network initialized with pre-trained weights for feature extraction to obtain the feature set of the human instance picture.
[0130] Exemplarily, the weights used for initializing the deep learning backbone network are pre-trained on the ImageNet-1k dataset using the Masked Autoencoders (MAE) self-supervised method.
[0131] Among them, ImageNet-1k is a large visual database for visual object recognition software research. It contains more than 14 million images, and these images are manually annotated to represent the objects in the pictures, and bounding boxes are provided in at least one million images
[0132] S4. Input the feature set output by the deep learning backbone network into the human key point decoder for decoding to generate the human key point heat map prediction set.
[0133] In some embodiments, the human key point decoder mainly includes two transposed convolution modules and a 1×1 convolution module. The two transposed convolution modules gradually restore a higher-resolution feature map from a low-resolution feature map to generate a more detailed key point heat map subsequently. Each transposed convolution module can double the height and width of the input feature map. The final 1×1 convolution module adjusts the number of output channels to the number of human key points annotated in the dense pose dataset.
[0134] S5. Calculate the loss function value, train the deep learning backbone network, and obtain a final human pose estimation model for human pose estimation.
[0135] In some embodiments, during the training process, the trained deep learning network is evaluated according to the validation set. The AdamW optimizer is used to update the network parameters and the polynomial decay learning rate strategy is used to update the learning rate. The training is iterated for 270 epochs. The model weights with the smallest loss function value on the validation set are tested and retained every 10 epochs to obtain a final human pose estimation model.
[0136] Embodiment 2
[0137] This embodiment provides a human pose estimation device in a dense pose scenario, including:
[0138] A data acquisition module for acquiring a dense pose scenario dataset, performing data preprocessing and data augmentation to obtain human instance pictures;
[0139] A model construction module for constructing a deep learning backbone network, which includes an image chunking module, a position embedding module, and an encoder;
[0140] A feature extraction module for inputting the human instance pictures into the deep learning backbone network initialized with pre-trained weights for feature extraction to obtain a feature set of the human instance pictures;
[0141] A key point extraction module for inputting the feature set output by the deep learning backbone network into a human key point decoder for decoding to generate a predicted set of human key point heat maps;
[0142] A model training module for calculating the loss function value and training the deep learning backbone network to obtain a final human pose estimation model.
[0143] Since this device is a human pose estimation device in a dense pose scenario of an embodiment of the present invention, and the principle of solving problems by this device is similar to that of this method, the implementation of this device can refer to the implementation process of the above method embodiment, and the repeated parts will not be elaborated.
[0144] Embodiment 3
[0145] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a human pose estimation method in a dense pose scenario as shown in Figure 1 Figure.
[0146] It can be understood that the memory may include a random access memory (RAM), or may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above method embodiments, etc.; the data storage area can store data created according to the use of the server, etc.
[0147] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a central processing unit (CPU) and a modem, etc. in a combination of one or several. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a single chip.
[0148] Since this electronic device is an electronic device corresponding to a human pose estimation method in a dense pose scenario in an embodiment of the present invention, and the principle of the electronic device to solve problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0149] Example 4
[0150] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the method for human pose estimation in a dense pose scenario as Figure 1 shown.
[0151] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data.
[0152] Since this storage medium is the storage medium corresponding to the method for human pose estimation in a dense pose scenario of an embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0153] Example 5
[0154] In some possible embodiments, various aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a human pose estimation method in a dense pose scenario according to various exemplary embodiments described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0155] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0156] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0157] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for estimating human posture in a dense posture scene, characterized in that: The following steps are involved: Obtain a dense posture scene dataset, perform data preprocessing and data enhancement, and obtain human body instance images; Constructing a deep learning backbone network, wherein the deep learning backbone network includes an image segmentation module, a position embedding module and an encoder; Input the human body instance image into the deep learning backbone network initialized with pre-trained weights to extract features and obtain a feature set of the human body instance image; The feature set output by the deep learning backbone network is input into the human key point decoder for decoding to generate a human key point heat map prediction set; The deep learning backbone network is trained to obtain the final human posture estimation model.
2. The method for estimating human body posture in a dense posture scene according to claim 1, characterized in that: The working mode of the deep learning backbone network is: The human body instance image is input into the deep learning backbone network. The image segmentation module segments the input human body instance image into a series of small blocks of fixed size. Each small block is linearly expanded and represented as a fixed-length tag to obtain the image segmentation sequence tag. The image block sequence tag input position embedding module adds the corresponding sequence position encoding information to each tag; The position-encoded image block sequence is input into the encoder for image block relationship modeling and feature extraction.
3. The method for estimating human body posture in a dense posture scene according to claim 2, characterized in that: The encoder includes a multi-head attention module and a feedforward network module; wherein the multi-head attention module models the tags based on the feature similarity of the dense posture scene image to be estimated; the feedforward network module models the dense posture scene task to enhance the extraction capability of global features and local features; The feedforward network module includes a local and global information enhanced feedforward network unit and an expert feedforward network unit.
4. The method for estimating human body posture in a dense posture scene according to claim 3, characterized in that: The global information enhanced feedforward network unit includes two branches, branch 1 maintains linear transformation, and branch 2 introduces convolution and channel attention. The specific calculation process expression is as follows: Branch 1: X out1 =Linear2(GELU(Linear1(X in1 ))) Where, X in1 It represents the image block sequence token of input branch 1, with the dimension of B×tokens×dims; Linear1 represents the linear connection layer No. 1 in branch 1; GELU represents the GELU activation function; Linear2 represents the linear connection layer No. 2 in branch 1; x out1 Represents the output of branch 1; Branch 2: x cbh1 =Hardswigh(BN(Conv 1×1 (X B×C×H×W ))) X dwcbh =Hardswish(BN(DWConv 3×3 (X cbh1 ))) X cbh2 =Hardswish(BN(Conv 1×1 (X dwcbh ))) X se =Sigmoid(Linear4(ReLU(Linear3(AvgPool(X cbh2 ))))) Where, X in2 Indicates the input of branch 2, X out2 Represents the output of branch 2; reshape1 represents X in1 The dimension size is adjusted from B×tokens×dims to B×C×H×W; CBH1 represents CBH module No. 1, X cbh1 is the output after passing through CBH module No. 1; DWCBH represents DWCBH module, X dwcbh is the output after passing through the DWCBH module; CBH2 represents CBH module No. 2, X cbh2 is the output after passing through the CBH module No. 2; SE-Layer represents the channel attention layer, X se is the output after the channel attention layer; reshape2 means to reshape X se The dimension size is adjusted from B×C×H×W to B×tokens×dims; Hardswish represents the Hard-Swish activation function; BN represents the batch normalization layer; Conv 1×1 Represents a 1×1 convolutional layer; DWConv 3×3 Represents a 3×3 depth-wise separable convolutional layer; Sigmoid represents a Sigmoid activation function; ReLU represents a ReLU activation function; Linear4 and Linear3 represent two linear connection layers; AvgPool represents a global average pooling layer; Fusion output: X out =X out1 +X out2 In the formula, X out It is the output image block sequence token.
5. The method for estimating human posture in a dense posture scene according to claim 3, characterized in that: The expert feedforward network unit includes a trunk and two branches. After the trunk, it is further subdivided into two branches. Branch 1 maintains linear transformation, branch 2 introduces convolution and channel attention, and finally the two branches are spliced and output. The specific calculation process expression is as follows: Trunk Road: X1=YELLOW(Linear1(X in )) Where, X in is the input image block sequence token; Linear1 represents the linear connection layer in the trunk; GELU represents the GELU activation function; Branch 1: X out3 =Linear2(X1) In the formula, X out3 is the output of branch 1; Linear2 represents the linear connection layer in branch 1; Branch 2: X cbh2 =Hardswish(BN(Conv 1×1 (X dwcbh ))) X se =Sigmoid(Linear4(ReLU(Linear3(AvgPool(X cbh2 ))))) Fusion output: X out =concat(X out3 ,X out4 ) In the formula, X out4 represents the output of branch 2, X out It is the output image block sequence token.
6. The method for estimating human posture in a dense posture scene according to claim 1, characterized in that: The method of obtaining a dense posture scene data set, performing data preprocessing and data enhancement, and obtaining a human body instance image includes: The real human body detection frame annotated in the input dense posture scene image is adjusted, the height or width of the human body detection frame is expanded to a fixed aspect ratio, and multiple human body instance images are cropped from the image according to the annotated real human body detection frame. And adjust the human body instance image to a fixed size; Perform data enhancement and Gaussian blur processing on human body instance images.
7. The method for estimating human posture in a dense posture scene according to claim 1, characterized in that: The human key point decoder includes two deconvolution modules and a 1×1 convolution module; the two deconvolution modules are used to gradually restore higher resolution feature maps from low resolution feature maps so as to subsequently generate more detailed key point heat maps; the 1×1 convolution module is used to adjust the number of output channels to the number of human key points annotated by the dense posture data set.
8. A human body posture estimation device in a dense posture scene, characterized in that: include: The data acquisition module is used to acquire dense posture scene data sets, perform data preprocessing and data enhancement, and obtain human body instance images; A model building module, used to build a deep learning backbone network, wherein the deep learning backbone network includes an image segmentation module, a position embedding module and an encoder; A feature extraction module is used to input a human body instance image into a deep learning backbone network initialized with pre-trained weights to perform feature extraction and obtain a feature set of the human body instance image; The key point extraction module is used to input the feature set output by the deep learning backbone network into the human key point decoder for decoding, and generate a human key point heat map prediction set; The model training module is used to train the deep learning backbone network to obtain the final human posture estimation model.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.