A human body reconstruction method and device based on attention mechanism

By introducing attention mechanisms into the human body reconstruction network, the problem of difficulty in accurately reconstructing a three-dimensional mannequin in the presence of occlusion in the prior art is solved, and a more accurate reconstruction effect of human body posture and shape is achieved.

CN114067057BActive Publication Date: 2025-05-23ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111382077.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-22
Publication Date
2025-05-23
Estimated Expiration
2041-11-22

AI Technical Summary

Technical Problem

The prior art is difficult to accurately reconstruct a three-dimensional mannequin model with accurate posture and shape based on single human images of obscured bodies, especially when deep learning networks cannot effectively distinguish key information from redundant information.

Method used

The human body reconstruction method based on attention mechanism is adopted to construct a human body reconstruction network model including feature extraction module, attention module, fusion module, parameter inference module and SMPL submodule. The attention module generates attention maps through average pooling and maximum pooling operations, the fusion module generates body attention feature maps through multiplication of attention maps and the original feature maps, and the parameter inference module generates SMPL parameters based on these feature maps, and finally generates a three-dimensional human body model through the SMPL submodule.

Benefits of technology

Through the application of attention mechanism, the network can focus more effectively on important feature information in the image, reduce attention to unimportant information, thereby reducing the interference of occlusion on the network, and improving the accuracy of the posture and shape of the three-dimensional mannequin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067057B_ABST
    Figure CN114067057B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision, and specifically relates to a method, model, and device for human body reconstruction based on an attention mechanism. The reconstruction method comprises the following steps: Step 1: construct a human body reconstruction network model, the human body reconstruction network model comprises a feature extraction module, an attention module, a fusion module, a parameter inference module, and an SMPL submodule; Step 2: obtain a plurality of original images containing characters, pre-process the original images to form a training data set; Step 3: use the training data set of the above step to train the human body reconstruction network model by minimizing the network loss function; Step 4: input the human body image to be processed into the trained network model after pre-processing, and generate a three-dimensional human body model with a specific posture. The present invention solves the problem that it is difficult for existing methods to accurately reconstruct a three-dimensional human body model with accurate posture and morphology based on a single human body image with occlusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a method and device for human body reconstruction based on an attention mechanism. Background Art

[0002] Virtual reality technology is an emerging artificial intelligence technology and has been widely used in scenarios such as virtual fitting, body animation, and human motion simulation games. In the application of these technologies, using images to model the human body in three dimensions is an important link. Human body modeling is a core issue in computer vision and graphics. Existing methods for reconstructing human body three-dimensional models from images mainly include two categories, namely optimization-based methods and regression-based methods. The former fits the parameterized body model to the two-dimensional observation of a given image through an iterative optimization process, focusing on the use of two-dimensional joint position and contour to achieve the fitting and modeling process. The latter mainly builds a deep learning network, and extracts features from a single input image in the deep neural network to obtain human body model parameters, volume representation of the three-dimensional human body, model vertices and other information; the above information is used to generate a three-dimensional human body model.

[0003] The two methods mentioned above both have good model reconstruction effects when there are no obstructions or the obstructions are not obvious for the target person in the image. However, in practical applications, it is very common for the target person in the image to be obstructed by other people or objects; therefore, the application of the above methods has limitations. In particular, when using deep learning networks for 3D model reconstruction, deep neural networks cannot effectively distinguish between key information and redundant information in human images, but predict the parameters of the 3D model based on all pixel features in the human image. As a result, obvious errors will occur, and obstructions will seriously interfere with the actual 3D human model, resulting in the human posture and shape in the constructed 3D model not being consistent with the actual one. Summary of the invention

[0004] In order to solve the problem that the existing human body three-dimensional model reconstruction method is difficult to accurately reconstruct a three-dimensional human body model with accurate posture and shape based on a single human body image with occlusion, a human body reconstruction method and device based on an attention mechanism are provided.

[0005] The present invention is implemented by the following technical solutions:

[0006] A human body reconstruction method based on an attention mechanism, the human body reconstruction method comprising the following steps:

[0007] Step 1: Construct a human body reconstruction network model, which includes a feature extraction module, an attention module, a fusion module, a parameter inference module and a SMPL submodule. The feature extraction module is used to generate the corresponding original feature map according to the input human body image. The attention module includes two pooling layers, a convolutional layer and a Sigmoid operation layer; the two pooling layers are the average pooling layer and the maximum pooling layer. The attention module is used to generate an attention map according to the input original feature map. The fusion module is used to fuse the original feature map and the attention map to obtain a body attention feature map. The parameter inference module includes a pooling layer and three fully connected layers; the parameter inference module is used to generate the SMPL parameters of the corresponding target person in the human body image according to the input body attention feature map. The SMPL submodule is used to generate a three-dimensional human body model of the corresponding target person according to the SMPL parameters.

[0008] Step 2: Acquire multiple human images containing the target person as original images, pre-process the original images to form a training data set, and the original images in the training data set at least include some human images that are occluded by the person.

[0009] Step 3: Use the training data set in the previous step to train the human body reconstruction network model by minimizing the network loss function.

[0010] Step 4: Save the trained human body reconstruction network model; input the human body image to be processed into the saved network model after preprocessing to generate a three-dimensional human body model with a specific posture.

[0011] As a further improvement of the present invention, the feature extraction module is obtained by simplifying and repackaging the deep convolutional neural network Resnet50, and the simplification process only retains the convolution part in the original network model; the input human body image is convolved by the feature extraction module to obtain the original feature map.

[0012] As a further improvement of the present invention, the attention module takes the output of the feature extraction module as input. The input original feature map first passes through the average pooling layer and the maximum pooling layer in the attention module. The two pooling results are concatenated and then sequentially subjected to convolution processing and Sigmoid operation to obtain the attention map.

[0013] In the attention module, the pooling operation formula of the average pooling layer is:

[0014] F avg =AvgPool(F);

[0015] The pooling operation formula of the maximum pooling layer is:

[0016] F max =MaxPool(F);

[0017] In the above formula, F represents the original feature map, F avg represents the feature map after the average pooling operation, F max represents the feature map after the maximum pooling operation, MaxPool(·) represents the maximum pooling operation, and AvgPool(·) represents the average pooling operation.

[0018] The generation formula of the attention map is:

[0019] M(F)=σ(f(cat(F avg ,F max )));

[0020] In the above formula, M(F) represents the attention map; σ(·) represents the Sigmoid activation function; f(·) represents the convolution operation; cat(·) represents the concatenation operation of the feature map.

[0021] As a further improvement of the present invention, in the fusion module, the fused body attention feature map is obtained by performing corresponding element multiplication operation on the attention map and the original feature map. The formula of the fusion operation is:

[0022]

[0023] In the above formula, F′ represents the body attention feature map, and M(F) represents the attention map; represents the multiplication operation according to the corresponding elements; F represents the original feature map.

[0024] As a further improvement of the present invention, the pooling layer in the parameter inference module is an average pooling layer. The first two of the three fully connected layers each have 1024 neurons and are connected by a Dropout operation. The third fully connected layer has 85 neurons and is directly connected to the previous fully connected layer. Among them, the three fully connected layers constitute the iterative regression part in the parameter inference module.

[0025] As a further improvement of the present invention, in the parameter inference module, the generation process of SMPL parameters is as follows:

[0026] (1) The input body attention feature map F′ is average pooled to obtain a feature φ.

[0027] (2) Putting together the SMPL pose parameter θ, shape parameter β and camera parameter c, the formula is expressed as:

[0028] Θ=cat(θ,β,c);

[0029] In the above formula, θ represents the pose parameter of the SMPL model; β represents the shape parameter of the SMPL model; c represents the camera parameter; Θ represents the concatenated parameter set of the pose parameter θ, the shape parameter β and the camera parameter c.

[0030] (3) Use the average pose parameters, average shape parameters and average camera parameters to form the initialization parameter set θ 0 , the feature φ and the parameter set Θ 0 The concatenation is used as the input to the iterative regression part in the parameter inference module.

[0031] (4) Generate the residual of the parameter set corresponding to the current input, and then update the current parameter set. The update formula is:

[0032] Θ t+1 =Θ t +ΔΘ t ;

[0033] In the above formula, Θ t Represents the parameter set corresponding to the current input, Θ t+1 Represents the parameter set Θ t The updated state, ΔΘ t Represents the parameter set Θ t The residual.

[0034] (5) Iterate the update operation of the previous step three times; during each iterative update process, the parameter set obtained in the previous update is concatenated with the feature φ as the input of the iterative regression part of this parameter inference module to update the parameter set.

[0035] (6) After the iterative operation is completed, the SMPL parameters including the final posture parameter θ, morphological parameter β and the corresponding camera parameter c are obtained.

[0036] As a further improvement of the present invention, in the SMPL submodule, the SMPL parameters are input into the SMPL function to obtain a three-dimensional human body model; the expression of the SMPL function M(β,θ) is:

[0037]

[0038] In the above formula, B is the vertex coordinate of the 3D human body model in T posture; P (θ) and B s (β) represents the offset of the vertex vector relative to the SMPL standard template caused by the pose parameter θ and the morphological parameter β; J(β) is the joint point position of the model corresponding to the morphological parameter β; W(·) is the linear mixed skinning function; is the mixing weight.

[0039] As a further improvement of the present invention, the preprocessing process of the original image includes:

[0040] (1) Locate the target person in the human image and crop the image so that the target person is located in the central area of ​​the human image.

[0041] (2) The size of the cropped human body image is adjusted, and the pixel value of the adjusted image is unified to 224×224.

[0042] (3) Normalize the adjusted image to obtain the data element in the training data set.

[0043] As a further improvement of the present invention, during the training process of the network model, under the condition of minimizing the loss function, the Adam algorithm is used to adjust all parameters in the network model to train the network.

[0044] The expression for minimizing the loss function is as follows:

[0045] L=λ 2D L 2D joint +λ 3D L 3D joint +λ para L SMPL ;

[0046] In the above formula, L 2D joint represents the 2D joint loss function; L 3D joint represents the 3D joint point loss function; L SMPL represents the SMPL parameter loss function; λ 2D Represents the weight coefficient of the 2D joint loss function; λ 3D Represents the weight coefficient of the 3D joint loss function; λ para Represents the weight coefficient of the SMPL parameter loss function.

[0047] Among them, the 2D joint point loss function L 2D joint The expression is:

[0048]

[0049] In the above formula, v i Indicates the visibility of the i-th 2D joint point, with a value of 0 or 1, 0 means invisible, and 1 means visible; N represents the number of 2D joint points; represents the predicted value of the i-th 2D joint point; k i represents the true value of the i-th 2D joint point; where the predicted value of the 2D joint It is obtained by projecting the predicted 3D joint points.

[0050] 3D joint point loss function L 3D joint The expression is:

[0051]

[0052] In the above formula, M represents the number of images involved in the 3D joint point calculation; represents the 3D joint point prediction value of the i-th image; J i Represents the true value of the 3D joint point of the i-th image.

[0053] SMPL parameter loss function L SMPL The expression is:

[0054]

[0055] In the above formula, O represents the number of images involved in the SMPL parameter calculation. and Respectively represent the predicted values ​​of posture parameters and morphological parameters of the i-th image, θ i and β i They represent the true values ​​of the pose parameters and morphological parameters of the i-th image respectively.

[0056] The present invention also includes a human body reconstruction model, and the aforementioned human body reconstruction method based on the attention mechanism uses the human body reconstruction model to process the input human body image with occlusion, thereby generating a three-dimensional human body model of the target task in the human body image. The human body reconstruction model includes the following: a preprocessing module, a feature extraction module, an attention module, a fusion module, a parameter inference module, and an SMPL submodule.

[0057] The preprocessing module is used to: (1) locate the target person in the human image and crop the image so that the target person is located in the central area of ​​the human image; (2) adjust the size of the cropped human image, and unify the pixel values ​​of the adjusted image to 224×224; (3) normalize the adjusted image.

[0058] The feature extraction module uses the convolution part of the deep convolutional neural network Resnet50 as the backbone network. The output of the preprocessing module is used as the input of the feature extraction module; the feature extraction module is used to extract the features of the preprocessed human body image through convolution operations, and then generate the corresponding original feature map.

[0059] The attention module includes a maximum pooling submodule, an average pooling submodule, a feature splicing submodule, a convolution submodule, and a sigmoid operation submodule. The output of the feature extraction module is used as the input of the attention module; the original feature map is processed by the maximum pooling submodule and the average pooling submodule in the attention module to obtain two feature maps, and the two feature maps are feature spliced ​​in the feature splicing submodule; after the convolution processing in the convolution submodule and the sigmoid operation in the sigmoid operation submodule, the attention map is obtained.

[0060] The fusion module uses the original feature map output by the feature extraction module and the attention map output by the attention module, and then multiplies the corresponding elements of the original feature map and the attention map to obtain the fused body attention feature map.

[0061] The parameter inference module includes an average pooling layer, a fully connected layer 1, a fully connected layer 2, and a fully connected layer 3. Among them, the fully connected layer 1 and the fully connected layer 2 have 1024 neurons and are connected through the Dropout operation; the fully connected layer 3 has 85 neurons, and the fully connected layer 2 is directly connected to the fully connected layer 3. The fully connected layer 1, the fully connected layer 2, and the fully connected layer 3 constitute the iterative regression part of the network model. The output of the fusion module is used as the input of the parameter inference module; the parameter inference module generates iteratively updated SMPL parameters based on different input data.

[0062] The SMPL submodule is used to generate a three-dimensional human body model of the target person corresponding to the human body image according to the SMPL parameters output by the parameter inference submodule.

[0063] The technical solution provided by the present invention has the following beneficial effects:

[0064] In the three-dimensional human body reconstruction method based on the attention mechanism provided by the present invention, the introduced attention mechanism can process the features in the human body image, allowing the network to focus on the features containing important information, ignoring and reducing the attention to unimportant information. Among them, the original feature map is weighted by the attention map generated by the attention mechanism, so that the network focuses on the information related to the human body part in the image, reduces the attention to other information, and thus reduces the interference of the occlusion information on the network. At the same time, the network model can use the features of the visible part of the body to infer the situation of the occluded body part, thereby ensuring the integrity of the extracted characteristic information. It is ensured that the human posture and shape reflected in the finally constructed three-dimensional human body model are more in line with reality. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a flowchart of the steps of a human body reconstruction method based on the attention mechanism in Example 1 of the present invention.

[0066] Figure 2 Schematic diagram of the attention module in Example 1 of the present invention.

[0067] Figure 3 This is a flowchart of the steps of the process of generating SMPL parameters of a target person in the parameter inference module in Example 1 of the present invention.

[0068] Figure 4 This is a flowchart of the steps of the three-dimensional human body model reconstruction process in Example 1 of the present invention.

[0069] Figure 5 This is a module schematic diagram of a human body reconstruction model provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0071] Example 1

[0072] This embodiment provides a human body reconstruction method based on the attention mechanism. Figure 1 As shown, the human body reconstruction method comprises the following steps:

[0073] S1: Build a human body reconstruction network model, which includes feature extraction module, attention module, fusion module, parameter inference module and SMPL submodule. Specifically, it includes the following steps:

[0074] S11: The deep convolutional neural network Resnet50 is streamlined, and the streamlining process only retains the convolution part in the original network model. The streamlined network is then repackaged to obtain the required feature extraction module; the feature extraction module is used to generate the corresponding original feature map based on the input human body image.

[0075] S12: Construct an attention module including two pooling layers, a convolution layer and a sigmoid operation layer. The two pooling layers in the attention module are the average pooling layer and the maximum pooling layer. The attention module is used to generate an attention map based on the original input feature map. Figure 2 As shown in Figure 2, the processing of the attention module is as follows:

[0076] The input original feature map first passes through the average pooling layer and the maximum pooling layer in the attention module. The two pooling results are concatenated and then successively subjected to convolution processing and Sigmoid operation to obtain the attention map.

[0077] Among them, in the attention module, the pooling operation formula of the average pooling layer is:

[0078] F avg =AvgPool(F);

[0079] The pooling operation formula of the maximum pooling layer is:

[0080] F max =MaxPool(F);

[0081] In the above formula, F represents the original feature map, F avg represents the feature map after the average pooling operation, F max represents the feature map after the maximum pooling operation, MaxPool(·) represents the maximum pooling operation, and AvgPool(·) represents the average pooling operation.

[0082] The generation formula of the attention map is:

[0083] M(F)=σ(f(cat(F avg ,F max )));

[0084] In the above formula, M(F) represents the attention map; σ(·) represents the Sigmoid activation function; f(·) represents the convolution operation; cat(·) represents the concatenation operation of the feature map.

[0085] S13: Construct a fusion module to fuse the original feature map with the attention map to obtain the body attention feature map. The feature fusion method is to perform element-wise multiplication operation on the attention map and the original feature map.

[0086] The formula for the fusion operation is:

[0087]

[0088] In the above formula, F′ represents the body attention feature map, and M(F) represents the attention map; represents the multiplication operation according to the corresponding elements; F represents the original feature map.

[0089] S14: Construct a parameter inference module including a pooling layer and three fully connected layers. The parameter inference module is used to generate the SMPL parameters of the corresponding target person in the human image according to the input body attention feature map and the original feature map.

[0090] The pooling layer in the parameter inference module is an average pooling layer. The first two of the three fully connected layers have 1024 neurons each and are connected through the Dropout operation. The third fully connected layer has 85 neurons and is directly connected to the previous fully connected layer. The three fully connected layers constitute the iterative regression part of the parameter inference module.

[0091] Specifically, in this embodiment, if Figure 3 As shown in the figure, the generation process of the SMPL parameters of the target person is as follows:

[0092] (1) The input body attention feature map F′ is average pooled to obtain a feature φ.

[0093] (2) Putting together the SMPL pose parameter θ, shape parameter β and camera parameter c, the formula is expressed as:

[0094] Θ=cat(θ,β,c);

[0095] In the above formula, θ represents the pose parameter of the SMPL model; β represents the shape parameter of the SMPL model; c represents the camera parameter; Θ represents the concatenated parameter set of the pose parameter θ, the shape parameter β and the camera parameter c.

[0096] (3) Use the average pose parameters, average shape parameters and average camera parameters to form the initialization parameter set θ 0 , the feature φ and the parameter set Θ 0 The concatenation is used as the input to the iterative regression part in the parameter inference module.

[0097] (4) Generate the residual of the parameter set corresponding to the current input, and then update the current parameter set. The update formula is:

[0098] Θ t+1 =Θ t +ΔΘ t ;

[0099] In the above formula, Θ t Represents the parameter set corresponding to the current input, Θ t+1 Represents the parameter set Θ t The updated state, ΔΘ t Represents the parameter set Θ t The residual.

[0100] (5) Iterate the update operation of the previous step three times; during each iterative update process, the parameter set obtained in the previous update is concatenated with the feature φ as the input of the iterative regression part of this parameter inference module to update the parameter set.

[0101] (6) After the iterative operation is completed, the SMPL parameters including the final posture parameter θ, morphological parameter β and the corresponding camera parameter c are obtained.

[0102] S15: A SMPL submodule is connected after the parameter inference module. The SMPL (Skinned Multi-Person Linear Model) in this embodiment is a vertex-based three-dimensional naked human model, which can accurately represent different shapes and poses of the human body.

[0103] In the SMPL submodule, the SMPL parameters are input into the SMPL function to obtain a three-dimensional human body model; the expression of the SMPL function M(β,θ) is:

[0104]

[0105] In the above formula, B is the vertex coordinate of the 3D human body model in T posture; P (θ) and B S (β) represents the offset of the vertex vector relative to the SMPL standard template caused by the pose parameter θ and the morphological parameter β; J(β) is the joint point position of the model corresponding to the morphological parameter β; W(·) is the linear mixed skinning function; is the mixing weight.

[0106] S2: Acquire multiple human images containing target persons as original images, pre-process the original images to form a training data set, wherein the original images in the training data set at least include a portion of human images that are occluded by the person.

[0107] In this embodiment, the preprocessing process of the original image includes:

[0108] (1) Locate the target person in the human image and crop the image so that the target person is located in the central area of ​​the human image.

[0109] (2) The size of the cropped human body image is adjusted, and the pixel value of the adjusted image is unified to 224×224.

[0110] (3) Normalize the adjusted image to obtain the data element in the training data set.

[0111] S3: Using the training data set in the previous step, the human body reconstruction network model is trained by minimizing the network loss function.

[0112] During the training process of the network model, under the condition of minimizing the loss function, the Adam algorithm is used to adjust all the parameters in the network model and train the network.

[0113] The expression of the loss function is as follows:

[0114] L=λ 2D L 2Djoint +λ 3D L 3Djoint +λ para L SMPL ;

[0115] In the above formula, L 2Djoint represents the 2D joint point loss function; L 3Djoint represents the 3D joint point loss function; L SMPL represents the SMPL parameter loss function; λ 2D Represents the weight coefficient of the 2D joint loss function; λ 3D Represents the weight coefficient of the 3D joint loss function; λ para Represents the weight coefficient of the SMPL parameter loss function.

[0116] Among them, the 2D joint point loss function L 2Djoint The expression is:

[0117]

[0118] In the above formula, v i Indicates the visibility of the i-th 2D joint point, with a value of 0 or 1, 0 means invisible, and 1 means visible; N represents the number of 2D joint points; represents the predicted value of the i-th 2D joint point; k i represents the true value of the i-th 2D joint point; where the predicted value of the 2D joint point, It is obtained by orthogonal projection of the predicted 3D joint points.

[0119] Specifically in this embodiment, the projection formula is:

[0120]

[0121] In the above formula, represents the predicted value of 3D joint points, express The corresponding 2D joint point prediction value, (·) represents the projection function based on the camera parameter c.

[0122] 3D joint point loss function L 3Djoint The expression is:

[0123]

[0124] In the above formula, M represents the number of images involved in the 3D joint point calculation; represents the 3D joint point prediction value of the i-th image; J i Represents the true value of the 3D joint point of the i-th image.

[0125] SMPL parameter loss function L SMPL The expression is:

[0126]

[0127] In the above formula, O represents the number of images involved in the SMPL parameter calculation. and Respectively represent the predicted values ​​of posture parameters and morphological parameters of the i-th image, θ i and β i They represent the true values ​​of the pose parameters and morphological parameters of the i-th image respectively.

[0128] S4: Save the trained human body reconstruction network model; input the human body image to be processed into the saved network model after preprocessing to generate a three-dimensional human body model with a specific posture.

[0129] The processing process of human body reconstruction network model is as follows Figure 4 As shown, the pre-processed human body image is first subjected to convolution processing by a feature extraction model to extract the features related to the target person and generate an original feature map. Then the original feature map is transferred backward in two ways to the attention module and the parameter inference module respectively. The original feature map input into the attention module is first subjected to average pooling processing and maximum pooling processing to obtain an average pooling feature map and a maximum pooling feature map respectively. After the two types of pooling feature maps are feature spliced, they are sequentially subjected to convolution processing and Sigmoid operation to obtain an attention map. Then the fusion module simultaneously receives the attention map output by the attention module and the original feature map output by the feature extraction module; the attention map and the original feature map are fused to obtain a body attention feature map; the body attention feature map is input into the parameter inference module, and the parameter inference module generates a corresponding SMPL parameter according to the body attention feature map, and iteratively updates the SMPL parameter. Among them, the SMPL parameters include posture parameters and morphological parameters. Finally, the SMPL parameters that have completed the iterative update are input into the SMPL submodule to generate a three-dimensional human body model of the target person.

[0130] The method provided by the present invention can process the features in the human body image by introducing an attention mechanism, so that the network can focus on the features containing important information and ignore and reduce the attention to unimportant information. Among them, the original feature map is weighted by the attention map generated by the attention mechanism, so that the network focuses on the information related to the human body part in the image and reduces the attention to other information, thereby reducing the interference of the occlusion information on the network. At the same time, the network model can use the features of the visible part of the body to infer the situation of the occluded body part. It is ensured that the human posture and shape reflected in the finally constructed three-dimensional human body model are more in line with reality.

[0131] Example 2

[0132] This embodiment provides a human body reconstruction model, which uses the human body reconstruction method based on the attention mechanism as in Embodiment 1 to process the input human body image with occlusion, and then generates a three-dimensional human body model of the target task in the human body image. Figure 5 As shown, the human body reconstruction model includes the following: a preprocessing module, a feature extraction module, an attention module, a fusion module, a parameter inference module, and a SMPL sub-model.

[0133] The preprocessing module is used to: (1) locate the target person in the human image and crop the image so that the target person is located in the central area of ​​the human image; (2) adjust the size of the cropped human image, and unify the pixel values ​​of the adjusted image to 224×224; (3) normalize the adjusted image.

[0134] The feature extraction module uses the convolution part of the deep convolutional neural network Resnet50 as the backbone network. The output of the preprocessing module is used as the input of the feature extraction module; the feature extraction module is used to extract the features of the preprocessed human body image through convolution operations, and then generate the corresponding original feature map.

[0135] The attention module includes a maximum pooling submodule, an average pooling submodule, a feature splicing submodule, a convolution submodule and a Sigmoid operation submodule. The output of the feature extraction module is used as the input of the attention module; the original feature map is processed by the maximum pooling submodule and the average pooling submodule in the attention module to obtain two feature maps, and the two feature maps are feature spliced ​​in the feature splicing submodule; after the convolution processing in the convolution submodule and the Sigmoid operation in the Sigmoid operation submodule, the attention map is obtained. The detailed generation process of the attention map has been described in detail in Example 1, and will not be repeated here.

[0136] The fusion module uses the original feature map output by the feature extraction module and the attention map output by the attention module, and then multiplies the original feature map and the attention map by corresponding elements to obtain the fused body attention feature map. The feature fusion method is to multiply the attention map and the original feature map by corresponding elements. The formula for the fusion operation is:

[0137]

[0138] In the above formula, F′ represents the body attention feature map, and M(F) represents the attention map; represents the multiplication operation according to the corresponding elements; F represents the original feature map.

[0139] The parameter inference module includes an average pooling layer, a fully connected layer 1, a fully connected layer 2, and a fully connected layer 3. Among them, the fully connected layer 1 and the fully connected layer 2 have 1024 neurons and are connected through the Dropout operation; the fully connected layer 3 has 85 neurons, and the fully connected layer 2 is directly connected to the fully connected layer 3. The fully connected layer 1, the fully connected layer 2, and the fully connected layer 3 constitute the iterative regression part of the network model. The output of the fusion module is used as the input of the parameter inference module; the parameter inference module generates iteratively updated SMPL parameters based on different input data.

[0140] The SMPL submodule is used to generate a three-dimensional human body model of the target person corresponding to the human body image according to the SMPL parameters output by the parameter inference submodule.

[0141] In other embodiments, the preprocessing module and the SMPL sub-model may be part of the human body reconstruction model or not. When the preprocessing module is not part of the human body reconstruction model, each human body image may be manually processed before being input into the human body reconstruction model, so that the input human body image is more in line with the requirements. At the same time, the manual processing may also make the target person in the center of the image, and make the proportion of the occluder in the image relatively reduced. In this way, a more accurate three-dimensional human body model result may be obtained.

[0142] When the SMPL sub-model is not part of the human body reconstruction model, the existing SMPL model can be called through the relevant module call program, the corresponding SMPL parameters generated by the human body reconstruction model are input into the SMPL model, and the three-dimensional human body model generated by the SMPL model is obtained at the same time. When this method is used, the structure and scale of the human body reconstruction model are simplified, which can save computing power and reduce the requirements for hardware equipment. At the same time, distributed computing can also be used in this architecture to improve the generation rate of the three-dimensional human body model.

[0143] In order to verify the performance of a human body reconstruction model provided in this embodiment, this embodiment also simulates the processing of the model. The simulation experiment environment uses Intel(R)Xeon(R)CPU E5-2609 V4@1.70GHz, 16G memory, Ubuntu18.04 system, GTX1080Ti graphics card, Pycharm programming environment, pytorch1.1.0 deep learning framework, and the data sets used are 2D data sets Leeds Sports Pose (LSP) data sets, MPII data sets and 3D data sets 3DPW data sets, Human3.6M data sets.

[0144] Through simulation, it is found that the human body reconstruction model provided in this embodiment still has good three-dimensional model reconstruction performance for various human body images with occlusions. The human body posture and shape of the constructed three-dimensional human body model are very realistic, and therefore has good practical value. It is suitable for application in various scenarios that rely on three-dimensional modeling of the human body, such as virtual fitting, body animation, and human motion simulation games.

[0145] Example 3

[0146] This embodiment provides a human body reconstruction device based on an attention mechanism. The reconstruction device is a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the human body reconstruction method based on the attention mechanism as described in Example 1 are implemented.

[0147] The computer device may be a smart phone, tablet computer, laptop computer, desktop computer, rack server, blade server, tower server or cabinet server (including an independent server or a server cluster composed of multiple servers) that can execute programs. The computer device of this embodiment includes at least but is not limited to: a memory and a processor that can communicate with each other through a system bus.

[0148] In this embodiment, the memory (i.e., readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of a computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Of course, the memory may also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the memory is generally used to store an operating system and various application software installed on the computer device. In addition, the memory may also be used to temporarily store various types of data that have been output or are to be output.

[0149] The processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor is generally used to control the overall operation of a computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data to implement the processing of the human body reconstruction method based on the attention mechanism in the aforementioned embodiment, so as to construct a three-dimensional human body model corresponding to the target task according to a single human body image of the target person given.

[0150] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A human body reconstruction method based on an attention mechanism, characterized in that, it includes the following steps: Step 1: Construct a human body reconstruction network model, which includes a feature extraction module, an attention module, a fusion module, a parameter inference module, and an SMPL sub-module; the feature extraction module is used to generate a corresponding original feature map according to the input human body image; the attention module includes two pooling layers, a convolutional layer, and a Sigmoid operation layer; the two pooling layers are respectively an average pooling layer and a max pooling layer; the attention module is used to generate an attention map according to the input original feature map; The fusion module is used to perform a fusion operation on the original feature map and the attention map to obtain a body attention feature map; the parameter inference module includes a pooling layer and three fully connected layers; the parameter inference module is used to generate SMPL parameters of the corresponding target person in the human body image according to the input body attention feature map; the SMPL sub-module is used to generate a three-dimensional human body model of the corresponding target person according to the SMPL parameters; In the parameter inference module, the generation process of SMPL parameters is as follows: (1) The input body attention feature map F′ is averaged pooled to obtain a feature φ; (2) The SMPL pose parameter θ, shape parameter β, and camera parameter c are concatenated together, and the formula is expressed as: Θ = cat(θ, β, c); In the above formula, θ represents the pose parameter of the SMPL model; β represents the shape parameter of the SMPL model; c represents the camera parameter; Θ represents the parameter set concatenated by the pose parameter θ, shape parameter β, and camera parameter c; (3) Use the average pose parameters, average shape parameters and average camera parameters to form the initialization parameter set θ 0 , the feature φ and the parameter set Θ 0 The splicing is performed as an input to the iterative regression part in the parameter inference module; (4) Generate the residual of the parameter set corresponding to the current input, and then update the current parameter set. The update formula is: I t+1 =Θ t +DTH t ; In the above formula, Θ t Represents the parameter set corresponding to the current input, Θ t+1 Represents the parameter set Θ t The updated state, ΔΘ t Represents the parameter set Θ t The residual of (5) Iterate the update operation in the above step 3 times; In each iteration update process, the parameter set obtained from the previous update is concatenated with the feature φ as the input of the iterative regression part of the parameter inference module in this iteration to update the parameter set; (6) After the iteration operation is completed, the SMPL parameters including the final pose parameter θ, shape parameter β, and the corresponding camera parameter c are obtained; Step 2: Obtain multiple human body images containing the target person as original images, preprocess the original images to form a training data set, and the original images in the training data set at least include some human body images with people occluded; Step 3: Use the training data set in the above step to train the human body reconstruction network model by minimizing the network loss function; Step 4: Save the trained human body reconstruction network model; input the preprocessed human body image to be processed into the saved network model to generate a human body three-dimensional model with a specific pose.

2. The human body reconstruction method based on an attention mechanism according to claim 1, characterized in that: The feature extraction module is obtained by streamlining and repackaging the deep convolutional neural network Resnet50, and only the convolutional part of the original network model is retained during the streamlining process; the input human body image is convolutionally processed by the feature extraction module to obtain the original feature map.

3. The human body reconstruction method based on the attention mechanism as claimed in claim 1, Features: The attention module takes the output of the feature extraction module as input. The input original feature map first passes through the average pooling layer and the maximum pooling layer in the attention module respectively. The two pooling results are concatenated and then sequentially subjected to convolution processing and Sigmoid operation to obtain the attention map. In the attention module, the pooling operation formula of the average pooling layer is: F avg =AvgPool(F); The pooling operation formula of the maximum pooling layer is: F max =MaxPool(F); In the above formula, F represents the original feature map, F avg represents the feature map after the average pooling operation, F max Represents the feature map after the maximum pooling operation, MaxPool(·) represents the maximum pooling operation, and AvgPool(·) represents the average pooling operation; The generation operation formula of the attention map is: M(F)=σ(f(cat(F avg ,F max ))); In the above formula, M(F) represents the attention map; σ(·) represents the Sigmoid activation function; f(·) represents the convolution operation; cat(·) represents the concatenation operation of the feature map.

4. The human body reconstruction method based on the attention mechanism as claimed in claim 1, Features: In the fusion module, the fused body attention feature map is obtained by performing corresponding element multiplication operation on the attention map and the original feature map; wherein the formula of the fusion operation is: In the above formula, F′ represents the body attention feature map, and M(F) represents the attention map; represents the multiplication operation of corresponding elements; F represents the original feature map.

5. The human body reconstruction method based on the attention mechanism as claimed in claim 1, Features: The pooling layer in the parameter inference module is an average pooling layer; the first two of the three fully connected layers each have 1024 neurons and are operated through Dropout; the third fully connected layer has 85 neurons and is directly connected to the previous fully connected layer; wherein the three fully connected layers constitute the iterative regression part in the parameter inference module.

6. The human body reconstruction method based on the attention mechanism as claimed in claim 1, Features: In the SMPL submodule, the SMPL parameters are input into the SMPL function, and the SMPL function maps the morphological parameters and posture parameters into the vertices of the model to obtain the three-dimensional human body model; the expression of the SMPL function M(β,θ) is: In the above formula, B is the vertex coordinate of the 3D human body model in T posture; P (θ) and B S (β) represents the offset of the vertex vector relative to the SMPL standard template caused by the pose parameter θ and the morphological parameter β; J(β) is the joint point position of the model corresponding to the morphological parameter β; W(·) is the linear mixed skinning function; is the mixing weight.

7. The human body reconstruction method based on the attention mechanism as claimed in claim 1, Features: The preprocessing process of the original image includes: Locate the target person in the human image and crop the image so that the target person is located in the central area of ​​the human image; The size of the cropped human body image is adjusted, and the pixel value of the adjusted image is unified to 224×224; The adjusted image is normalized to obtain the data element in the training data set.

8. The human body reconstruction method based on the attention mechanism as claimed in claim 1, Features: During the training of the network model, the Adam algorithm is used to adjust all parameters in the network model and train the network under the condition of minimizing the loss function. The expression of the loss function is as follows: L=λ 2D L 2Djoint +λ 3D L 3Djoint +λ para L SMPL ; In the above formula, L 2Djoint represents the 2D joint point loss function; L 3Djoint represents the 3D joint point loss function; L SMPL represents the SMPL parameter loss function; λ 2D Represents the weight coefficient of the 2D joint loss function; λ 3D Represents the weight coefficient of the 3D joint loss function; λ para Represents the weight coefficient of the SMPL parameter loss function; Among them, the 2D joint point loss function L 2Djoint has the following expression: In the above formula, v i Indicates the visibility of the i-th 2D joint point, with a value of 0 or 1, 0 means invisible, and 1 means visible; N represents the number of 2D joint points; represents the predicted value of the i-th 2D joint point; k i represents the true value of the i-th 2D joint point; where the predicted value of the 2D joint It is obtained by projecting the predicted 3D joint points; 3D joint point loss function L 3Djoint The expression is as follows: In the above formula, M represents the number of images involved in the 3D joint point calculation; represents the 3D joint point prediction value of the i-th image; J i Represents the true value of the 3D joint point of the i-th image; SMPL parameter loss function L SMPL The expression is: In the above formula, O represents the number of images involved in the SMPL parameter calculation. and Respectively represent the predicted values ​​of posture parameters and morphological parameters of the i-th image, θ i and β i They represent the true values ​​of the pose parameters and morphological parameters of the i-th image respectively.

9. A human body reconstruction device based on an attention mechanism, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the human body reconstruction method based on the attention mechanism as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN111640119A

  • Picture-based SMPL parameter prediction and human body model generation method

    CN111968217A