Single-view three-dimensional dressed human body reconstruction method based on pixel and voxel feature fusion

By using a method of integrating pixels and voxel features in single-view three-dimensional dress human body reconstruction, combined with multi-layer perceptron and Marching Cubes module, the problem of low reconstruction accuracy in the existing technology is solved, and a higher precision three-dimensional human body model reconstruction is achieved.

CN120047617APending Publication Date: 2025-05-27SHAANXI NORMAL UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510112785.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately extract joint nodes and process clothing details in single-view three-dimensional dress reconstruction, resulting in low reconstruction accuracy.

Method used

Using a method based on the fusion of pixel and voxel features, a three-dimensional dress mannequin model is reconstructed through a pixel feature extraction network and a voxel feature extraction network combined with a multi-layer perceptron and Marching Cubes module. This method uses coordinate attention modules and stacked hourglass network to extract high-precision pixel features, and extracts precise voxel features through three-dimensional channels and spatial attention modules.

Benefits of technology

The accuracy of human body reconstruction in three-dimensional clothing has been significantly improved, especially in terms of limb integrity and clothing details, which have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047617A_ABST
    Figure CN120047617A_ABST
Patent Text Reader

Abstract

The invention discloses a single-view three-dimensional dressed human body reconstruction method based on pixel and voxel feature fusion. The method comprises the steps of data set preprocessing, data set dividing, three-dimensional dressed human body reconstruction network construction, three-dimensional dressed human body reconstruction network training, three-dimensional dressed human body reconstruction network testing and reconstruction result displaying. In the pixel feature extraction network, the coordinate attention and the stacking hourglass network are combined, so that the pixel feature extraction precision is improved; in the voxel feature extraction network, a three-dimensional channel and space attention are introduced into the voxelization human body parameterization model network, and a multi-scale residual feature extraction module is combined to extract voxel features. And splicing the extracted pixels and voxel feature vectors, inputting the spliced pixels and voxel feature vectors into a multi-layer perceptron, and reconstructing a three-dimensional dressed human body model through Marking Cubes, thereby solving the technical problem of insufficient precision of the existing single-view three-dimensional dressed human body reconstruction method, and being applicable to the technical fields of artificial intelligence and computer vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and computer vision, and specifically relates to single-view Figure 3 dressed human body reconstruction. Background Art

[0002] Single-view Figure 3 dressed human body reconstruction is a challenging task in the field of computer vision. Its goal is to reconstruct a high-fidelity three-dimensional human model from a single image. This technology is of great significance in many application fields such as 3D printing, virtual reality, and game design. Traditional three-dimensional reconstruction methods, such as those based on geometric models or template matching, although able to complete the reconstruction task to a certain extent, often have difficulty accurately capturing and understanding the complex semantic information in the image, and show strong dependence when dealing with human poses, lighting changes, and clothing details. This is because human images usually contain rich details and texture information, and parsing these information requires more advanced computer vision and deep learning technologies to achieve more accurate and high-quality three-dimensional reconstruction.

[0003] Existing deep learning methods, such as convolutional neural networks, have made significant progress in single-view Figure 3 dressed human body reconstruction methods. Neural networks can effectively extract hierarchical features from images and capture local details and global structure information of the human body. Traditional neural network architectures are usually limited to the processing of two-dimensional image information and lack a profound understanding of three-dimensional spatial relationships. To solve this problem, researchers have introduced more complex network architectures, such as convolutional neural networks, generative adversarial networks, and graph convolutional networks. In addition, human parametric models such as the Skinned Multi-Person Linear Model are also widely used. By integrating prior knowledge of the human body into the network, the pose, shape, and clothing fit of the reconstruction results can be effectively constrained, improving the reconstruction accuracy and effect. The human parametric model provides a flexible and efficient way to perform human body reconstruction by establishing a parametric representation for the three-dimensional human body, playing an important role in the deep learning framework and significantly enhancing the detail expression and diversity processing capabilities during the reconstruction process.

[0004] Existing deep learning methods still face some technical problems to be solved, such as achieving a more accurate single-view Figure 3 dressed human body reconstruction model, more delicate clothing details, etc. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the above-mentioned shortcomings of the existing technology and provide a single-view based on the fusion of pixel and voxel features with accurate joint point extraction and high reconstruction accuracy for processing clothing details. Figure 33D Dressed Human Reconstruction Method

[0006] The technical solution steps to solve the above technical problems are as follows:

[0007] (1) Dataset Preprocessing

[0008] Utilize the open-source dataset THUman2.0 from Tsinghua University. The THUman2.0 dataset contains 526 3D scanned human models, and the data of each 3D scanned human model consists of a 3D file in obj format, a texture image, and a material file. The following process is used for preprocessing this dataset.

[0009] Normalize and randomly scale each 3D scanned human model; pre-compute Photometric Rendering Transfer (PRT) to improve photo-realism; rotate the 3D model 360° around the y-axis using the OpenGL (Open Graphics Library), and collect the original image, segmentation mask image, and texture image every 1°. Generate the joint point information corresponding to each original image with the help of the pose estimation library OpenPose, and generate the human parametric model corresponding to the original image using this joint point information.

[0010] (2) Dataset Partitioning

[0011] Partition the preprocessed THUman2.0 dataset into a training set and a test set according to 8:2.

[0012] (3) Constructing a 3D Dressed Human Reconstruction Network

[0013] The 3D dressed human reconstruction network is composed of a pixel feature extraction network, a voxel feature extraction network, a multi-layer perceptron, and a Marching Cubes module connected in sequence. The output ends of the pixel feature extraction network and the voxel feature extraction network are connected in series with the multi-layer perceptron and the Marching Cubes module in turn.

[0014] The pixel feature extraction network is composed of Coordinate Attention Module 1, Convolution Layer 1, BN Layer 1, ReLU Activation Function Layer 1, Coordinate Attention Module 2, Residual Module 1, Max Pooling Layer 1, Coordinate Attention Module 3, Residual Module 2, Coordinate Attention Module 4, Residual Module 3, Stacked Hourglass Network Module 1, Stacked Hourglass Network Module 2, Stacked Hourglass Network Module 3, Stacked Hourglass Network Module 4, Fourth-order Hourglass Module 1, Residual Module 4, Convolution Layer 2, BN Layer 2, ReLU Activation Function Layer 2, and Convolution Layer 3 connected in series in sequence.

[0015] The voxel feature extraction network is composed of a voxelized human parametric model, a 3D channel and spatial attention module, and a multi-scale residual feature extraction module connected in series in sequence.

[0016] (4) Training the 3D Dressed Human Reconstruction Network

[0017] 1) Construct the depth blur reconstruction loss function

[0018] Construct the depth blur reconstruction loss function according to Equation (1)

[0019]

[0020] C(p i +Δp i ) = (S(F I , π(p i +Δp i ), S(F V , (p i +Δp i ))) T

[0021] Δp i = (0, 0, Δz i ) T

[0022]

[0023] F I = E I (I)

[0024] F V = E V (V O )

[0025]

[0026] where n p is the number of point samples, n p is a finite positive integer, p i is the three-dimensional point sample indexed by i, Δp i is the compensation translation along the z-axis, F * (p i ) is the true occupancy value of p i , C(p i +Δp i ) is the parametric model conditional implicit, F I represents the image feature map from the depth image encoder E I (·), I represents the input image, π(p i +Δp i ) is the two-dimensional projection of the point p i +Δp i after z-axis compensation on the image feature map F I , S(F I , π(pi +Δp i ) represents sampling F at pixel π(p i +Δp i ), where F I is the feature volume, and S(F V ,(p V +Δp i )) represents the voxel-aligned volume feature sampled from F i . E V is the 3D encoder, V V is the occupancy volume after converting the human parametric model, and F(C(p O +Δp i )) represents classifying whether the point p i +Δp i is inside or outside the surface. Δz i is determined according to the apparent correspondence between the predicted human parametric model and the real model during network training. ω i is the corresponding mixing weight, ||p j→i -v i || is the three-dimensional distance between the point p j and its neighboring vertex v i of the human parametric model in 3D space. ω j is the weight normalization factor, and v i and are the j-th vertices of the predicted human parametric model and the j-th vertex of the corresponding real human parametric model respectively. σ is the standard deviation. Z(v j j ) and are the depth values of v j and in the camera coordinate space respectively.

[0027] 2) Construct evaluation metrics

[0028] The evaluation metrics consist of the chamfer distance d CD and the point-to-plane distance d PSD .

[0029] The chamfer distance d CD is constructed according to Equation (2):

[0030]

[0031] where m = |P|, P is the set of target points, P ∈ {p 1 , p 2 ,..., p m}, n = |Q|, Q is the set of reconstructed points, Q ∈ {q 1 , q2 ,..., q n},where m and n are finite positive integers, ||p i - q j || is the three - dimensional distance between point p i and point q j in 3D space, ||q j - p i || is the three - dimensional distance between point q j and point p i in 3D space.

[0032] Construct the distance d from a point to a plane according to the following formula PSD :

[0033]

[0034] (3) Train the three - dimensional dressed human body reconstruction network

[0035] Input the training set into the three - dimensional dressed human body reconstruction network for training.

[0036] The parameter settings for the training process are as follows: the server graphics card is RTX 3090, the adaptive gradient method optimizer is Adam, the initial learning rate is 0.001, the learning rate of the network is dynamically adjusted during the training process, the batch size is 3, the number of epochs is 10, the number of sampling points for each human body model is 5000, and in every 10,000 iterations, the learning rate decays by 0.1 times until the loss function converges.

[0037] (5) Test the three - dimensional dressed human body reconstruction network

[0038] Input the test set into the trained three - dimensional dressed human body reconstruction network for testing, and evaluate according to the evaluation index to obtain the optimal reconstruction result.

[0039] (6) Display the reconstruction result

[0040] Input a single human picture to obtain the reconstruction result of the human body model.

[0041] In the construction of the human body reconstruction network in step (3) of the present invention, the structures and connection relationships of the stacked hourglass network module 1 and the stacked hourglass network module 2 are as follows:

[0042] The described stacked hourglass network module 1 is composed of a fourth-order hourglass network module 2, a residual module 5, a convolutional layer 4, a BN layer 3, a ReLU activation function layer 3, a convolutional layer 5, a Concat layer 1, a coordinate attention module 5, a convolutional layer 6, and a convolutional layer 7. The fourth-order hourglass network module 2 is serially connected to the residual module 5, the convolutional layer 4, the BN layer 3, the ReLU activation function layer 3, the convolutional layer 5, the Concat layer 1, and the coordinate attention module 5 in sequence. The other output end of the ReLU activation function layer 3 is connected to the second input end of the Concat layer 1 through the convolutional layer 6 and the convolutional layer 7. The input end of the fourth-order hourglass network module 2 is connected to the third input end of the Concat layer 1. The other output end of the convolutional layer 7 is connected to the stacked hourglass network module 2, and the output end of the coordinate attention module 5 is connected to the stacked hourglass network module 2.

[0043] The described stacked hourglass network module 2 is composed of a fourth-order hourglass network module 3, a residual module 6, a convolutional layer 8, a BN layer 4, a ReLU activation function layer 4, a convolutional layer 9, a Concat layer 2, a coordinate attention module 6, a convolutional layer 10, and a convolutional layer 11. The fourth-order hourglass network module 3 is serially connected to the residual module 6, the convolutional layer 8, the BN layer 4, the ReLU activation function layer 4, the convolutional layer 9, the Concat layer 2, and the coordinate attention module 6 in sequence. The other output end of the ReLU activation function layer 4 is connected to the second input end of the Concat layer 2 through the convolutional layer 10 and the convolutional layer 11. The input end of the fourth-order hourglass network module 3 is connected to the third input end of the Concat layer 2. The other output end of the convolutional layer 11 is connected to the stacked hourglass network module 3, and the output end of the coordinate attention module 6 is connected to the stacked hourglass network module 3.

[0044] The structures of the stacked hourglass network module 3 and the stacked hourglass network module 4 are the same as the structure of the stacked hourglass network module 1.

[0045] The connection relationship between the described stacked hourglass network module 2 and the stacked hourglass network module 3 is the same as the connection relationship between the stacked hourglass network module 1 and the stacked hourglass network module 2. The connection relationship between the stacked hourglass network module 3 and the stacked hourglass network module 4 is the same as the connection relationship between the stacked hourglass network module 1 and the stacked hourglass network module 2.

[0046] In the voxel feature extraction network for constructing the human body reconstruction network in step (3) of the present invention, the multi-scale residual feature extraction module is composed of convolutional layer 12, convolutional layer 13, convolutional layer 14, convolutional layer 15, convolutional layer 16, convolutional layer 17, convolutional layer 18, convolutional layer 19, convolutional layer 20, convolutional layer 21, and Concat layer 3; convolutional layer 13 and convolutional layer 14 are connected in series, convolutional layer 15, convolutional layer 16, and convolutional layer 17 are connected in series in sequence, convolutional layer 18, convolutional layer 19, convolutional layer 20, convolutional layer 21, and Concat layer 3 are connected in series in sequence, the input ends of convolutional layer 12, convolutional layer 13, convolutional layer 15, and convolutional layer 18 are connected to the second input end of Concat layer 3, and the output ends of convolutional layer 12, convolutional layer 14, and convolutional layer 17 are connected to the third input end of Concat layer 3.

[0047] In the present invention, the convolutional kernel size of convolutional layer 12 is 1×1, the stride is 1, the sizes of convolutional layer 13, convolutional layer 15, and convolutional layer 18 are the same as that of convolutional layer 12, the convolutional kernel size of convolutional layer 14 is 3×3, the stride is 1, the padding is 2, the dilation is 2, the sizes of convolutional layer 16 and convolutional layer 19 are the same as that of convolutional layer 14, the convolutional kernel size of convolutional layer 17 is 3×3, the stride is 1, the padding is 1, and the sizes of convolutional layer 20 and convolutional layer 21 are the same as that of convolutional layer 17.

[0048] In equation (1) for training the three-dimensional dressed human body reconstruction network in step (4) of the present invention, the n p is the number of point samples, and n p ∈ [5000, 10000].

[0049] In equation (2), m = |P|, where P is the target point set, P ∈ {p 1 , p 2 ,..., p m}, n = |Q|, where Q is the point set obtained by reconstruction, Q ∈ {q 1 , q 2 ,..., q n}, m, n ∈ [5000, 10000], ||p i - q j || is the three-dimensional distance between point p i and point q j in 3D space, and ||q j - p i || is the three-dimensional distance between point q j and point p i in 3D space.

[0050] In equation (1) for training the three-dimensional dressed human body reconstruction network in step (4) of the present invention, the n p is the number of point samples, and n pThe optimal value for the extraction is 7000.

[0051] In formula (2), m = |P|, where P is the target point set, P ∈ {p 1 , p 2 ,..., p m}}, n = |Q|, where Q is the point set obtained by reconstruction, Q ∈ {q 1 , q 2 ,..., q n}}, the optimal values for m and n are 7000, ||p i - q j || is the three-dimensional distance between point p i and point q j in 3D space, ||q j - p i || is the three-dimensional distance between point q j and point p i in 3D space.

[0052] Since the present invention adopts a pixel feature extraction network, combines coordinate attention and a stacked hourglass network, the pixel feature extraction accuracy is improved; in the voxel feature extraction network, a three-dimensional channel and spatial attention module, as well as a multi-scale residual feature extraction module, are adopted to extract more accurate voxel feature information. The extracted pixel and voxel features are spliced and input into a multi-layer perceptron, and a three-dimensional dressed human body model is reconstructed through Marching Cubes. A comparison experiment with Fangzheng was carried out using the method of Embodiment 1 of the present invention. The experimental results show that the integrity of the limbs and the wrinkle details of the clothing of the reconstructed three-dimensional human body model have been significantly improved. Compared with the prior art, the present invention can extract joint points more accurately, process clothing details, and improve the accuracy of three-dimensional dressed human body reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is the flowchart of Embodiment 1 of the present invention.

[0054] Figure 2 is the structural schematic diagram of the three-dimensional dressed human body reconstruction network.

[0055] Figure 3 is Figure 2 the structural schematic diagram of the pixel feature extraction network in

[0056] Figure 4 is Figure 3 the structural and connection schematic diagram of Stacked Hourglass Network Module 1 and Stacked Hourglass Network Module 2.

[0057] Figure 5 is Figure 2 the structural schematic diagram of the voxel feature extraction network in

[0058] Figure 6 is Figure 5 The structural schematic diagram of the multi-scale residual feature extraction module in

[0059] Figure 7 is the result graph of the comparative experiment. Specific implementation manners

[0060] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments, but the present invention is not limited to the following embodiments.

[0061] Embodiment 1

[0062] In Figure 1 the single-view Figure 3 dimensional dressed human body reconstruction method based on pixel and voxel feature fusion in this embodiment consists of the following steps:

[0063] (1) Dataset preprocessing

[0064] Using the open-source dataset THUman2.0 from Tsinghua University, the THUman2.0 dataset contains 526 3D scanned human models, and the data of each 3D scanned human model consists of a 3D file in obj format, a texture image, and a material file. This dataset is preprocessed through the following process:

[0065] Normalize and randomly scale each 3D scanned human model; pre-calculate the precomputed radiance transfer (PRT) to improve the photorealism; use the Open Graphics Library (OpenGL) to rotate the 3D model 360° around the y-axis, and collect the original image, segmentation mask image, and texture image every 1°; generate the joint point information corresponding to each original image with the help of the pose estimation library OpenPose, and generate the human parametric model corresponding to the original image with this joint point information.

[0066] (2) Dataset division

[0067] Divide the preprocessed THUman2.0 dataset into a training set and a test set according to 8:2.

[0068] (3) Construct a 3D dressed human body reconstruction network

[0069] Figure 2 shows the structural schematic diagram of the 3D dressed human body reconstruction network in this embodiment. In Figure 2 the 3D dressed human body reconstruction network in this embodiment is composed of a pixel feature extraction network, a voxel feature extraction network, a multi-layer perceptron, and a MarchingCubes module connected in series. The output ends of the pixel feature extraction network and the voxel feature extraction network are connected in series with the multi-layer perceptron and the Marching Cubes module in sequence.

[0070] Figure 3shows Figure 2 the schematic structural diagram of the pixel feature extraction network in Figure 3 In

[0071] Figure 4 shows Figure 3 the structural diagram and connection relationship of stacked hourglass network module 1 and stacked hourglass network module 2 in Figure 4 In

[0072] The stacked hourglass network module 1 of this embodiment is composed of a fourth-order hourglass network module 2, a residual module 5, a convolutional layer 4, a BN layer 3, a ReLU activation function layer 3, a convolutional layer 5, a Concat layer 1, a coordinate attention module 5, a convolutional layer 6, and a convolutional layer 7 connected in series; the fourth-order hourglass network module 2 is connected in series with the residual module 5, the convolutional layer 4, the BN layer 3, the ReLU activation function layer 3, the convolutional layer 5, the Concat layer 1, and the coordinate attention module 5 in sequence. The other output end of the ReLU activation function layer 3 is connected to the second input end of the Concat layer 1 through the convolutional layer 6 and the convolutional layer 7. The input end of the fourth-order hourglass network module 2 is connected to the third input end of the Concat layer 1. The other output end of the convolutional layer 7 is connected to the stacked hourglass network module 2, and the output end of the coordinate attention module 5 is connected to the stacked hourglass network module 2.

[0073] The stacked hourglass network module 2 of this embodiment is composed of a fourth-order hourglass network module 3, a residual module 6, a convolutional layer 8, a BN layer 4, a ReLU activation function layer 4, a convolutional layer 9, a Concat layer 2, a coordinate attention module 6, a convolutional layer 10, and a convolutional layer 11 connected in series; the fourth-order hourglass network module 3 is connected in series with the residual module 6, the convolutional layer 8, the BN layer 4, the ReLU activation function layer 4, the convolutional layer 9, the Concat layer 2, and the coordinate attention module 6 in sequence. The other output end of the ReLU activation function layer 4 is connected to the second input end of the Concat layer 2 through the convolutional layer 10 and the convolutional layer 11. The input end of the fourth-order hourglass network module 3 is connected to the third input end of the Concat layer 2. The other output end of the convolutional layer 11 is connected to the stacked hourglass network module 3, and the output end of the coordinate attention module 6 is connected to the stacked hourglass network module 3.

[0074] The structures of stacked hourglass network module 3 and stacked hourglass network module 4 are the same as that of stacked hourglass network module 1.

[0075] The connection relationship between the stacked hourglass network module 2 and the stacked hourglass network module 3 is the same as that between the stacked hourglass network module 1 and the stacked hourglass network module 2, and the connection relationship between the stacked hourglass network module 3 and the stacked hourglass network module 4 is the same as that between the stacked hourglass network module 1 and the stacked hourglass network module 2.

[0076] Figure 5 The Figure 2 structural schematic diagram of the voxel feature extraction network in Figure 5 In

[0077] Figure 6 The Figure 5 structural schematic diagram of the multi-scale residual feature extraction module in Figure 6 In

[0078] The multi-scale residual feature extraction module of this embodiment is composed of convolutional layer 12, convolutional layer 13, convolutional layer 14, convolutional layer 15, convolutional layer 16, convolutional layer 17, convolutional layer 18, convolutional layer 19, convolutional layer 20, convolutional layer 21, and Concat layer 3 connected in series; convolutional layer 13 is connected in series with convolutional layer 14, convolutional layer 15 is connected in series with convolutional layer 16 and convolutional layer 17 in sequence, convolutional layer 18 is connected in series with convolutional layer 19, convolutional layer 20, convolutional layer 21, and Concat layer 3 in sequence, the input ends of convolutional layer 12, convolutional layer 13, convolutional layer 15, and convolutional layer 18 are connected to the second input end of Concat layer 3, and the output ends of convolutional layer 12, convolutional layer 14, and convolutional layer 17 are connected to the third input end of Concat layer 3.

[0079] (4) Training the 3D dressed human body reconstruction network

[0080] 1) Constructing the depth blur reconstruction loss function

[0081] Construct the depth blur reconstruction loss function according to Equation (1)

[0082]

[0083] C(p i +Δp i )=(S(F I ,π(p i +Δp i )),S(F V ,(p i +Δp i ))) T

[0084] Δp i =(0,0,Δz i ) T

[0085]

[0086] F I =E I (I)

[0087] F V =E V (V O )

[0088]

[0089]

[0090] Among them, n p is the number of point samples, n p ∈[5000,10000], and in this embodiment, n p takes the value of 7000, p i is a three-dimensional point sample indexed by i, Δp i is the compensation translation along the z-axis, F * (p i ) is the true value occupancy of p i , C(p i +Δp i ) is the parametric model conditional implicit, F I represents the image feature map from the depth image encoder E I (·), I represents the input image, π(p i +Δp i ) is the two-dimensional projection of the point p i +Δp i on the image feature map F I , and S(F I ,π(p i +Δp i )) represents bilinear interpolation at the pixel π(pi +Δp i ) sample F at I The value of F V is the feature volume, S(F V ,(p i +Δp i )) represents the voxel-aligned volume feature sampled from F V E V is the 3D encoder, V O is the occupancy volume after converting the human parametric model, F(C(p i +Δp i )) represents classifying the point p after z-axis compensation i +Δp i whether it is inside or outside the surface, Δz i is based on the apparent correspondence between the predicted human parametric model and the real model during network training, ω j→i is the corresponding mixing weight, ||p i -v j || is the three-dimensional distance in 3D space between the point p i and its neighboring human parametric model vertex v j ω i is the weight normalization factor, v j and are the j-th vertex of the predicted human parametric model and the j-th vertex of the corresponding real human parametric model respectively, σ is the standard deviation, Z(v j ) and are the depth values of v j and in the camera coordinate space respectively.

[0091] 2) Construct evaluation metrics

[0092] The evaluation metrics consist of the chamfer distance d CD and the point-to-plane distance d PSD .

[0093] Construct the chamfer distance d according to Equation (2) CD :

[0094]

[0095] where m = |P|, P is the set of target points, P ∈ {p 1 ,p 2 ,...,p m}, n = |Q|, Q is the set of points obtained by reconstruction, Q ∈ {q 1 ,q 2 ,...,q n}, m, n ∈ [5000, 10000], where m and n in this embodiment are taken as 7000, ||p i -q j || is the three-dimensional distance between point p i and point q j in 3D space, ||q j -p i || is the three-dimensional distance between point q j and point p i in 3D space.

[0096] Construct the distance d from a point to a plane according to the following formula PSD :

[0097]

[0098] (3) Train the 3D dressed human body reconstruction network

[0099] Input the training set into the 3D dressed human body reconstruction network for training.

[0100] The parameter settings for the training process are as follows: the server graphics card is RTX 3090, the adaptive gradient method optimizer is Adam, the initial learning rate is 0.001, the learning rate of the network is dynamically adjusted during the training process, the batch size is 3, the number of epochs is 10, the number of sampling points for each character model is 5000, and in every 10,000 iterations, the learning rate decays by 0.1 times, and train until the loss function converges.

[0101] (5) Test the 3D dressed human body reconstruction network

[0102] Input the test set into the trained 3D dressed human body reconstruction network for testing, and evaluate according to the evaluation index to obtain the optimal reconstruction result.

[0103] (6) Display the reconstruction result

[0104] Input a single human picture to obtain the reconstruction result of the human body model.

[0105] Complete the single-view Figure 3 dressed human body reconstruction method based on the fusion of pixel and voxel features.

[0106] Embodiment 2

[0107] The single-view dressed human body reconstruction method based on the fusion of pixel and voxel features in this embodiment consists of the following steps: Figure 3 :

[0108] (1) Dataset preprocessing

[0109] This step is the same as that in Embodiment 1.

[0110] (2) Divide the dataset

[0111] This step is the same as that in Embodiment 1.

[0112] (3) Construct a three-dimensional dressed human body reconstruction network

[0113] This step is the same as that in Embodiment 1.

[0114] (4) Train the three-dimensional dressed human body reconstruction network

[0115] 1) Construct the depth blur reconstruction loss function

[0116] Construct the depth blur reconstruction loss function according to Equation (1)

[0117] The expression of Equation (1) is the same as that in Embodiment 1.

[0118] In Equation (1), n p is the number of point samples, n p ∈ [5000, 10000], and n in this embodiment p takes the value of 5000. The meanings and value ranges represented by other parameters and variables are the same as those in Embodiment 1.

[0119] 2) Construct evaluation metrics

[0120] The evaluation metrics consist of the chamfer distance d CD and the point-to-plane distance d PSD .

[0121] Construct the chamfer distance d CD according to Equation (2):

[0122] The expression of Equation (2) is the same as that in Embodiment 1.

[0123] In Equation (2), m = |P|, P is the target point set, P ∈ {p 1 , p 2 ,..., p m}, n = |Q|, Q is the point set obtained by reconstruction, Q ∈ {q 1 , q 2 ,..., q n}, m, n ∈ [5000, 10000], and m and n in this embodiment take the value of 5000. The meanings and value ranges represented by other parameters and variables are the same as those in Embodiment 1.

[0124] Other steps are the same as those in Embodiment 1. Complete the single-view Figure 3 dressed human body reconstruction method based on pixel and voxel feature fusion.

[0125] Embodiment 3

[0126] The single-view three-dimensional dressed human body reconstruction method based on the fusion of pixel and voxel features in this embodiment Figure 3 consists of the following steps:

[0127] (1) Dataset preprocessing

[0128] This step is the same as that in Embodiment 1.

[0129] (2) Dataset division

[0130] This step is the same as that in Embodiment 1.

[0131] (3) Construct a three-dimensional dressed human body reconstruction network

[0132] This step is the same as that in Embodiment 1.

[0133] (4) Train the three-dimensional dressed human body reconstruction network

[0134] 1) Construct a depth blur reconstruction loss function

[0135] Construct a depth blur reconstruction loss function according to Equation (1)

[0136] The expression of Equation (1) is the same as that in Embodiment 1.

[0137] In Equation (1), n p is the number of point samples, n p ∈[5000, 10000], and n p in this embodiment takes the value of 10000. The meanings and value ranges represented by other parameters and variables are the same as those in Embodiment 1.

[0138] 2) Construct evaluation indicators

[0139] The evaluation indicators consist of the chamfer distance d CD and the point-to-plane distance d PSD .

[0140] Construct the chamfer distance d CD according to Equation (2).

[0141] The expression of Equation (2) is the same as that in Embodiment 1.

[0142] In Equation (2), m = |P|, P is the target point set, P ∈ {p 1 , p 2 ,..., p m}}, n = |Q|, Q is the point set obtained by reconstruction, Q ∈ {q 1 , q 2 ,..., q n}, where \(m,n\in[5000,10000]\), and in this embodiment, \(m\) and \(n\) are taken as 10000. The meanings and value ranges represented by other parameters and variables are the same as those in Embodiment 1.

[0143] Other steps are the same as those in Embodiment 1. Complete the single-view Figure 3 dimensional dressed human body reconstruction method based on the fusion of pixel and voxel features.

[0144] To verify the beneficial effects of the present invention, the single-view Figure 3 dimensional dressed human body reconstruction method based on the fusion of pixel and voxel features in Embodiment 1 of the present invention and the Parametric Model-Conditioned Implicit Representation for Image-based Human Reconstruction method (referred to as the comparative experimental method) were used to conduct computer comparative simulation experiments. The experimental results are shown in Figure 7 . In Figure 7 , \(a\) represents the input single-person picture, \(b\) represents the reconstruction result of \(a\) by the comparative experimental method, and \(c\) represents the reconstruction result of \(a\) by the method in Embodiment 1 of the present invention; \(d\) represents the input single-person picture, \(e\) represents the reconstruction result of \(d\) by the comparative experimental method, and \(f\) represents the reconstruction result of \(d\) by the method in Embodiment 1 of the present invention. As Figure 7 can be seen, by using the method in Embodiment 1 of the present invention, the integrity of the limbs of the reconstructed three-dimensional human body model and the wrinkle details of the clothing are significantly improved.

Claims

1. A single-view 3D clothed human reconstruction method based on pixel and voxel feature fusion, characterized in that It consists of the following steps: (1) Dataset preprocessing The open source data set THUman2.0 from Tsinghua University is used. The THUman2.0 data set contains 526 3D scanned human models, each of which consists of a 3D file in obj format, a texture image, and a material file. The data set is preprocessed by the following process: Each 3D scanned human body model is normalized and randomly scaled; the light radiation transmission PRT is pre-calculated to improve the light realism; the 3D model is rotated 360° on the y-axis using the open graphics library OpenGL, and the original image, segmentation mask map, and texture map are collected every 1°; the pose estimation library OpenPose is used to generate the joint point information corresponding to each original image, and the human body parameterized model corresponding to the original image is generated using this joint point information; (2) Dividing the Dataset The preprocessed THUman2.0 dataset is divided into training set and test set according to 8:2; (3) Constructing a 3D clothed human body reconstruction network The 3D clothed human body reconstruction network is composed of a pixel feature extraction network, a voxel feature extraction network, a multi-layer perceptron, and a Marching Cubes module. The output ends of the pixel feature extraction network and the voxel feature extraction network are connected in series with the multi-layer perceptron and the Marching Cubes module in sequence. The pixel feature extraction network is composed of a coordinate attention module 1 and a convolution layer 1, a BN layer 1, a ReLU activation function layer 1, a coordinate attention module 2, a residual module 1, a maximum pooling layer 1, a coordinate attention module 3, a residual module 2, a coordinate attention module 4, a residual module 3, a stacked hourglass network module 1, a stacked hourglass network module 2, a stacked hourglass network module 3, a stacked hourglass network module 4, a fourth-order hourglass module 1, a residual module 4, a convolution layer 2, a BN layer 2, a ReLU activation function layer 2, and a convolution layer 3 connected in series in sequence; The voxel feature extraction network is composed of a voxelized human body parameterized model, a three-dimensional channel and a spatial attention module, and a multi-scale residual feature extraction module connected in series in sequence; (4) Training the 3D clothed human body reconstruction network 1) Constructing the depth blur reconstruction loss function According to formula (1), the depth blur reconstruction loss function is constructed C(p i +Δp i )=(S(F I ,π(p i +Δp i )),S(F V ,(p i +Δp i ))) T Δp i =(0,0,Δz i ) T F I =E I (I) F V =E V (V O ) Among them, n p is the number of point samples, n p is a finite positive integer, p i is a 3D point sample indexed by i, Δp i is the compensating translation along the z-axis, F * (p i ) is p i The true value of occupation value, C(p i +Δp i ) is the conditional implicit parameterization model, F I Represents the image from the deep image encoder E I (·) image feature map, I represents the input image, π(p i +Δp i ) is the point p after z-axis compensation i +Δp i In the image feature map F I The two-dimensional projection on I ,π(p i +Δp i )) represents the bilinear interpolation at pixel π(p i +Δp i ) at sampling F I The value of F V is the characteristic volume, S(F V ,(p i +Δp i )) indicates that from F V The sampled voxel-aligned volume features, E V For 3D encoder, V O is the occupied volume after the human body parameterized model is converted, F(C(p i +Δp i )) indicates that the point p is classified after z-axis compensation i +Δp i Whether it is inside or outside the surface, Δz i It is the displayed correspondence between the predicted human body parameterized model and the real model during network training, ω j→i is the corresponding mixing weight, p i -v j It's point p i The adjacent human body parametric model vertex v j The three-dimensional distance in 3D space, ω i is the weight normalization factor, v j and are the jth vertex of the predicted human parametric model and the jth vertex of the corresponding true human parametric model, σ is the standard deviation, Z(v j )and They are v in the camera coordinate space j and The depth value of 2) Constructing evaluation indicators The evaluation index is the chamfer distance d CD and the distance d from the point to the surface PSD composition; According to formula (2), the chamfer distance d is constructed CD : Among them, m = P, P is the target point set, P∈{p1,p2,,p m }, n = Q, Q is the reconstructed point set, Q∈{q1,q2,,q n }, m and n are finite positive integers, p i -q j It's point p i and dot q j The three-dimensional distance in 3D space, q j -p i It's Q j and point p i Three-dimensional distance in 3D space; Construct the distance d from the point to the surface as follows PSD : 3) Training the 3D clothed human body reconstruction network Input the training set into the 3D clothed human body reconstruction network for training; The parameters of the training process are set as follows: the server graphics card is RTX 3090, the adaptive gradient method optimizer is Adam, the initial learning rate is 0.001, the learning rate of the network is dynamically adjusted during training, the batch size is 3, the number of epochs is 10, the number of sampling points for each character model is 5000, and the learning rate is decayed by 0.1 times every 10,000 iterations. Training to the loss function convergence; (5) Testing the 3D clothed human body reconstruction network The test set is input into the trained 3D clothed human body reconstruction network for testing, and evaluated according to the evaluation indicators to obtain the optimal reconstruction result; (6) Display the reconstruction results Input a single person picture and get the human body model reconstruction result.

2. The single-view 3D clothed human body reconstruction method based on pixel and voxel feature fusion according to claim 1, characterized in that: In step (3) of constructing a human body reconstruction network, the structure and connection relationship of the stacked hourglass network module 1 and the stacked hourglass network module 2 are as follows: The stacked hourglass network module 1 is composed of a fourth-order hourglass network module 2, a residual module 5, a convolutional layer 4, a BN layer 3, a ReLU activation function layer 3, a convolutional layer 5, a Concat layer 1, a coordinate attention module 5, a convolutional layer 6, and a convolutional layer 7; the fourth-order hourglass network module 2 is connected in series with the residual module 5, the convolutional layer 4, the BN layer 3, the ReLU activation function layer 3, the convolutional layer 5, the Concat layer 1, and the coordinate attention module 5, the other output end of the ReLU activation function layer 3 is connected to the second input end of the Concat layer 1 through the convolutional layer 6 and the convolutional layer 7, the input end of the fourth-order hourglass network module 2 is connected to the third input end of the Concat layer 1, the other output end of the convolutional layer 7 is connected to the stacked hourglass network module 2, and the output end of the coordinate attention module 5 is connected to the stacked hourglass network module 2; The stacked hourglass network module 2 is composed of a fourth-order hourglass network module 3, a residual module 6, a convolutional layer 8, a BN layer 4, a ReLU activation function layer 4, a convolutional layer 9, a Concat layer 2, a coordinate attention module 6, a convolutional layer 10, and a convolutional layer 11; the fourth-order hourglass network module 3 is connected in series with the residual module 6, the convolutional layer 8, the BN layer 4, the ReLU activation function layer 4, the convolutional layer 9, the Concat layer 2, and the coordinate attention module 6, the other output end of the ReLU activation function layer 4 is connected to the second input end of the Concat layer 2 through the convolutional layer 10 and the convolutional layer 11, the input end of the fourth-order hourglass network module 3 is connected to the third input end of the Concat layer 2, the other output end of the convolutional layer 11 is connected to the stacked hourglass network module 3, and the output end of the coordinate attention module 6 is connected to the stacked hourglass network module 3; The structures of the stacked hourglass network module 3 and the stacked hourglass network module 4 are the same as the structure of the stacked hourglass network module 1; The connection relationship between the stacked hourglass network module 2 and the stacked hourglass network module 3 is the same as the connection relationship between the stacked hourglass network module 1 and the stacked hourglass network module 2, and the connection relationship between the stacked hourglass network module 3 and the stacked hourglass network module 4 is the same as the connection relationship between the stacked hourglass network module 1 and the stacked hourglass network module 2.

3. The single-view 3D clothed human body reconstruction method based on pixel and voxel feature fusion according to claim 1, characterized in that: In the step (3) of constructing the voxel feature extraction network of the human body reconstruction network, the multi-scale residual feature extraction module is composed of convolution layer 12, convolution layer 13, convolution layer 14, convolution layer 15, convolution layer 16, convolution layer 17, convolution layer 18, convolution layer 19, convolution layer 20, convolution layer 21, and Concat layer 3 connected together; convolution layer 13 is connected in series with convolution layer 14, convolution layer 15 is connected in series with convolution layer 16 and convolution layer 17 in sequence, convolution layer 18 is connected in series with convolution layer 19, convolution layer 20, convolution layer 21, and Concat layer 3 in sequence, the input ends of convolution layer 12, convolution layer 13, convolution layer 15, and convolution layer 18 are connected to the second input end of Concat layer 3, and the output ends of convolution layer 12, convolution layer 14, and convolution layer 17 are connected to the third input end of Concat layer 3.

4. The single-view 3D clothed human body reconstruction method based on pixel and voxel feature fusion according to claim 1, characterized in that: The convolution kernel size of the convolution layer 12 is 1×1 and the step size is 1. The sizes of the convolution layers 13, 15 and 18 are the same as those of the convolution layer 12. The convolution kernel size of the convolution layer 14 is 3×3, the step size is 1, the padding is 2, and the expansion is 2. The sizes of the convolution layers 16 and 19 are the same as those of the convolution layer 14. The convolution kernel size of the convolution layer 17 is 3×3, the step size is 1, and the padding is 1. The sizes of the convolution layers 20 and 21 are the same as those of the convolution layer 17.

5. The single-view 3D clothed human body reconstruction method based on pixel and voxel feature fusion according to claim 1, characterized in that: In step (4), in formula (1) of training the 3D clothed human body reconstruction network, the n p is the number of point samples, n p ∈[5000,10000]; In formula (2), m = P, P is the target point set, P∈{p1,p2,,p m }, n = Q, Q is the reconstructed point set, Q∈{q1,q2,,q n },m,n∈[5000,10000],p i -q j It's point p i and dot q j The three-dimensional distance in 3D space, q j -p i It's Q j and point p i Three-dimensional distance in 3D space.

6. The single-view 3D clothed human body reconstruction method based on pixel and voxel feature fusion according to claim 1 or 5, characterized in that: In step (4), in formula (1) of training the 3D clothed human body reconstruction network, the n p is the number of point samples, n p The value is 7000; In formula (2), m = P, P is the target point set, P∈{p1,p2,,p m }, n = Q, Q is the reconstructed point set, Q∈{q1,q2,,q n }, m and n are 7000, p i -q j It's point p i and dot q j The three-dimensional distance in 3D space, q j -p i It's Q j and point p i Three-dimensional distance in 3D space.

Citation Information

Cited By

  • Liver and tumor recognition method and device fusing attention mechanism and multi-scale convolution and readable storage medium thereof

    CN120563529A

  • Autonomous data set training-based nnUNet zebra fish juvenile fish whole cerebral vessel system segmentation method

    CN120997829A

  • nnunet segmentation method for zebrafish larva whole brain vasculature based on self-contained dataset training

    CN120997829B

  • Three-dimensional dressing human body generation method based on text driving

    CN122473365A

  • A text-driven based method for generating a three-dimensional clothed human

    CN122473365B