A 3D Human Body Reconstruction Method and System Based on Deep Learning
Through the human feature extraction model and three-dimensional reconstruction algorithm based on deep learning, combined with texture processing technology, the existing human body three-dimensional reconstruction technology is solved, and high-quality and efficient human body three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202411007528.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-07-25
AI Technical Summary
The human body models generated by existing three-dimensional reconstruction techniques are of low quality and difficult to apply on a large scale, especially in complex backgrounds.
The human feature extraction model, occupancy network and traveling cube modeling algorithm based on deep learning are used, and the target image is processed to generate a textured three-dimensional model of the human body.
Improves the quality and accuracy of the 3D reconstruction model of the human body, can handle non-frontal images and complex backgrounds, reduces analysis time and improves reconstruction efficiency.
Smart Images

Figure CN119006742B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional reconstruction technology, and in particular to a method and system for human three-dimensional reconstruction based on deep learning. Background Art
[0002] Human three-dimensional reconstruction technology is a technology that uses computer vision and image processing technologies to reconstruct the three-dimensional geometric shape and surface texture of the human body through information extracted from two-dimensional images or videos. It is widely used in fields such as computer graphics, human-computer interaction, virtual reality, augmented reality, medical imaging, virtual digital humans, and human comfort analysis.
[0003] Currently, human three-dimensional reconstruction technology mainly includes:
[0004] 1. First, construct a human pose estimation and three-dimensional reconstruction network in advance. Crop the sub-image containing the human body from the training pictures through the object detection and segmentation network, and determine the global position parameters of the sub-image relative to the training picture. Then, predict the pose-related parameters of the human body in the sub-image through the position parameter prediction network. Next, construct a three-dimensional human body model based on the predicted pose-related parameters through the human body model construction network, and then use the two-dimensional projection model to project the constructed three-dimensional human body model onto a two-dimensional plane. Finally, train the parameter prediction network, the human body model construction network, and the two-dimensional projection model based on the obtained two-dimensional projection results and the constructed three-dimensional human body model. The disadvantage of this technology is that it has high requirements for the background of the training pictures. When the background is complex, a cumbersome human body recognition process is required, and the accuracy of human body recognition cannot be guaranteed;
[0005] 2. Build a digital human (skinned multi-person linear, SMPL) model for the human body based on vertex offset to obtain the lengths of the limbs and bones of the human body. Then, add a new camera to the camera array. According to the lengths of the limbs and bones and the 3D human body joint points obtained from the initial binocular camera intersection and the 2D human body joint points under the new camera, continuously iterate and optimize the calibration relationship between the new camera and the camera array to obtain optimized 3D human body joint points. Calculate the distances between the human body joint points in the optimized 3D human body joint points to optimize the morphological parameters and obtain the best morphological parameters. Secondly, calculate the pose parameters in the SMPL model according to the changes in the relative rotation angles of the optimized 3D human body joint points, solve the problems of the complex external calibration process of the multi-camera system and the error coupling in human body pose estimation, and estimate a more accurate human body pose. The disadvantage of this technology is that it has high requirements for equipment, requires the construction of a complex camera array, is difficult to be applied on a large scale, and the generation process requires continuous iteration and update, which takes a long time;
[0006] 3. Input the human body image into the global encoder to obtain the first body parameters and global two-dimensional features. Input the first hand feature, the first head feature, and the human body image into the local decoder to obtain the first hand parameters and the first head parameters. The first hand feature and the first head feature are separated from the global two-dimensional features. Then, input the first body parameters, the first hand parameters, and the first head parameters into the component interaction module for component interaction to obtain the three-dimensional human body reconstruction result. The disadvantage of this technology is that its network model belongs to the traditional self-attention network method of directly regressing parameters, with insufficient learning of discriminative features, resulting in a poor effect of the generated human body model and the need for cumbersome processing of different parts of the human body, leading to a long generation time.
[0007] It can be seen from this that the current three-dimensional human body reconstruction technology requires multiple high-precision frontal photos of the human body, or requires cumbersome human body recognition methods for complex background environments, with a large amount of calculation and a long reconstruction time. In addition, the traditional self-attention network method of directly regressing parameters has insufficient learning of discriminative features, resulting in a low quality of the generated human body model, which also makes it difficult for the three-dimensional human body reconstruction technology to be widely applied. Summary of the Invention
[0008] To solve the problem of the low quality of the human body model generated by the three-dimensional human body reconstruction technology, this application provides a three-dimensional human body reconstruction method and system based on deep learning.
[0009] In the first aspect of this application, a three-dimensional human body reconstruction method based on deep learning is provided. The method includes:
[0010] Receive a target image, where the target image is an image containing the human body to be reconstructed;
[0011] Process the target image using a human body feature extraction model to obtain target human body features, where the target human body features are used to reflect the human body to be reconstructed in the target image;
[0012] Process the target human body features using a three-dimensional human body reconstruction model to obtain an initial three-dimensional human body model;
[0013] Process the target image using a texture acquisition model to obtain two-dimensional human body textures;
[0014] Combine the initial three-dimensional human body model and the human body textures using a texture mapping model to obtain a target three-dimensional human body model, where the target three-dimensional human body model is a three-dimensional human body model with textures.
[0015] By adopting the above technical solution, first, the human feature extraction model receives the target image and extracts the target person's features, and then the human three-dimensional reconstruction model constructs an initial human three-dimensional model based on the target person's features. At the same time, the texture acquisition model is also used to process the target image to obtain the two-dimensional human texture of the human body to be reconstructed. Finally, the initial human three-dimensional model and the two-dimensional human texture are combined to obtain the target human three-dimensional model. It can be seen that multiple models are set in advance in this application. Through the cooperation of multiple models, it is possible to receive target images with non-frontal human bodies and complex backgrounds, and generate a textured human three-dimensional model based on the target image, improving the quality of the obtained human three-dimensional model.
[0016] In a possible implementation manner: before using the human feature extraction model to process the target image to obtain the target human features, the method includes: training the human feature extraction model;
[0017] The method for training the human feature extraction model includes:
[0018] Obtain a human image set, where the human image set includes multiple training images, and the training images are images containing human bodies; calculate the training features of each training image, where the training features are data predicted by the human feature extraction model and used to reflect the human body in the training image;
[0019] Obtain the actual features of each training image, where the actual features are data that truly reflect the human body in the training image; construct a loss function according to the training features and actual features of multiple training images;
[0020] When training the human feature extraction model according to the loss function until the value output by the loss function is less than a preset value, output the human feature extraction model.
[0021] By adopting the above technical solution, first construct a loss function according to the training features and actual features of the training images, and then train the human feature extraction model according to the loss function until the value output by the loss function is less than a preset value, and then output the trained human feature extraction model, so as to ensure the accuracy of the extraction results of the human feature extraction model.
[0022] In a possible implementation manner: the calculation of the training features of each training image set includes:
[0023] Use the perception model to process the training image to obtain the first feature;
[0024] Linearly splice the training image and the first feature to obtain the second feature;
[0025] Use the channel selection model to process the second feature to obtain the third feature;
[0026] Linearly splice the first feature and the third feature to obtain a fourth feature;
[0027] Process the fourth feature using a feature interaction model to obtain a fifth feature;
[0028] Linearly splice the third feature and the fifth feature to obtain a training feature.
[0029] By adopting the above technical solution, using a perception model, a channel selection model, and a feature interaction model according to the context information of the training image, and also adopting a multi-level and multi-angle feature extraction and fusion technology, adaptively learn the important features related to the human body on the training image, and improve the accuracy of the extraction result of the human feature extraction model obtained by training.
[0030] In a possible implementation: the perception model includes a convolutional layer, an activation function layer, and an attention layer;
[0031] The convolutional layer is used to access the training image and perform convolution on the training image;
[0032] The attention layer is connected to the convolutional layer, and the attention layer is used to extract the key features in the training image, and the key features are used to reflect the human body in the training image;
[0033] The activation function layer is arranged between the convolutional layer and the attention layer, and the activation function layer is used to adjust the weights of the key features in the training image.
[0034] By adopting the above technical solution, the perception model can increase the weights of the key features in the training image, so that the key features can be more effectively captured in the subsequent processing process, providing data support for promoting the training features to continuously approach the actual features, thereby improving the performance of the human feature extraction model.
[0035] In a possible implementation: the obtaining of the human body image set includes:
[0036] Receive multiple training images;
[0037] Add Gaussian noise to each training image and crop the invalid area, where the invalid area is the area that does not contain the human body, to obtain the final training image;
[0038] The set of the final training images is used as the human body image set.
[0039] By adopting the above technical solutions, in order to improve the generalization performance of the human feature extraction model obtained by training, this application introduces preprocessing steps of regional noise injection and cropping invalid regions, which helps the human feature extraction model capture more local details and process occluded parts, so that the human feature extraction model can extract training features from the training images even if the received training images are non-frontal views of the human body with complex backgrounds.
[0040] In a possible implementation: the human three-dimensional reconstruction model is configured with an occupancy network and a marching cubes modeling algorithm;
[0041] The process of using the human three-dimensional reconstruction model to process the target human features to obtain an initial human three-dimensional model includes:
[0042] Using the occupancy network to calculate the occupancy probability of sampling points in the three-dimensional space according to the target human features, where the three-dimensional space is a pre-set three-dimensional space to be reconstructed, and the sampling points are the vertices of the grid cells after dividing the three-dimensional space into multiple grid cells;
[0043] Retaining the sampling points with an occupancy probability greater than a preset probability;
[0044] Using the marching cubes modeling algorithm to construct an initial human three-dimensional model based on the retained sampling points.
[0045] In a possible implementation: the occupancy probability of the sampling points is calculated through the following calculation formula: f θ (F 目 , V) = τ, where θ is the parameter of the occupancy network, F 目 is the target human feature, V is the sampling point, and τ is the occupancy probability of the sampling point V.
[0046] In a possible implementation: the process of using the marching cubes modeling algorithm to construct an initial human three-dimensional model based on the retained sampling points includes:
[0047] Constructing a new three-dimensional space based on the retained sampling points;
[0048] Dividing the new three-dimensional space to obtain multiple cube units, where each cube unit contains a voxel point or a voxel value, and both the voxel point and the voxel value are used to reflect the target human features;
[0049] Extracting an isosurface, which is composed of multiple cube units with the same voxel points or the same voxel values;
[0050] Locking the intersection points of the isosurface and each cube unit that makes up the isosurface;
[0051] Use a triangulation method to connect the intersection points to form a triangular mesh;
[0052] Combine the triangular meshes to obtain the initial 3D human model.
[0053] By adopting the above technical solutions, the occupancy network has certain advantages in tasks such as 3D reconstruction and 3D object recognition. Especially in the 3D reconstruction task from a single perspective, it can simplify the data acquisition and processing process, and can better handle problems such as occlusion and perspective change, providing technical guarantee for obtaining an accurate 3D human model of the target. The marching cubes modeling algorithm is a classic and effective 3D surface reconstruction algorithm, which can process various types of volume data and has good robustness and performance at different resolutions, further providing technical guarantee for obtaining an accurate 3D human model of the target.
[0054] In a possible implementation: the texture mapping model is configured with UV mapping;
[0055] Combining the initial 3D human model and the human texture by using the texture mapping model to obtain the target 3D human model includes:
[0056] Use the UV mapping to map the initial 3D human model to the UV coordinates;
[0057] Obtain the UV coordinates corresponding to the human texture;
[0058] Overlap the UV coordinates corresponding to the initial 3D human model and the UV coordinates corresponding to the human texture to form a reconstructed UV coordinate;
[0059] Use inverse UV mapping to map the reconstructed UV coordinates to the 3D model to obtain the target 3D human model.
[0060] By adopting the above technical solutions, the initial 3D human model and the human texture are respectively mapped to different UV coordinates, then the two UV coordinates are overlapped to obtain a UV coordinate, and finally the obtained UV coordinate is converted to the 3D space by inverse UV mapping, so as to obtain a textured 3D human model.
[0061] In the second aspect of the present application, a 3D human reconstruction system based on deep learning is provided. The system includes:
[0062] A data acquisition module for receiving a target image, where the target image is an image containing the human body to be reconstructed;
[0063] The first processing module is configured to process the target image using a human feature extraction model to obtain target human features, where the target human features are used to reflect the human body to be reconstructed in the target image;
[0064] The second processing module is configured to process the target human features using a human three-dimensional reconstruction model to obtain an initial human three-dimensional model; The third processing module is configured to process the target image using a texture acquisition model to obtain a two-dimensional human texture;
[0065] The data combination module is configured to combine the initial human three-dimensional model and the human texture using a texture mapping model to obtain a target human three-dimensional model, where the target human three-dimensional model is a textured human three-dimensional model.
[0066] In summary, the present application includes at least one of the following beneficial technical effects:
[0067] First, a perception model, a channel selection model, and a feature interaction model are used to adaptively learn important features related to the human body on the training image according to the context information of the training image, and a multi-level and multi-angle feature extraction and fusion technology is used to improve the accuracy of the extraction results of the human feature extraction model obtained by training. Based on the trained human feature extraction model, the human feature extraction model receives the target image and extracts the target person features, and then the occupancy network and marching cubes modeling algorithm in the human three-dimensional reconstruction model are used to construct an initial human three-dimensional model according to the target person features. Secondly, a texture acquisition model is also used to process the target image to obtain a two-dimensional human texture of the human body to be reconstructed. Finally, the initial human three-dimensional model is mapped to the UV coordinates, the initial human three-dimensional model located on the UV coordinates and the two-dimensional human texture are combined, and then the reverse UV mapping is performed to obtain a textured human three-dimensional model. It can be seen that based on the human feature extraction model with guaranteed accuracy, regardless of whether the input target image is a non-front view of the human body to be reconstructed or the background is complex, more local details in the target image can be captured and the occluded parts can be processed to obtain accurate target features, thereby providing data support for obtaining a high-quality target human three-dimensional model subsequently. In addition, a variety of current mature models are used to reconstruct the target human three-dimensional model based on the target features, reducing the analysis time, improving the reconstruction efficiency, and further improving the quality of the generated target human three-dimensional model. Description of the Drawings
[0068] Figure 1 is a flowchart of a method for human three-dimensional reconstruction based on deep learning according to an embodiment of the present application.
[0069] Figure 2 is an example diagram of the composition structure of the perception model in the method embodiment of the present application.
[0070] Figure 3It is an exemplary diagram of the composition structure of the channel selection model in the method embodiment of the present application.
[0071] Figure 4 It is an exemplary diagram of the composition structure of the feature interaction model in the method embodiment of the present application.
[0072] Figure 5 It is a block diagram of a three-dimensional human body reconstruction system based on deep learning in the embodiment of the present application.
[0073] Explanation of reference numerals: 1. Data acquisition module; 2. First processing module; 3. Second processing module; 4. Third processing module; 5. Data combination module; 6. Data training module. Detailed implementation manners
[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0075] In order to accurately capture human body features when the input human body image is non-frontal or has a complex background, the present application introduces a deep learning algorithm, uses a deep neural network and a large-scale data set to adaptively learn the relationship between the region where the human body is located and the human body features in the human body image, and improves the accuracy of three-dimensional human body reconstruction. As Figure 1 shown, the embodiment of the present application proposes a three-dimensional human body reconstruction method based on deep learning, and the main process of this method is described as follows.
[0076] Step S100: Receive a target image, where the target image is an image containing the human body to be reconstructed.
[0077] For the user's three-dimensional human body reconstruction requirement, the image selected by the user and containing the human body to be reconstructed is set as the target image, and this target image is received and the following steps are performed.
[0078] Step S200: Process the target image with a human body feature extraction model to obtain target human body features, where the target human body features are used to reflect the human body to be reconstructed in the target image.
[0079] The human body feature extraction model is a pre-trained model, which can automatically identify and extract the target human body features used to reflect the human body to be reconstructed when receiving the target image. Specifically, the training process of the human body feature extraction model is as shown in Step S210 - Step S250.
[0080] Step S210: Obtain a human body image set, which includes multiple training images. The training images are images containing a human body. Specifically, the human body image set can be sourced from open-source data sets, such as the AFLW2000-3D database, the AFLW face database, or the Helen face database, or it can also be sourced from a database organized or supplemented by experts.
[0081] Based on the obtained human body image set, perform regional noise injection on each training image in the human body image set, specifically by adding Gaussian noise for data augmentation. Then, screen out the training images with a relatively large area of invalid regions in the human body image set. The invalid regions are regions that do not contain a human body. Crop the invalid regions of the training images and only retain the regions where the human body is located in the training images. Finally, merge the same training images as the final training images to make the obtained training images contain more human body information.
[0082] Step S220: Calculate the training features of each training image.
[0083] First, use a perception model to process the training image to obtain a first feature. As Figure 2 shown, the perception model includes a convolutional layer, an activation function layer, and an attention layer. The convolutional layer is used to receive the training image and perform a convolution operation on the training image. In this example, there are multiple convolutional layers, that is, there are multiple layers of convolutional layers in the perception model, and each layer of convolutional layer is used to perform a convolution operation on the training image. In this example, the activation function layer uses a non-linear activation function (rectified linear unit, ReLU). The ReLU activation function is one of the commonly used activation functions in deep learning. It can introduce non-linear factors into the neural network of deep learning, enabling the neural network to learn and represent more complex functions. Here, the ReLU activation function is mainly used in cooperation with the attention layer. Specifically, the attention layer identifies and extracts the key features in the training image that reflect the human body, and the ReLU activation function is used to adjust the weights of the key features in the training image.
[0084] The attention layer adopts the attention mechanism (convolutional block attention module, CBAM), also known as the CBAM attention layer. Among them, the CBAM attention layer is further divided into a channel attention module (CAM) and a spatial attention module (SAM). The channel attention module aims to extract the correlation between different channels in the training image. The channel represents the output dimension of the training image. The working process of the channel attention module is as follows: the global average pooling (GAP) technology is used to capture the global information carried by each channel. Then, according to the global information carried by each channel, the ReLU activation function is used to adjust the weight of each channel. For example, if the global information carried by channel a is more important than the global information carried by channel b, the weight of channel a is increased. The importance of the global information is predefined. Then, the weights are used to perform weighted summation on the training image to generate a feature representation with stronger correlation on the training image. The spatial attention module is connected in series with the channel attention module. Based on the processing result of the channel attention module, the spatial attention module further highlights the correlation between different positions on the training image. The position represents the pixel unit on the training image, also known as the pixel position. The working process of the spatial attention module is as follows: the channel weights output by the channel attention module are used to perform a one-dimensional convolution operation on the training image to capture the correlation between different positions on the training image. Then, the built-in activation function is used to normalize the convolution output to the range of [0,1], so as to obtain the weight of each pixel position on the training image. It should be noted that the built-in activation function of the spatial attention module adopts the Sigmoid activation function.
[0085] Combining the channel attention module and the spatial attention module can simultaneously focus on the correlations between multiple channels and the correlations between spatial positions on the training image, enabling the subsequent human feature extraction model to more effectively capture the important features related to the human body on the training image and improve the model performance. In this example, the features identified and extracted by the attention layer are referred to as the first features, denoted as F1. The first feature F1 is the feature identified after adjusting the channel weights and spatial weights. Simply put, the training image passes through the convolutional layer, and the convolutional layer extracts local features and outputs them as feature out11. The ReLU activation function layer promotes the non-linear transformation of feature out11 to obtain feature out12. The CBAM attention layer identifies the global information and position information of the channels on the training image, then adjusts the channel weights of the training image through the ReLU activation function, and adjusts the spatial weights of the training image through the built-in activation function to change feature out12 and generate the first feature F1. It should be noted that feature out11 - feature out12 refers to the features extracted from the training image at different processing stages and is not used to limit the data output by the perception model.
[0086] The first feature F1 is composed of F ij . For example, if i ≤ 2 and j ≤ 2, then F = (F 11 , F 12 , F 21 , F 22 ). This combination method can also be in matrix form. Specifically, the calculation formula of F ij is as follows:
[0087]
[0088] Among them, Q and K respectively refer to Query and Key in the attention layer, and both Q and K are in matrix form. The Exp(,) function means that for each Query, the similarity between the Query and all Keys is calculated, and the similarity is normalized to obtain the attention weights. F ij represents the attention weight learned by the CABM attention layer between the i-th feature vector of Q and the j-th feature vector of K based on feature out12, and N represents the number of feature vectors in the feature matrix of K.
[0089] Then, based on the obtained first feature F1, the training image and the first feature F1 are linearly concatenated to obtain the second feature F2. In this example, linear concatenation is a method of feature fusion, which means simply connecting multiple features together to form a larger feature vector or feature matrix.
[0090] After that, based on the obtained second feature F2, the channel selection model is used to process the second feature F2 to obtain the third feature F3. AsFigure 3 As shown, in the direction from the input end to the output end of the channel selection model, there are, in sequence, a first convolutional layer, an average pooling layer, a second convolutional layer, a first activation layer, a third convolutional layer, a first concatenation layer, and a second activation layer. The ReLU activation function is used for both the first activation layer and the second activation layer. Specifically, the working process of the channel selection model is as follows: The first convolutional layer performs a convolutional operation on the second feature F2 to obtain the feature out21. The feature out21 is successively processed by the average pooling layer, the second convolutional layer, and the first activation layer to obtain the feature out22. The third convolutional layer convolves the feature out21 and the feature out22 to obtain the feature out23. The first concatenation layer uses the linear concatenation method to linearly concatenate the feature out21 and the feature out23 to obtain the feature out24. Finally, the second activation function processes the feature out24 to obtain the third feature F3.
[0091] It should be noted that the above features out21 - out24 refer to the second feature F2 at different processing stages and are not used to define the data output by the channel selection model.
[0092] In addition, based on the obtained third feature F3, the first feature F1 and the third feature F3 are linearly concatenated to obtain the fourth feature F4. The linear concatenation used here is also a method of feature fusion, which also means simply connecting multiple features together to form a larger feature vector or feature matrix.
[0093] Finally, based on the fourth feature F4, the feature interaction model processes the fourth feature F4 to obtain the fifth feature F5. As Figure 4 shown, in the direction from the input end to the output end of the feature interaction model, there are, in sequence, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a second concatenation layer, a normalization layer, and a third concatenation layer. The linear concatenation method is used for both the second concatenation layer and the third concatenation layer. Specifically, the working process of the feature interaction model is as follows: The fourth convolutional layer performs a convolutional operation on the fourth feature F4 to obtain three identical feature matrices, denoted as X, Y, and Z respectively. The calculation formulas for the feature matrices X, Y, and Z are all as follows: where H and W are the height and width of the training image, respectively, and X ijis the value at position (i, j) on the training image. The fifth convolutional layer performs convolution on the feature matrices Y and Z, and then the sixth convolutional layer performs convolution again based on the output of the fifth convolutional layer. That is, after the convolution of the feature matrices Y and Z, convolution is performed with the feature matrix X. The second concatenation layer linearly concatenates the feature matrix X and the data output by the sixth convolutional layer to obtain an attention map, which is denoted by A. Then A = X + (Y × Z) × X. The normalization layer uses the Softmax function to decompose the attention map A into a horizontal attention map and a vertical attention map, which are denoted by AW and AH respectively. The third concatenation layer performs weighted calculation on the fourth feature F4, the horizontal attention map AW, and the vertical attention map AH to obtain the fifth feature F5, that is, F5 = F4 + AW × ω1 + AH × ω2, where ω1 and ω2 are the weights of the horizontal attention map AW and the vertical attention map AH respectively, and both ω1 and ω2 are predefined in advance.
[0094] Based on the obtained fifth feature F5, the third feature F3 and the fifth feature F5 are linearly concatenated to obtain a training feature, which is denoted by F 训 Here, the linear concatenation adopted is also a method of feature fusion, which also means simply connecting multiple features together to form a larger feature vector or feature matrix.
[0095] It should be noted that through the above calculation process, the training features of each training image in the human body image set can be calculated, and after calculating the training features of all training images, it enters step S240. In actual training, the more training images used, the higher the accuracy of the subsequent trained human feature extraction model. Therefore, all training images in the human body image set are used in this example. In other examples, the number of training images used can also be set according to requirements.
[0096] Step S230: Obtain the actual features of each training image. When establishing the human body image set in step S210, each collected training image has its own real human body features, which are also called actual features, and the actual features are denoted by F 实 So when there is a need for the actual feature F 实 it can be directly called from the human body image set.
[0097] Step S240: Construct a loss function according to the training features and actual features of multiple training images. Specifically, the calculation formula of the loss function is: where the superscript T is used to indicate the transpose operation of the matrix, and this representation method is used for row-column interchange.
[0098] Step S250: When the human feature extraction model is trained according to the loss function until the value output by the loss function is less than a preset value, output the human feature extraction model. Specifically, when the value output by the loss function is less than the preset value, it indicates that the training features obtained by the human feature extraction model have approached or equal to the actual features. Therefore, output the trained human feature extraction model.
[0099] Based on the trained human feature extraction model, input the target image into the human feature extraction model. The human feature extraction model automatically identifies and extracts the target human features used to reflect the human body to be reconstructed, and then proceeds to step S300.
[0100] Step S300: Use the human three-dimensional reconstruction model to process the target human features to obtain an initial human three-dimensional model.
[0101] The human three-dimensional reconstruction model uses the Occupancy Networks. The Occupancy Networks is a method based on deep learning for directly reconstructing three-dimensional objects from single or multiple two-dimensional images. Traditional human three-dimensional reconstruction methods usually require image sequences taken from multiple perspectives or three-dimensional point cloud data obtained using sensors such as laser scanning. However, the Occupancy Networks attempts to directly learn the representation of three-dimensional objects from single or multiple two-dimensional images without explicitly reconstructing from multiple perspectives or using other sensors. The basic idea of the Occupancy Networks is to divide the three-dimensional space into a grid (usually a cubic grid), and then use a neural network to predict whether each grid cell is occupied by an object. In this way, a representation of the three-dimensional object can be established in the three-dimensional space without explicitly reconstructing the geometry or surface.
[0102] Therefore, after receiving the target human features, the specific working process of the Occupancy Networks is as follows: Preset the three-dimensional space to be reconstructed, divide the three-dimensional space into multiple identical grid cells, and use one vertex of each grid cell as a sampling point. The sampling point is represented by V. Therefore, a grid cell has four sampling points. Then, based on the target human features, predict the occupancy probability of the sampling point V in the three-dimensional space. The calculation formula for the occupancy probability is as follows:
[0103] f θ (F 目 ,V) = τ, where θ is the parameter of the Occupancy Networks, F 目 is the target human feature, τ is the occupancy probability of the sampling point V, τ is between 0 and 1. When τ is closer to 1, it indicates that the possibility of this sampling point being occupied is greater.
[0104] Then, retain the sampling points with the occupancy probability τ greater than the preset probability. The preset probability is a predefined probability value.
[0105] Finally, the remaining sampled points are collected, and the Marching Cube modeling algorithm is used to construct a three-dimensional human model based on the set of resampled points. The three-dimensional human model obtained here is called the initial three-dimensional human model. The Marching Cube modeling algorithm is a classic and effective three-dimensional surface reconstruction algorithm that can handle various types of human data and has good robustness and performance at different resolutions. Due to its simple and intuitive implementation and good results, the Marching Cube modeling algorithm has been widely used in the fields of three-dimensional modeling and visualization.
[0106] The working principle of the above-mentioned Marching Cube modeling algorithm is as follows: First, divide the three-dimensional space: The three-dimensional space is divided into a series of cubic cells. Since this step is to further generate a three-dimensional human model based on the remaining sampled points, when dividing the three-dimensional space, it is a re-division of the three-dimensional space composed of multiple grids where the occupancy probability of the sampled points is greater than the preset probability. So each cubic cell contains a voxel or voxel value, and both the voxel and voxel value are representations of the target human features. Then, extract the isosurface: Traverse each cubic cell, and determine the isosurface (i.e., the surface where the value is equal to or exceeds the threshold) by comparing the value inside the cubic cell with the preset threshold. Usually, interpolation methods are used to estimate the position of the isosurface. Secondly, perform triangulation: For each isosurface, determine the multiple cubic cells that make up the isosurface, lock the intersection points of the isosurface and each cubic cell that makes up the isosurface, and use the triangulation method to connect the intersection points to form a triangular mesh. The obtained triangular mesh finally forms the mesh representation of the entire surface. Finally, merge the surface meshes: Merge the triangular meshes in the space into an overall surface mesh to obtain the initial three-dimensional human model. Simply put, the Marching Cube modeling algorithm connects the intersection points of the isosurface and the small cubes in multiple small cubes to form one or more triangular meshes, and the triangular meshes in all small cubes are connected to form the initial three-dimensional human model.
[0107] Step S400: Use the texture acquisition model to process the target image to obtain a two-dimensional human texture.
[0108] The texture acquisition model is configured with UV mapping, and here the UV mapping is to map the two-dimensional target image onto the surface with the initial three-dimensional human model as the object. Specifically, the texture acquisition model first identifies the texture of the human body to be reconstructed on the target image, removes impurities or noises to obtain a two-dimensional human texture, and then uses UV mapping to map the two-dimensional human texture onto the UV coordinates.
[0109] Step S500: Use the texture mapping model to combine the initial three-dimensional human model and the human texture to obtain the target three-dimensional human model, and the target three-dimensional human model is a three-dimensional human model with texture.
[0110] The texture mapping model is also configured with UV mapping. Here, the UV mapping maps each vertex on multiple cube units that make up the initial three-dimensional human model to a UV coordinate on a two-dimensional plane, thereby obtaining a flattened human body, that is, a two-dimensional human body image.
[0111] Specifically, the working process of the texture mapping model is as follows: First, the initial three-dimensional human model is mapped to UV coordinates by using UV mapping, that is, the triangular mesh obtained by the above-mentioned marching cubes modeling algorithm, and the UV coordinates of the vertices of each triangular mesh are mapped to a two-dimensional plane to form a UV coordinate. Then, the UV coordinates corresponding to the initial three-dimensional human model and the UV coordinates corresponding to the two-dimensional human texture are overlapped to form a reconstructed UV coordinate, and the reconstructed UV coordinate contains the human body to be reconstructed and the human texture. Finally, the reconstructed UV coordinate is mapped back to the three-dimensional model by using inverse UV mapping, thereby obtaining a three-dimensional human model with texture. In this example, the three-dimensional human model with texture is called the target three-dimensional human model.
[0112] In summary, the implementation principle of a three-dimensional human body reconstruction method based on deep learning in an embodiment of the present application is as follows: First, a perception model, a channel selection model, and a feature interaction model are used to focus on multiple dimensions (such as channel information and spatial information) in the training image according to the context information of the training image, and a multi-level and multi-angle feature extraction and fusion technology is adopted to adaptively learn the important features related to the human body on the training image, so as to improve the accuracy of the extraction result of the human body feature extraction model obtained by training. Based on the trained human body feature extraction model, the human body feature extraction model receives the target image and extracts the target person's features, and then the occupancy network and the marching cubes modeling algorithm in the three-dimensional human body reconstruction model are used to construct an initial three-dimensional human model according to the target person's features. Secondly, a texture acquisition model is also used to process the target image to obtain a two-dimensional human texture of the human body to be reconstructed. Finally, the initial three-dimensional human model is mapped to UV coordinates, the initial three-dimensional human model located on the UV coordinates and the two-dimensional human texture are combined, and then inverse UV mapping is used to obtain a three-dimensional human model with texture. It can be seen that based on the human body feature extraction model with guaranteed accuracy, no matter whether the input target image is a non-front view of the human body to be reconstructed or the background is complex, it can capture more local details in the target image and process the occluded parts to obtain accurate target features, thereby providing data support for obtaining a high-quality target three-dimensional human model subsequently. In addition, the current mature technology is used to reconstruct the target three-dimensional human model based on the target features, which reduces the analysis time, improves the reconstruction efficiency, and further improves the quality of the generated target three-dimensional human model.
[0113] Figure 5The block diagram of a 3D human body reconstruction system based on deep learning according to an embodiment of the present application is shown. The system includes a data acquisition module 1, a first processing module 2, a second processing module 3, a third processing module 4, and a data combination module 5.
[0114] The data acquisition module 1 is configured to receive a target image, where the target image is an image containing the human body to be reconstructed.
[0115] The first processing module 2 is configured to process the target image using a human feature extraction model to obtain target human body features, and the target human body features are used to reflect the human body to be reconstructed in the target image.
[0116] The second processing module 3 is configured to process the target human body features using a human 3D reconstruction model to obtain an initial human 3D model.
[0117] The third processing module 4 is configured to process the target image using a texture acquisition model to obtain a 2D human body texture.
[0118] The data combination module 5 is configured to combine the initial human 3D model and the human body texture using a texture mapping model to obtain a target human 3D model, and the target human 3D model is a textured human 3D model.
[0119] In addition, the 3D human body reconstruction system based on deep learning further includes a data training module 6. The data training module 6 is configured to train the human feature extraction model based on a human body image set, and the trained human feature extraction model can be called by the first processing module 2.
[0120] The modules described in the embodiments of the present application can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a data acquisition module 1, a first processing module 2, a second processing module 3, a third processing module 4, a data combination module 5, and a data training module 6. Among them, the names of these modules do not limit the modules themselves in some cases. For example, the data acquisition module 1 can also be described as "a module for receiving a target image".
[0121] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0122] To better execute the program of the above method, the present application further provides a device, and the device includes a memory and a processor.
[0123] Among them, the memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. The program storage area can store instructions for implementing an operating system, instructions for at least one function, and instructions for implementing the above-mentioned deep learning-based human three-dimensional reconstruction method, etc.; the data storage area can store data involved in the above-mentioned deep learning-based human three-dimensional reconstruction method, etc.
[0124] The processor may include one or more processing cores. By running or executing instructions, programs, code sets or instruction sets stored in the memory, the processor calls the data stored in the memory and executes various functions of this application and processes the data. The processor may be at least one of an application specific integrated circuit, a digital signal processor, a digital signal processing device, a programmable logic device, a field programmable gate array, a central processing unit, a controller, a microcontroller and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above-mentioned processor functions may also be others, and the embodiments of this application do not make specific limitations.
[0125] This application also provides a computer-readable storage medium, for example, including: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical discs. The computer-readable storage medium stores a computer program that can be loaded and executed by the processor to implement the above-mentioned deep learning-based human three-dimensional reconstruction method.
[0126] The above description is only a preferred embodiment of this application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A method for three-dimensional reconstruction of human body based on deep learning, characterized in that: include: receiving a target image, wherein the target image is an image containing a human body to be reconstructed; Acquire a human body image set, and use the human body image set to train a human body feature extraction model to obtain a trained human body feature extraction model; The human body image set includes a plurality of training images, wherein the training images are images containing human bodies. When the human body image set is used to train the human body feature extraction model, the training features of each training image are obtained, and the process is as follows: the training image is processed by using the perception model to obtain a first feature; Linearly concatenate the training image and the first feature to obtain a second feature; process the second feature using a channel selection model to obtain a third feature; linearly concatenate the first feature and the third feature to obtain a fourth feature; Processing the fourth feature using a feature interaction model to obtain a fifth feature; Linearly concatenate the third feature and the fifth feature to obtain a training feature; The perception model includes a convolution layer, an activation function layer and an attention layer; The attention layer is divided into a channel attention module and a spatial attention module. The channel attention module extracts the correlation between different channels in the training image, and the channel represents the output dimension of the training image. The channel selection model is composed of the first convolution layer, the average pooling layer, the second convolution layer, the first activation layer, the third convolution layer, the first splicing layer and the second activation layer from the input end to the output end. The feature interaction model includes a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a second splicing layer, a normalization layer and a third splicing layer, wherein the process of using the feature interaction model to process the fourth feature to obtain the fifth feature is: using the fourth convolution layer to perform a convolution operation on the fourth feature to obtain three identical feature matrices, represented by X, Y, and Z respectively, the fifth convolution layer convolves the feature matrix Y and the feature matrix Z, and then the sixth convolution layer convolves again on the basis of the fifth convolution layer, the second splicing layer linearly splices the feature matrix X and the data output by the sixth convolution layer to obtain an attention map, and uses the normalization layer to decompose the attention map into a horizontal attention map and a vertical attention map; using the third splicing layer to weight the fourth feature, the horizontal attention map and the vertical attention map to obtain the fifth feature; Processing the target image with a human body feature extraction model to obtain target human body features, wherein the target human body features are used to reflect the human body to be reconstructed in the target image; The target human body features are processed by a human body three-dimensional reconstruction model to obtain an initial human body three-dimensional model; the human body three-dimensional reconstruction model is configured with an occupancy network and a marching cube modeling algorithm; the occupancy network is used to calculate the occupancy probability of sampling points in a three-dimensional space according to the target human body features, the three-dimensional space is a pre-set three-dimensional space to be reconstructed, the sampling points are the vertices of the grid units after the three-dimensional space is evenly divided into a plurality of grid units, the sampling points with the occupancy probability greater than the preset probability are retained, a new three-dimensional space is constructed based on the retained sampling points, the new three-dimensional space is divided into a plurality of cube units, and marching cube modeling is performed; Processing the target image using a texture acquisition model to obtain a two-dimensional human body texture; Using a texture mapping model to combine the initial human body 3D model and the human body texture to obtain a target human body 3D model, wherein the target human body 3D model is a human body 3D model with texture; The texture mapping model is configured with UV mapping, and each vertex on multiple cubic units constituting the initial human body three-dimensional model is mapped to a UV coordinate on a two-dimensional plane to obtain a flattened two-dimensional human body image. A triangular mesh is obtained through a marching cube modeling algorithm, and the UV coordinates of the vertices of each triangular mesh are mapped to the two-dimensional plane to form the UV coordinates corresponding to the initial human body three-dimensional model; then, the UV coordinates corresponding to the initial human body three-dimensional model and the UV coordinates corresponding to the human body texture are overlapped to form reconstructed UV coordinates, and the reconstructed UV coordinates contain the human body to be reconstructed and the human body texture. Finally, the reconstructed UV coordinates are mapped to the three-dimensional model by using inverse UV mapping to obtain the human body three-dimensional model with texture.
2. The method for three-dimensional reconstruction of a human body based on deep learning according to claim 1, characterized in that: The method for training the human body feature extraction model also includes: The training features are data predicted by the human feature extraction model and used to reflect the human body in the training image; Acquire actual features of each training image, where the actual features are data that truly reflect the human body in the training image; A loss function is constructed according to the training features and actual features of the plurality of training images; a human feature extraction model is trained according to the loss function until a value output by the loss function is less than a preset value, and the human feature extraction model is output.
3. The method for three-dimensional reconstruction of a human body based on deep learning according to claim 1, characterized in that: In the perception model, the attention layer is connected to the convolution layer, and the attention layer is used to extract key features in the training image, and the key features are used to reflect the human body in the training image; the activation function layer is arranged between the convolution layer and the attention layer, and the activation function layer is used to adjust the weights of the key features in the training image.
4. The method for three-dimensional reconstruction of a human body based on deep learning according to claim 2, characterized in that: The step of acquiring a human body image set comprises: Receive multiple training images; add Gaussian noise to each training image and crop invalid areas to obtain a final training image, wherein the invalid areas are areas that do not contain a human body; and use the final set of training images as the human body image set.
5. The method for three-dimensional reconstruction of a human body based on deep learning according to claim 1, characterized in that: The occupancy probability of the sampling point is calculated by the following formula: f θ (F 目 ,V)= τ ,in, θ is the parameter of the occupied network, F 目 is the target human feature, V is the sampling point, τ is the occupancy probability of the sampling point V.
6. The method for three-dimensional reconstruction of a human body based on deep learning according to claim 1, characterized in that: The marching cube modeling algorithm is used to build an initial human 3D model based on the retained sampling points, including: The cubic unit includes a volume data point or a volume data value, and both the volume data point and the volume data value are used to reflect the target human body characteristics; Extracting an isosurface, wherein the isosurface is composed of a plurality of cubic units of the same volume data point, or a plurality of cubic units of the same volume data value; Locking the intersection of the isosurface and each cubic unit constituting the isosurface; Connecting the intersection points using a triangulation method to form a triangular mesh; The triangular meshes are combined to obtain the initial human body three-dimensional model.
7. A human body three-dimensional reconstruction system based on deep learning, characterized in that: The system is used to implement the method described in any one of claims 1 to 6, including: A data acquisition module (1) is used to receive a target image, wherein the target image is an image containing a human body to be reconstructed; A first processing module (2) is used to process the target image using a human feature extraction model to obtain target human features, wherein the target human features are used to reflect the human body to be reconstructed in the target image; A second processing module (3) is used to process the target human body features using a human body three-dimensional reconstruction model to obtain an initial human body three-dimensional model; A third processing module (4) is used to process the target image using a texture acquisition model to obtain a two-dimensional human body texture; The data combination module (5) is used to combine the initial human body three-dimensional model and the human body texture by using a texture mapping model to obtain a target human body three-dimensional model, wherein the target human body three-dimensional model is a human body three-dimensional model with texture.
Citation Information
Patent Citations
Three-dimensional face model reconstruction method and device, storage medium and computer equipment
CN114037802A
Three-dimensional human body model reconstruction method and apparatus, electronic device, and storage medium
WO2022156533A1
Three-dimensional human body reconstruction method and apparatus, and device and storage medium
WO2022205760A1