Human body reconstruction network training, human body reconstruction, fitting method and related device
By training the human body to reconstruct the network and using the three-dimensional model parameters to build binary maps and heat maps, the problem of poor accuracy in building the human body’s three-dimensional model is solved, and the cost-effectiveness and business processing quality are improved.
Patent Information
- Application Number
- CN202111305818.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-05
AI Technical Summary
The prior art is difficult to improve the accuracy of building a three-dimensional model of the human body on the basis of maintaining costs, especially when users are wearing loose clothing.
By training the human body to reconstruct the network, the sampled human body three-dimensional model parameters are used to construct binary maps and heat maps, as training samples, self-supervised learning, get rid of the dependence on manual marking, and reduce costs.
It improves the accuracy of the three-dimensional model of the human body, enhances the quality of business processing in different scenarios, and takes into account the accuracy and generalization capabilities of the network.
Smart Images

Figure CN114119861B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of computer vision, and in particular to a human body reconstruction network training, human body reconstruction, fitting method and related devices. Background Art
[0002] Three-dimensional modeling of the human body can be used to estimate the body shape and has increasingly wide applications in animation, fashion, health and other scenarios.
[0003] Currently, users can use professional equipment such as 3D cameras to reconstruct more accurate 3D models of the human body, but professional equipment is expensive and difficult to promote.
[0004] Therefore, users mostly use non-professional devices such as mobile terminals to capture two-dimensional images and reconstruct a three-dimensional model of the human body based on the two-dimensional images. However, these methods basically reconstruct the consistency of the projection contour of the human body with the contour of the human body in the original image through fitting, while users often wear loose clothing such as skirts in daily life, which basically conceals the shape of the human body, resulting in poor accuracy in the construction of the three-dimensional model of the human body. As a result, the quality of business processing in these scenarios is relatively low. Summary of the invention
[0005] The embodiment of the present invention proposes a human body reconstruction network training, human body reconstruction, fitting method and related devices to solve the problem of how to improve the accuracy of constructing a three-dimensional model of the human body while maintaining the cost.
[0006] In a first aspect, an embodiment of the present invention provides a training method for a human body reconstruction network, comprising:
[0007] Sampling parameters in the three-dimensional human body model as original model parameters;
[0008] Constructing a binary image representing the two-dimensional contour of the human body when wearing clothes according to the original model parameters as the original clothing binary image;
[0009] Constructing a thermal map representing a two-dimensional skeleton of a human body according to the original model parameters as an original thermal map;
[0010] The human body reconstruction network is trained using the original clothing binary image and the original heat map as samples and the original model parameters as labels.
[0011] In a second aspect, an embodiment of the present invention further provides a human body reconstruction method, comprising:
[0012] Acquire two-dimensional target image data collected from the user;
[0013] Convert the target image data into a binary image representing the two-dimensional outline of a human body when dressed, as a target dressing binary image, and a thermal image representing the two-dimensional skeleton of a human body, as a target thermal image;
[0014] Input the target clothing binary image and the target heat map into the human body reconstruction network trained by the method described in the first aspect, and predict the parameters of the three-dimensional model of the human body as the target model parameters;
[0015] A three-dimensional human body model is constructed for the user using the target model parameters.
[0016] In a third aspect, an embodiment of the present invention further provides a fitting method, which is applied to a mobile terminal, and the method includes:
[0017] Determining clothes selected by a user, wherein the three-dimensional model of the clothes has clothes model parameters;
[0018] Collecting two-dimensional target image data from the user;
[0019] Convert the target image data into a binary image representing the two-dimensional outline of a human body when dressed, as a target dressing binary image, and a thermal image representing the two-dimensional skeleton of a human body, as a target thermal image;
[0020] Input the target clothing binary image and the target heat map into the human body reconstruction network trained by the method described in the first aspect, and predict the parameters of the three-dimensional model of the human body as the target model parameters;
[0021] A three-dimensional clothing model is constructed and displayed using the target model parameters and the clothing model parameters, wherein the clothing model is used to represent the user wearing the clothing.
[0022] In a fourth aspect, an embodiment of the present invention further provides a training device for a human body reconstruction network, comprising:
[0023] The original model parameter sampling module is used to sample the parameters in the three-dimensional human body model as the original model parameters;
[0024] An original clothing binary image construction module is used to construct a binary image representing the two-dimensional contour of the human body when wearing clothes according to the original model parameters as the original clothing binary image;
[0025] An original thermal map construction module, used to construct a thermal map representing a two-dimensional skeleton of a human body according to the original model parameters as an original thermal map;
[0026] The human body reconstruction network training module is used to train the human body reconstruction network using the original clothing binary image and the original thermal map as samples and the original model parameters as labels.
[0027] In a fifth aspect, an embodiment of the present invention further provides a human body reconstruction device, comprising:
[0028] A target image data acquisition module, used to acquire two-dimensional target image data collected from a user;
[0029] A target image data conversion module, used to convert the target image data into a binary image representing the two-dimensional contour of a human body when dressed, as a target dressing binary image, and a thermal image representing the two-dimensional skeleton of a human body, as a target thermal image;
[0030] A target model parameter prediction module, used for inputting the original clothing binary image and the target heat map into the human body reconstruction network trained by the method described in the first aspect, and predicting the parameters of the three-dimensional model of the human body as target model parameters;
[0031] The human body three-dimensional model construction module is used to construct a three-dimensional model of the human body of the user using the target model parameters.
[0032] In a sixth aspect, an embodiment of the present invention further provides a fitting device, comprising:
[0033] Clothes determination parameters, used to determine the clothes selected by the user, the three-dimensional model of the clothes having clothes model parameters;
[0034] A target image data acquisition module is used to collect two-dimensional target image data from a user;
[0035] A target image data conversion module, used to convert the target image data into a binary image representing the two-dimensional contour of a human body when dressed, as a target dressing binary image, and a thermal image representing the two-dimensional skeleton of a human body, as a target thermal image;
[0036] A target model parameter prediction module, used for inputting the target clothing binary image and the target heat map into the human body reconstruction network trained by the method described in the first aspect, and predicting the parameters of the three-dimensional model of the human body as target model parameters;
[0037] A dressing model construction module is used to construct a dressing model using the target model parameters and the clothing model parameters, and render the dressing model for display, wherein the dressing model is used to represent the user wearing the clothing.
[0038] In a seventh aspect, an embodiment of the present invention further provides a computer device, the computer device comprising:
[0039] one or more processors;
[0040] a memory for storing one or more programs,
[0041] When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the human body reconstruction network as described in the first aspect, the human body reconstruction method as described in the second aspect, or the fitting method as described in the third aspect.
[0042] In an eighth aspect, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the training method of the human body reconstruction network as described in the first aspect, the human body reconstruction method as described in the second aspect, or the fitting method as described in the third aspect is implemented.
[0043] In this embodiment, the parameters in the three-dimensional model of the human body are sampled as the original model parameters, and a binary image representing the two-dimensional contour of the human body when dressed is constructed according to the original model parameters as the original dressing binary image. A heat map representing the two-dimensional skeleton of the human body is constructed according to the original model parameters as the original heat map. The original dressing binary image and the original heat map are used as samples and the original model parameters are used as labels to train the human body reconstruction network. Since the parameters in the three-dimensional model of the human body as the real labels already exist, the source is simple and easy to obtain, the dependence on manual labeling can be eliminated, which greatly reduces the cost. By simulating real scenes through data enhancement technology, constructing the original dressing binary image of the two-dimensional contour of the human body when dressed and the original heat map projected to represent the two-dimensional skeleton of the human body, replacing the real two-dimensional image as the input, a large number of training sets can be constructed, and self-supervised learning of the human body reconstruction network is realized. At the same time, the accuracy and generalization ability of the human body reconstruction network are taken into account, that is, the accuracy of constructing the three-dimensional model of the human body is guaranteed, so that the quality of business processing is guaranteed in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A flowchart of a training method for a human body reconstruction network provided in Embodiment 1 of the present invention;
[0045] FIG. 2A to FIG. 2C This is an example diagram of simulating a human body putting on clothes provided in the first embodiment of the present invention;
[0046] Figure 3 A thermal map of a human skeleton provided in Embodiment 1 of the present invention;
[0047] Figure 4 is a flow chart of a human body reconstruction method provided by Embodiment 2 of the present invention;
[0048] Figure 5 is a flow chart of a fitting method provided by Embodiment 3 of the present invention;
[0049] Figure 6 A schematic diagram of the structure of a training device for a human body reconstruction network provided in a fourth embodiment of the present invention;
[0050] Figure 7 A schematic diagram of the structure of a human body reconstruction device provided in Embodiment 5 of the present invention;
[0051] Figure 8 A schematic structural diagram of a fitting device provided in Embodiment 6 of the present invention;
[0052] Fig. 9 A schematic diagram of the structure of a computer device provided in Embodiment 7 of the present invention. DETAILED DESCRIPTION
[0053] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0054] Embodiment 1
[0055] Figure 1 A flowchart of a training method for a human body reconstruction network provided in Example 1 of the present invention. This embodiment is applicable to the situation where a human body reconstruction network is trained by self-supervised learning. The so-called self-supervised learning may refer to the self-supervision of the output by mining the internal features of the input through an automated algorithm, thereby being able to use a large number of training sets that do not rely on manual labeling to train the human body reconstruction network.
[0056] The method can be performed by a training device for a human body reconstruction network, which can be implemented by software and / or hardware and can be configured in a computer device, such as a personal computer, a server, a workstation, etc., and specifically includes the following steps:
[0057] Step 101: Sample parameters in the three-dimensional human body model as original model parameters.
[0058] In this embodiment, open source three-dimensional human body data sets can be downloaded, and parameters in human body three-dimensional models can be randomly sampled from these open source three-dimensional human body data sets, for example, parameters of SMPL (Skinned Multi-Person Linear Model) model, parameters of STAR model, parameters of SMPL-X model, etc. These human body three-dimensional models are usually three-dimensional models of real human bodies. For ease of distinction, these parameters are recorded as original model parameters.
[0059] Taking the SMPL model as an example of a three-dimensional model of the human body, the SMPL model is a vertex-based three-dimensional model of the human body that can accurately represent different shapes and poses of the human body.
[0060] It divides body shapes into identity-dependent shape and non-rigid pose-dependent shape. The human body can be understood as the sum of a basic model and deformations based on the model. PCA is performed on the deformation to obtain low-dimensional parameters that describe the shape - shape parameters (shape); at the same time, the motion tree is used to represent the posture of the human body, that is, the rotation relationship between each joint point and the parent node of the motion tree. This relationship can be expressed as a three-dimensional vector. Finally, the local rotation vector of each joint point constitutes the posture parameter (pose) of the SPML model.
[0061] Step 102: construct a binary image representing the two-dimensional contour of the human body when wearing clothes according to the original model parameters, as the original clothing binary image.
[0062] Since the two-dimensional images of the human body in real scenes have already caused information loss due to the projection from three-dimensional to two-dimensional, it is difficult to obtain the parameters of the real three-dimensional human body model from daily two-dimensional images of the human body.
[0063] In this embodiment, the parameters of the real human body three-dimensional model are defined as labels, and corresponding two-dimensional data are generated to generate data pairs, forming a training set for training the human body reconstruction network.
[0064] Since the body shape estimation of the human body is basically based on the information of the human body contour, other redundant information such as clothing texture and background will not only interfere with the body shape estimation, but also require a larger capacity for the human body reconstruction network, which is not easy to be lightweight and efficiently calculated. Therefore, this embodiment can construct a binary image (two-dimensional image) representing the two-dimensional contour of the human body when dressed based on the original model parameters, define the binary image representing the two-dimensional contour of the human body when dressed as one of the training sets, and for ease of distinction, record the binary image representing the two-dimensional contour of the human body when dressed as the original dressing binary image.
[0065] Among them, a binary image means that the pixel value is two numerical values, such as 0 and 1, which is used to express the foreground area and the background area.
[0066] In this embodiment, if Figure 2A As shown, the foreground area in the binary image can refer to the area where the human body is located.
[0067] In one embodiment of the present invention, step 102 may include the following steps:
[0068] Step 1021: Use the original model parameters to construct a three-dimensional model of the human body.
[0069] In this embodiment, rendering can be performed according to the original model parameters to obtain a three-dimensional model of the human body, and information such as the surface mesh of the human body can be obtained. Generally, the three-dimensional model of the human body is a three-dimensional model without clothes.
[0070] Step 1022: Sample the camera parameters as original camera parameters.
[0071] The projection of the three-dimensional model of the human body into two dimensions depends on the parameters of the camera (also known as the camera). That is, the camera can project objects in the three-dimensional world (the three-dimensional model of the human body) into a two-dimensional image by taking images. The imaging model is to establish a projection mapping relationship from three-dimensional space to two-dimensional space.
[0072] In this regard, the present embodiment can randomly sample the parameters of the camera, and for the sake of distinction, the parameters of the camera are recorded as the original camera parameters. For the same 3D model, different original camera parameters can be randomly sampled to increase the number of projections to 2D images, thereby enriching the training set.
[0073] For the parameters of the human body three-dimensional model such as the SMPL model, its posture already includes the global rotation parameters. Therefore, for the camera, displacement and size can be defined to characterize the mapping relationship from the three-dimensional model to the two-dimensional image.
[0074] In a specific implementation, under the condition that the original model parameters include rotation parameters, the projection of the three-dimensional model to the two-dimensional model can be determined as a weak projection. In the weak projection, the displacement and scale of the camera are set as the original camera parameters.
[0075] Furthermore, pinhole imaging is when an object is larger when it is closer and smaller when it is farther away, and parallel lines intersect. The projection process corresponding to pinhole imaging is called perspective projection. Weak perspective projection is a further simplification of pinhole imaging. When projecting an object, the distance between each point and the pinhole is replaced by the average distance of the object, that is, there is no such thing as larger when it is closer and smaller when it is farther away, because they are all at the same distance. In this way, the size of the image is determined by the ratio between the distance d between the imaging plane and the pinhole, and the distance D between the object and the pinhole.
[0076] This embodiment uses weak perspective projection to model the process of projecting a three-dimensional model to a two-dimensional image, which can not only more accurately characterize the actual camera imaging, but also relatively reduce the original camera parameters, that is, the displacement in two directions and a scale, thereby reducing the output dimension during the prediction of the human body reconstruction network, reducing the amount of calculation, and reducing the resource usage, thus providing conditions for the deployment of the human body reconstruction network on devices with limited resources such as mobile terminals.
[0077] Step 1023: Project the surface mesh of the human body in the three-dimensional model into a binary image representing the two-dimensional contour of the human body according to the original camera parameters, as the original human body binary image.
[0078] Generally speaking, in order to obtain more realistic three-dimensional human body shape data, we can rely on two-dimensional body shape data from at least two different perspectives (such as the front perspective and the side perspective). Therefore, we can project the surface mesh of the human body in the three-dimensional model according to the original camera parameters at least in two perspectives (such as the front perspective and the side perspective) to obtain a two-dimensional image, which represents the original human body binary image of the two-dimensional contour of the human body.
[0079] Taking the SMPL model as an example, the value of the 0th dimension posture parameter (controlling the global rotation of the three-dimensional model) in the SMPL model can be controlled within the range corresponding to the front and side viewing angles.
[0080] Of course, projecting the surface mesh of the human body in the three-dimensional model according to the original camera parameters in a single perspective can also be applied, but the accuracy of the human body reconstruction network will be degraded.
[0081] Step 1024: Obtain a binary image representing the clothes as the clothes binary image.
[0082] In this embodiment, an open-source clothing segmentation dataset may be downloaded, and binary images representing clothing may be obtained by random sampling from the clothing segmentation dataset. For ease of distinction, these images are recorded as clothing binary images.
[0083] like Figure 2B As shown, the foreground area in the binary image can refer to the area where the clothes are located.
[0084] Step 1025: superimpose the clothing binary image on the original human body binary image to generate a binary image representing the two-dimensional contour of the human body when wearing the clothing, as the original clothing binary image.
[0085] Since the human body in the original human body binary image is generally not wearing clothes (dressed), there is a big difference between the distribution of the binary image of the human body contour obtained by the real segmentation and the distribution of the binary image. Therefore, this embodiment can use data enhancement technology to simulate the shape (binary image) of the two-dimensional contour of the human body when wearing clothes (dressed) based on the original human body binary image, which is recorded as the original clothing binary image.
[0086] Among them, data enhancement technology uses various methods to change existing data in order to cover the distribution in real scenarios as much as possible, which is a method to improve the generalization ability of the model.
[0087] In one example, since clothes are worn on the upper body and / or the lower body, for example, shirts and sweatshirts are worn on the upper body, pants and skirts are worn on the lower body, and dresses and suspender skirts are worn on the whole body (upper body and lower body), the clothing binary image can be pre-labeled with a label that represents the clothing mark suitable for the upper body and / or lower body.
[0088] In this example, anchor points can be selected from the parameters of the human body three-dimensional model, thereby selecting anchor points in the original human body binary image, and the anchor points are used to roughly locate the upper body and / or lower body of the human body two-dimensional contour.
[0089] Thus, according to the labels of the clothing binary images, a clothing binary image suitable for being superimposed on the upper body and / or the lower body is selected as the superimposed binary image.
[0090] The superimposed binary image is subjected to affine transformation until the distance between the outline of the superimposed binary image and the anchor point meets the preset superposition condition, thereby obtaining a binary image representing the two-dimensional outline of the human body when dressed as the original dressing binary image.
[0091] Among them, affine transformation (also known as affine mapping) refers to a process in geometry where a vector space is transformed into another vector space by a linear transformation followed by a translation.
[0092] The superposition condition indicates that the superposition of clothes and human body is close to or meets the real condition. When the distance between the outline of the superimposed binary image and the anchor point meets the preset superposition condition, the effect of simulating human body wearing clothes is met.
[0093] Exemplarily, the superposition condition may include at least one of the following:
[0094] The sum of all distances is the smallest, the average of all distances is the smallest, the sum of all distances is less than a certain threshold, and the average of all distances is another threshold.
[0095] Since data augmentation is used very many times in the training stage, there is no need to demand the accuracy of matching the human body contour with the clothing contour each time it is superimposed. As long as there is a certain probability of simulating a real human body wearing clothes, it can be very effective for training. This eliminates the distribution difference between the binary image projected by the three-dimensional human body model and the real human body contour, so that the human body reconstruction network can adapt to the situation of wearing loose clothing such as skirts, thereby increasing the adaptability and generalization ability of the human body reconstruction network to real scenes. The effect is more robust than other schemes that only supervise based on contour consistency.
[0096] Furthermore, image enhancement processing is performed on the original clothing binary image. The image enhancement processing belongs to the basic data enhancement technology, wherein the image enhancement processing includes at least one of the following:
[0097] Randomly block the original clothing binary image, randomly add noise to the original clothing binary image (to simulate inaccurate segmentation), and perform Gaussian blur processing on the original clothing binary image (to remove edge jaggedness).
[0098] Of course, the above-mentioned method of generating the original clothing binary image is only used as an example. When implementing the embodiment of the present invention, other methods of generating the original clothing binary image can be set according to actual conditions, for example, distinguishing the head, torso, limbs and other parts of the human body, superimposing the clothing binary image on the original human body binary image, or using an adversarial generation network to generate the original clothing binary image, etc., and the embodiment of the present invention does not limit this. In addition, in addition to the above-mentioned judgment and processing method, those skilled in the art can also adopt other methods of generating the original clothing binary image according to actual needs, and the embodiment of the present invention does not limit this.
[0099] Step 103: construct a thermal map representing the two-dimensional skeleton of the human body according to the original model parameters as the original thermal map.
[0100] This embodiment can construct a thermal map (two-dimensional image) representing the two-dimensional skeleton of the human body based on the original model parameters, define the thermal map of the two-dimensional skeleton of the human body as one of the training sets, and for easy distinction, record the thermal map representing the two-dimensional skeleton of the human body as the original thermal map.
[0101] The heat map uses the original image data width and height multi-channel matrix to represent the anchor point position coordinates, and converts the coordinate information into a feature map structure that can be input into the neural network.
[0102] In a specific implementation, the original model parameters can be used to construct a three-dimensional model of the human body, and information such as the human body's bone points can be obtained.
[0103] The parameters of the sampled camera are used as the original camera parameters.
[0104] Further, under the condition that the original model parameters include rotation parameters, it is determined that the projection of the three-dimensional model to the two-dimensional is a weak projection, so that in the weak projection, the displacement and scale of the camera are set as the original camera parameters.
[0105] Orthogonal projection is performed on the skeleton points (key points representing the skeleton) of the human body in the three-dimensional model according to the original camera parameters in at least two viewing angles (such as the frontal view and the side vision, etc.), and a thermal map representing the two-dimensional skeleton of the human body is obtained as the target thermal map.
[0106] Among them, orthogonal projection is a further simplification of weak perspective projection, making the ratio in weak perspective projection 1, and its effect is equivalent to projecting parallel light onto the imaging plane.
[0107] After orthogonal projection, the skeleton points of the human body in the 3D model are aligned with the skeleton points of the human body in 2D.
[0108] In addition, the number of skeleton points can be adjusted according to business needs, including 14 skeleton points, 18 skeleton points, and so on.
[0109] For example, Figure 3 As shown in the figure, in the heat map, there are a total of 18 bone points (points numbered 0-17), the pixel value of the bone point is 1, and the pixel value of other positions is 0. For ease of understanding, the bone points are connected to represent the two-dimensional skeleton of the human body.
[0110] Of course, constructing a heat map representing the two-dimensional skeleton of the human body according to the original camera parameters in a single perspective can also be applied, but the accuracy of the human body reconstruction network will be degraded.
[0111] Step 104: Using the original clothing binary image and the original heat map as samples and the original model parameters as labels, train the human body reconstruction network.
[0112] In this embodiment, the original clothing binary image and the original heat map are used as training samples, and the original model parameters are marked as tags. Under the supervision of the tags (true values), the human body reconstruction network is trained so that the human body reconstruction network can at least output the parameters of the human body three-dimensional model.
[0113] In order to enable the human body reconstruction network to at least output the parameters of the human body three-dimensional model, the human body reconstruction network is a regression model.
[0114] If deployed on devices with abundant resources such as servers, the human body reconstruction network can be a larger model with stronger generalization performance, such as ResNet, VGG, etc.
[0115] If deployed on devices with limited resources such as mobile terminals, the human body reconstruction network can be a lightweight model based on separable convolution and residual networks, such as MobileNet.
[0116] Of course, the structure of the human body reconstruction network is not limited to artificially designed neural networks, but can also be a neural network optimized by a model quantization method, a neural network searched by a NAS (neural network structure search) method, and so on. This embodiment does not impose any restrictions on this.
[0117] Furthermore, the process of training the human body reconstruction network may be to retrain the human body reconstruction network, or to perform fine-tuning based on a previously trained human body reconstruction network, which is not limited in this embodiment.
[0118] In one embodiment of the present invention, step 104 may include the following steps:
[0119] Step 1041: Query original camera parameters.
[0120] In this embodiment, the original camera parameters used in constructing a target binary image representing the two-dimensional outline of a human body when dressed and an original thermal map representing the two-dimensional skeleton of a human body can be queried. The original camera parameters are used to project the three-dimensional model to two dimensions, such as the original binary image and the original thermal map.
[0121] Step 1042: Input the original clothing binary image and the original heat map into the human body reconstruction network to predict the parameters of the human body three-dimensional model and the parameters of the camera as the predicted model parameters and the predicted camera parameters.
[0122] The original clothing binary image and the original heat map under at least two perspectives are input into the human body reconstruction network. The human body reconstruction network processes the original clothing binary image and the original heat map under at least two perspectives according to its own network structure, and outputs the parameters of the human body three-dimensional model and the parameters of the camera. For easy distinction, the parameters of the human body three-dimensional model are recorded as predicted model parameters, and the parameters of the camera are recorded as predicted camera parameters.
[0123] Step 1043: Calculate the difference between the original model parameters, the original camera parameters and the predicted model parameters, the predicted camera parameters as the first loss value.
[0124] In this embodiment, the original model parameters, the original camera parameters are compared with the predicted model parameters, the predicted camera parameters, and the difference between the original model parameters, the original camera parameters and the predicted model parameters, the predicted camera parameters is calculated, and recorded as the first loss value LOSS 1 , the first loss value LOSS 1 Represents the loss of parameters in regression.
[0125] In one example, on the one hand, the norm distance L2 between each original model parameter and each predicted model parameter can be calculated as the first difference value.
[0126] On the other hand, a norm distance L2 between each original camera parameter and each predicted camera parameter may be calculated as a second difference value.
[0127] Calculate the average of the first difference value and the second difference value as the first loss value LOSS 1 .
[0128] Among them, the norm distance L2 is the Euclidean distance, which represents the distance between two points in space.
[0129] In addition, in addition to the above-mentioned method for calculating the first loss value, those skilled in the art may also adopt other methods for calculating the first loss value according to actual needs, for example, summing the first difference value and the second difference value, etc., and the embodiments of the present invention are not limited to this.
[0130] Step 1044: Update the human body reconstruction network according to the first loss value.
[0131] Back-propagation is performed on the human body reconstruction network, and the weights in the human body reconstruction network are updated based on the first loss value.
[0132] In some cases, the human body reconstruction network can be back-propagated, and the first loss value can be substituted into algorithms such as SGD (stochastic gradient descent) and Adam (Adaptive momentum) to calculate the update amplitude of the weights in the human body reconstruction network, thereby updating the weights in the human body reconstruction network according to the update amplitude.
[0133] In one embodiment of the present invention, considering that the relationship between the original model parameters, the original camera parameters and the predicted model parameters, the predicted camera parameters is relatively remote, the human body reconstruction network is relatively difficult to fit. Therefore, two additional loss values are added to supervise the training of the human body reconstruction network respectively. Since the rendering part is differentiable and can be back-propagated normally, the process of human body reconstruction and projection can be simulated by differentiable rendering, wherein differentiable rendering refers to a rendering method that calculates the gradient of each step, which can complete back-propagation as a module during the training process.
[0134] In this embodiment, step 1044 may further include the following steps:
[0135] Step 10441: construct a binary image representing the two-dimensional contour of the human body when wearing clothes according to the prediction model parameters as the predicted human body binary image.
[0136] In this embodiment, a binary image representing the two-dimensional contour of the human body when dressed can be constructed according to the prediction model parameters in reference to the method of generating the original clothing binary image, and recorded as the predicted human body binary image, thereby simulating the projection process.
[0137] In a specific implementation, the predicted model parameters can be used to construct a three-dimensional model of the human body as the predicted three-dimensional model; the parameters of the sampling camera are used as candidate camera parameters; (at least in two viewing angles) the surface mesh of the human body in the predicted three-dimensional model is projected into a binary image representing the two-dimensional contour of the human body according to the candidate camera parameters, which is used as a candidate human body binary image; a binary image representing clothes is obtained, which is used as a clothes binary image; the clothes binary image is superimposed on the candidate human body binary image to generate a binary image representing the two-dimensional contour of the human body when wearing clothes, which is used as the predicted human body binary image.
[0138] Among them, when sampling candidate camera parameters, it is possible to determine that the projection of the predicted three-dimensional model to two dimensions is a weak projection under the condition that the predicted model parameters include rotation parameters; in the weak projection, the displacement and scale of the camera are set as candidate camera parameters.
[0139] When generating a predicted human body binary map, anchor points can be selected in the candidate human body binary map, and the anchor points are used to locate the upper body and / or lower body of the two-dimensional contour of the human body; a clothing binary map suitable for superimposing on the upper body and / or lower body is selected as the superimposed binary map; an affine transformation is performed on the superimposed binary map until the distance between the contour of the superimposed binary map and the anchor points meets the preset superposition condition, so as to obtain a binary map representing the two-dimensional contour of the human body when wearing the clothes, which is used as the predicted human body binary map.
[0140] Further, performing image enhancement processing on the predicted human body binary image;
[0141] The image enhancement processing includes at least one of the following:
[0142] Randomly occlude the predicted human binary image, randomly add noise to the predicted human binary image, and perform Gaussian blur processing on the predicted human binary image.
[0143] Step 10442: construct a thermal map representing the two-dimensional skeleton of the human body according to the prediction model parameters as a prediction thermal map.
[0144] In this embodiment, a method of generating a target thermogram can be referred to, and a thermogram representing the two-dimensional skeleton of the human body can be constructed according to the prediction model parameters as a prediction thermogram, thereby simulating the projection process.
[0145] In a specific implementation, the predicted model parameters can be used to construct a three-dimensional model of the human body as a predicted three-dimensional model; the parameters of the sampling camera can be used as predicted camera parameters; (at least in two perspectives) the bone points of the human body in the predicted three-dimensional model are orthogonally projected according to the predicted camera parameters to obtain a heat map representing the two-dimensional skeleton of the human body as a predicted heat map.
[0146] Among them, when sampling candidate camera parameters, it is possible to determine that the projection of the predicted three-dimensional model to two dimensions is a weak projection under the condition that the predicted model parameters include rotation parameters; in the weak projection, the displacement and scale of the camera are set as candidate camera parameters.
[0147] In this embodiment, since the application of step 10441 and step 10442 is basically similar to that of step 101 and step 102, the description is relatively simple. For relevant details, please refer to the partial description of step 102 and step 103. The embodiment of the present invention is not described in detail here.
[0148] Step 10443: Calculate the difference between the original clothing binary image and the predicted human body binary image as the second loss value.
[0149] In this embodiment, the original clothing binary image is compared with the predicted human body binary image, and the difference between the original clothing binary image and the predicted human body binary image is calculated as the second loss value LOSS2 , the second loss value LOSS 2 Represents the loss of human contour in regression.
[0150] In one example, the norm distance L2 between each pixel in the original clothing binary image and each pixel in the predicted human body binary image is calculated as the third difference value; the average value of the third difference value is calculated as the second loss value LOSS 2 .
[0151] In addition, in addition to the above method for calculating the second loss value, those skilled in the art may also adopt other methods for calculating the second loss value according to actual needs, such as summing the third difference values, etc., and the embodiment of the present invention is not limited to this.
[0152] Step 10444: Calculate the difference between the original heat map and the predicted heat map as the third loss value.
[0153] In this embodiment, the original heat map is compared with the predicted heat map, and the difference between the original heat map and the predicted heat map is calculated as the third loss value LOSS. 3 , the third loss value LOSS 3 Represents the loss of skeleton points in regression.
[0154] In one example, the norm distance L2 between each bone point in the original heat map and each bone point in the predicted heat map is calculated as the fourth difference value, and the average value of the fourth difference value is calculated as the third loss value LOSS 3 .
[0155] In addition, in addition to the above method for calculating the third loss value, those skilled in the art may also adopt other methods for calculating the third loss value according to actual needs, for example, summing the fourth difference values, etc., and the embodiment of the present invention is not limited thereto.
[0156] Step 10445: merge the first loss value, the second loss value and the third loss value to obtain a total loss value.
[0157] In this embodiment, the first loss value, the second loss value and the third loss value may be integrated with reference to each other to obtain a total loss value.
[0158] By adding the loss of human contour in regression (the second loss value) and the loss of skeleton point in regression (the third loss value), which are the two types of inputs that supervise the training of the human reconstruction network, the input and output are more closely linked, the entire training optimization process is more interpretable, and the training difficulty of the human reconstruction network is reduced.
[0159] In one fusion method, the first loss value LOSS1 , the second loss value LOSS 2 and the third loss value LOSS 3 There is a linear fusion between them. Specifically, the first loss value LOSS 1 Multiply the first coefficient α by the preset value to obtain the first weighted loss value, and convert the second loss value LOSS 2 Multiply the preset second coefficient β to obtain the second weighted loss value, and convert the third loss value LOSS 3 Multiply by the preset third coefficient γ to obtain the third weighted adjustment loss value, and calculate the sum of the first weighted adjustment loss value, the second weighted adjustment loss value and the third weighted adjustment loss value as the total loss value LOSS 总 , which is expressed as follows:
[0160] LOSS 总 =α×LOSS 1 +β×LOSS 2 +γ×LOSS 3
[0161] Among them, the first loss value LOSS 1 Plays a major role, the second loss value LOSS 2 and the third loss value LOSS 3 It plays an auxiliary role, so the first coefficient α is greater than the third coefficient γ, and the third coefficient γ is greater than the second coefficient β, for example, α:β:γ=20:1:5.
[0162] Step 10446: adjust the weights in the human body reconstruction network according to the total loss value.
[0163] In this embodiment, the human body reconstruction network can be back-propagated, and the total loss value can be substituted into algorithms such as SGD and Adam to calculate the update amplitude of the weights in the human body reconstruction network, thereby updating the weights in the human body reconstruction network according to the update amplitude.
[0164] Step 1045, determine whether the number of steps of the current iteration reaches a preset threshold; if so, execute step 1046, if not, return to execute step 1042.
[0165] Step 1046: Determine that the human body reconstruction network training is completed.
[0166] In this embodiment, a threshold value may be set in advance for the iterative deployment as a stop condition. In each round of iterative training, the number of steps of the current iteration is counted to determine whether the threshold value is reached.
[0167] If the threshold is reached, the training of the human body reconstruction network can be considered complete, at which point the weights in the human body reconstruction network are recorded.
[0168] If the threshold is not reached, the next round of iterative training can be started, and the iterative training will be repeated until the human body reconstruction network training is completed.
[0169] In this embodiment, the parameters in the three-dimensional model of the human body are sampled as the original model parameters, and a binary image representing the two-dimensional contour of the human body when dressed is constructed according to the original model parameters as the original dressing binary image. A heat map representing the two-dimensional skeleton of the human body is constructed according to the original model parameters as the original heat map. The original dressing binary image and the original heat map are used as samples and the original model parameters are used as labels to train the human body reconstruction network. Since the parameters in the three-dimensional model of the human body as the real labels already exist, the source is simple and easy to obtain, the dependence on manual labeling can be eliminated, which greatly reduces the cost. By simulating real scenes through data enhancement technology, constructing the original dressing binary image of the two-dimensional contour of the human body when dressed and the original heat map projected to represent the two-dimensional skeleton of the human body, replacing the real two-dimensional image as the input, a large number of training sets can be constructed, and self-supervised learning of the human body reconstruction network is realized. At the same time, the accuracy and generalization ability of the human body reconstruction network are taken into account, that is, the accuracy of constructing the three-dimensional model of the human body is guaranteed, so that the quality of business processing is guaranteed in different scenarios.
[0170] Embodiment 2
[0171] Figure 4 This is a flow chart of a human body reconstruction method provided in Embodiment 2 of the present invention. This embodiment is applicable to the case of reconstructing a human body through a human body reconstruction network. The method can be executed by a human body reconstruction device, which can be implemented by software and / or hardware and can be configured in a computer device, such as a personal computer, a server, a workstation, a mobile terminal (such as a mobile phone, a tablet computer, a smart wearable device (such as a watch, a bracelet, etc.), etc., and specifically includes the following steps:
[0172] Step 401: Acquire two-dimensional target image data collected from the user.
[0173] In a specific implementation, if the computer device is a client-side device (such as a mobile terminal), its operating system may include Android, IOS, Windows, etc.
[0174] These operating systems support running applications that can perform image processing, such as shopping applications, live broadcast applications, image editing applications, camera applications, instant messaging tools, gallery applications, and so on.
[0175] Applications such as image editing applications, instant messaging tools, and gallery applications may have a UI (User Interface) that may provide imported controls. Users may operate the imported controls through touch or a mouse or other peripheral device to select locally stored two-dimensional target image data collected by the user (represented by a thumbnail or path), or may select network stored two-dimensional target image data collected by the user (represented by a URL (Uniform Resource Locators)).
[0176] Applications such as live broadcast applications, image editing applications, camera applications, instant messaging tools, etc., their UIs may provide controls for taking photos and recording videos. Users may operate the controls through touch or external devices such as a mouse to notify the application to call the camera to collect two-dimensional target image data for the user.
[0177] If the computer device is a device on the server side (such as a server), it can receive target image data uploaded by the client to capture two-dimensional data of the user.
[0178] The so-called two-dimensional target image data collected from the user may mean that the content of the two-dimensional target image data includes the user, and in particular, the body of the user.
[0179] Step 402: convert the target image data into a binary image representing the two-dimensional contour of the human body when dressed, as the target dressing binary image, and a thermal map representing the two-dimensional skeleton of the human body, as the target thermal map.
[0180] On the one hand, the target image data can be binarized to obtain a binary image representing the two-dimensional contour of the human body when dressed, which is recorded as the target clothing binary image.
[0181] On the other hand, the key points of the human skeleton (i.e., bone points) can be detected in the target image data, the posture of the human body can be detected, and a thermal map representing the two-dimensional skeleton of the human body can be obtained, which is recorded as the target thermal map.
[0182] Step 403: Input the target clothing binary image and the target heat map into the human body reconstruction network to predict the parameters of the three-dimensional model of the human body as the target model parameters.
[0183] In this embodiment, a human body reconstruction network may be pre-trained, and the human body reconstruction network is used to generate parameters of a three-dimensional model of a human body.
[0184] In the specific implementation, the training method of the human body reconstruction network is as follows:
[0185] The parameters in the human body three-dimensional model are sampled as the original model parameters. A binary image representing the two-dimensional contour of the human body when dressed is constructed according to the original model parameters as the original dressing binary image. A heat map representing the two-dimensional skeleton of the human body is constructed according to the original model parameters as the original heat map. The human body reconstruction network is trained with the original dressing binary image and the original heat map as samples and the original model parameters as labels.
[0186] In this embodiment, since the training method of the human body reconstruction network is basically similar to the application of the first embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the first embodiment. The embodiment of the present invention is not described in detail here.
[0187] The target clothing binary image and the target heat map are input into the human body reconstruction network. The human body reconstruction network processes the target clothing binary image and the target heat map according to its own network structure, outputs the parameters of the three-dimensional model of the human body, recorded as the target model parameters, and outputs the parameters of the camera.
[0188] Since the camera parameters are not used in the process of constructing the three-dimensional model of the human body in this embodiment, but are used in other business processing, the camera parameters can be ignored in this embodiment.
[0189] Step 404: Use the target model parameters to construct a three-dimensional human body model for the user.
[0190] In this embodiment, a three-dimensional model can be constructed and rendered according to the target model parameters to obtain a three-dimensional human body model of the current user.
[0191] Furthermore, considering that some target model parameters may omit certain data, such as the parameters of the SMPL model omit the user's height, the user can additionally input these data (such as the user's actual height) so as to use the target model parameters and these data to construct a three-dimensional model of the human body for the user, thereby improving the authenticity of the three-dimensional model of the human body.
[0192] In this embodiment, two-dimensional target image data collected from the user is obtained, and the target image data is converted into a binary image representing the two-dimensional contour of the human body when dressed as the target dressing binary image, and a heat map representing the two-dimensional skeleton of the human body as the target heat map. The target dressing binary image and the target heat map are input into the human body reconstruction network, and the parameters of the three-dimensional model of the human body are predicted as the target model parameters. The target model parameters are used to construct a three-dimensional model of the human body for the user. Since the parameters in the three-dimensional model of the human body as the real label already exist, the source is simple and easy to obtain, and the dependence on manual labeling can be eliminated, which greatly reduces the cost. By simulating real scenes through data enhancement technology, constructing the original dressing binary image of the two-dimensional contour of the human body when dressed and the original heat map projected to represent the two-dimensional skeleton of the human body, replacing the real two-dimensional image as the input, a large number of training sets can be constructed, and self-supervised learning of the human body reconstruction network is realized. At the same time, the accuracy and generalization ability of the human body reconstruction network are taken into account, that is, the accuracy of constructing the three-dimensional model of the human body is guaranteed, so that the quality of business processing is guaranteed in different scenarios.
[0193] Embodiment 3
[0194] Figure 5 This is a flowchart of a fitting method provided in Embodiment 3 of the present invention. This embodiment can be applied to reconstructing a human body through a human body reconstruction network for virtual fitting, and can be configured in a computer device, especially in a mobile terminal, such as a mobile phone, a tablet computer, a smart wearable device (such as a watch, a bracelet, etc.), etc., and specifically includes the following steps:
[0195] Step 501: Determine the clothes selected by the user.
[0196] In a specific implementation, the operating system of the mobile terminal may include Android, IOS, Windows, etc., and these operating systems support running various applications, such as shopping applications, live broadcast applications, short video applications, and the like.
[0197] In one case, these applications are configured with a page browsing component (such as WebView), which can access the server and load a page, in which a variety of clothes are provided for users to browse and select, and each piece of clothing represents a style, which is provided with multiple sizes.
[0198] The user can use touch or other methods to trigger a selection operation for a certain piece of clothing (including size), and at this time, the server is notified that the user has selected the piece of clothing (including size).
[0199] In another case, these applications download a database that records a variety of clothes, which are displayed on the UI for users to browse and select. Each piece of clothing represents a style that is provided with a variety of sizes.
[0200] The user can use touch or other methods to trigger a selection operation for a certain piece of clothing (including size), and at this time, it can be determined that the user has selected the piece of clothing (including size).
[0201] Step 502: Collect two-dimensional target image data from the user.
[0202] Import controls may be provided in the UI, and users may operate the imported controls through touch or external devices such as a mouse, and select locally stored two-dimensional target image data collected for the user (represented by a thumbnail or path), or select network stored two-dimensional target image data collected for the user (represented by a URL (Uniform Resource Locators)), and may also call a camera to collect two-dimensional target image data for the user.
[0203] Step 503: convert the target image data into a binary image representing the two-dimensional contour of the human body when dressed, as the target dressing binary image, and a thermal map representing the two-dimensional skeleton of the human body, as the target thermal map.
[0204] On the one hand, the target image data can be binarized to obtain a binary image representing the two-dimensional contour of the human body when dressed, which is recorded as the target clothing binary image.
[0205] On the other hand, the key points of the human skeleton (i.e., bone points) can be detected in the target image data, the posture of the human body can be detected, and a thermal map representing the two-dimensional skeleton of the human body can be obtained, which is recorded as the target thermal map.
[0206] Step 504: Input the target clothing binary image and the target heat map into the human body reconstruction network to predict the parameters of the three-dimensional model of the human body as the target model parameters.
[0207] In this embodiment, a human body reconstruction network may be pre-trained, and the human body reconstruction network is used to generate parameters of a three-dimensional model of a human body.
[0208] In the specific implementation, the training method of the human body reconstruction network is as follows:
[0209] The parameters in the human body three-dimensional model are sampled as the original model parameters. A binary image representing the two-dimensional contour of the human body when dressed is constructed according to the original model parameters as the original dressing binary image. A heat map representing the two-dimensional skeleton of the human body is constructed according to the original model parameters as the original heat map. The human body reconstruction network is trained with the original dressing binary image and the original heat map as samples and the original model parameters as labels.
[0210] In this embodiment, since the training method of the human body reconstruction network is basically similar to the application of the first embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the first embodiment. The embodiment of the present invention is not described in detail here.
[0211] The target clothing binary image and the target heat map are input into the human body reconstruction network. The human body reconstruction network processes the target clothing binary image and the target heat map according to its own network structure, outputs the parameters of the three-dimensional model of the human body, recorded as the target model parameters, and outputs the parameters of the camera.
[0212] The human body reconstruction network is deployed on the mobile terminal, which can avoid uploading its own privacy data (such as target image data, target clothing binary map, target heat map, etc.) to the server, avoiding the problem of user privacy leakage.
[0213] Since the camera parameters are not used in the process of constructing the three-dimensional model of the human body in this embodiment, but are used in other business processing, the camera parameters can be ignored in this embodiment.
[0214] Step 505: Use the target model parameters and the clothing model parameters to construct a clothing model and display it.
[0215] In this embodiment, the three-dimensional model of the clothes has clothes model parameters, and the three-dimensional model is constructed and rendered by combining the target model parameters with the clothes model parameters, which is recorded as a dressing model. The dressing model is used to represent the user putting on clothes.
[0216] Taking into account the large amount of computation required to construct and render a three-dimensional model, the mobile terminal can upload the target model parameters to the server, which will use the target model parameters and the clothing model parameters to construct and render the clothing model, and send the clothing model to the mobile terminal. The mobile terminal will display the clothing model on the UI for the user to browse.
[0217] Taking user privacy into consideration, the target model parameters and clothing model parameters may be used together to construct and render a dressing model locally on the mobile terminal, and the dressing model may be displayed on the UI for the user to browse.
[0218] Furthermore, considering that some target model parameters may omit certain data, such as the SMPL model parameters omit the user's height, the user can additionally input these data (such as the user's actual height) so as to use the target model parameters, clothing model parameters and these data to construct a three-dimensional clothing model for the user, thereby improving the authenticity of the clothing model.
[0219] In addition, rules for the degree of matching between the human body and clothes can be set in advance. For example, if the distance between the outline of the clothes and the outline of the human body is less than a first threshold or greater than a second threshold (the second threshold is greater than the first threshold), it indicates that there is a mismatch between the human body and the clothes; or, if the length of the clothes suitable for the upper body is less than the length of the upper body of the human body, it indicates that there is a mismatch between the human body and the clothes, and so on.
[0220] The dressing model is evaluated according to the rule to determine the matching degree between the current user's body and the clothes. If the current user's body and the clothes do not match, it means that the size of the current clothes is inappropriate. At this time, the size that matches the current user's body can be predicted and recommended to the user.
[0221] Although deep learning technology is developing rapidly, the cost of obtaining real labels of three-dimensional human bodies is relatively high at this stage, and open source related data is relatively scarce. Therefore, it is difficult to have data covering real scenes to train accurate human body reconstruction networks. In this context, human body reconstruction methods for two-dimensional images usually cannot take into account both accuracy and generalization in body shape.
[0222] In this embodiment, the clothes selected by the user are determined, and the three-dimensional model of the clothes has clothes model parameters. Two-dimensional target image data of the user is collected, and the target image data is converted into a binary image representing the two-dimensional contour of the human body when wearing clothes as the target clothing binary image, and a heat map representing the two-dimensional skeleton of the human body as the target heat map. The target clothing binary image and the target heat map are input into the human body reconstruction network, and the parameters of the three-dimensional model of the human body are predicted as the target model parameters. The target model parameters and the clothes model parameters are used to construct a three-dimensional dressing model and display it. The dressing model is used to represent the user putting on clothes. Since it is used as a real label The parameters in the human body three-dimensional model already exist, with a simple source and easy to obtain, which can get rid of the dependence on manual labeling and greatly reduce the cost. By simulating real scenes through data enhancement technology, constructing the original clothing binary image of the two-dimensional outline of the human body when dressing and the original thermal map of the projection representing the two-dimensional skeleton of the human body, instead of the real two-dimensional image as input, a large number of training sets can be constructed, and self-supervised learning of the human body reconstruction network is realized. At the same time, the accuracy and generalization ability of the human body reconstruction network are taken into account, that is, the accuracy of constructing the human body three-dimensional model is guaranteed, thereby ensuring the accuracy of constructing the dressing model, so that the user can try on clothes accurately.
[0223] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0224] Embodiment 4
[0225] Figure 6 A structural block diagram of a training device for a human body reconstruction network provided in Embodiment 4 of the present invention may specifically include the following modules:
[0226] The original model parameter sampling module 601 is used to sample parameters in the three-dimensional human body model as original model parameters;
[0227] The original clothing binary image construction module 602 is used to construct a binary image representing the two-dimensional contour of the human body when wearing clothes according to the original model parameters as the original clothing binary image;
[0228] An original thermal map construction module 603 is used to construct a thermal map representing a two-dimensional skeleton of a human body according to the original model parameters as an original thermal map;
[0229] The human body reconstruction network training module 604 is used to train the human body reconstruction network using the original clothing binary image and the original thermal map as samples and the original model parameters as labels.
[0230] In one embodiment of the present invention, the original clothing binary image construction module 602 includes:
[0231] A three-dimensional model building module, used to build a three-dimensional model of a human body using the original model parameters;
[0232] The original camera parameter sampling module is used to sample the camera parameters as the original camera parameters;
[0233] An original human body binary image projection module is used to project the surface mesh of the human body in the three-dimensional model into a binary image representing the two-dimensional contour of the human body as the original human body binary image according to the original camera parameters;
[0234] A clothing binary image acquisition module, used to acquire a binary image representing clothing as a clothing binary image;
[0235] The clothing binary image superposition module is used to superimpose the clothing binary image on the original human body binary image to generate a binary image representing the two-dimensional contour of the human body when wearing the clothing as the original clothing binary image.
[0236] In one embodiment of the present invention, the clothing binary image superposition module includes:
[0237] An anchor point selection module, used for selecting anchor points in the original human body binary image, wherein the anchor points are used for locating the upper body and / or lower body of the two-dimensional contour of the human body;
[0238] A superimposed binary image selection module, used for selecting the clothing binary image suitable for superimposing on the upper body and / or the lower body as the superimposed binary image;
[0239] The affine transformation module is used to perform affine transformation on the superimposed binary image until the distance between the contour of the superimposed binary image and the anchor point meets the preset superposition condition, thereby obtaining a binary image representing the two-dimensional contour of the human body when dressed as the original dressing binary image.
[0240] In one embodiment of the present invention, the clothing binary image superposition module further includes:
[0241] An image enhancement processing module, used for performing image enhancement processing on the original clothing binary image;
[0242] The image enhancement processing includes at least one of the following:
[0243] The original clothing binary image is randomly blocked, noise is randomly added to the original clothing binary image, and Gaussian blur processing is performed on the original clothing binary image.
[0244] In one embodiment of the present invention, the original heat map construction module 603 includes:
[0245] A three-dimensional model building module, used to build a three-dimensional model of a human body using the original model parameters;
[0246] The original camera parameter sampling module is used to sample the camera parameters as the original camera parameters;
[0247] The orthogonal projection module is used to perform orthogonal projection on the skeleton points of the human body in the three-dimensional model according to the original camera parameters to obtain a thermal map representing the two-dimensional skeleton of the human body as a target thermal map.
[0248] In one embodiment of the present invention, the original camera parameter sampling module includes:
[0249] A weak projection modeling module, used for determining that the projection of the three-dimensional model to two dimensions is a weak projection under the condition that the original model parameters include rotation parameters;
[0250] The camera parameter setting module is used to set the displacement and scale of the camera as the original camera parameters in the weak projection.
[0251] In one embodiment of the present invention, the human body reconstruction network training module 604 includes:
[0252] An original camera parameter query module, used to query original camera parameters, where the original camera parameters are used to project the three-dimensional model into two dimensions;
[0253] A parameter prediction module, used for inputting the original clothing binary image and the original thermal map into a human body reconstruction network to predict parameters of a human body three-dimensional model and parameters of a camera as predicted model parameters and predicted camera parameters;
[0254] A first loss value calculation module, used for calculating the difference between the original model parameters, the original camera parameters and the predicted model parameters, the predicted camera parameters as a first loss value;
[0255] A human body reconstruction network updating module, used for updating the human body reconstruction network according to the first loss value;
[0256] The iteration step number judgment module is used to judge whether the number of steps of the current iteration reaches a preset threshold; if so, the training completion determination module is called; if not, the parameter prediction module is returned to be called;
[0257] The training completion determination module is used to determine whether the training of the human body reconstruction network is completed.
[0258] In one embodiment of the present invention, the first loss value calculation module includes:
[0259] A first difference value calculation module, used for calculating the norm distance between each of the original model parameters and each of the prediction model parameters as a first difference value;
[0260] A second difference value calculation module, used for calculating a norm distance between each of the original camera parameters and each of the predicted camera parameters as a second difference value;
[0261] The first average value calculation module is used to calculate an average value of the first difference value and the second difference value as a first loss value.
[0262] In one embodiment of the present invention, the human body reconstruction network update module includes:
[0263] A predicted human body binary image construction module is used to construct a binary image representing the two-dimensional contour of the human body when wearing clothes according to the prediction model parameters as the predicted human body binary image;
[0264] A prediction heat map construction module, used to construct a heat map representing a two-dimensional skeleton of a human body according to the prediction model parameters as a prediction heat map;
[0265] A second loss value calculation module, used for calculating the difference between the original clothing binary image and the predicted human body binary image as a second loss value;
[0266] A third loss value calculation module, used to calculate the difference between the original heat map and the predicted heat map as a third loss value;
[0267] a total loss value calculation module, configured to merge the first loss value, the second loss value and the third loss value to obtain a total loss value;
[0268] A weight updating module is used to adjust the weights in the human body reconstruction network according to the total loss value.
[0269] In one embodiment of the present invention, the predicted human body binary image construction module includes:
[0270] A prediction three-dimensional model construction module, used to construct a three-dimensional model of a human body using the prediction model parameters as a prediction three-dimensional model;
[0271] A candidate camera parameter sampling module is used to sample camera parameters as candidate camera parameters;
[0272] A candidate human body binary image projection module is used to project the surface mesh of the human body in the predicted three-dimensional model into a binary image representing the two-dimensional contour of the human body as a candidate human body binary image according to the candidate camera parameters;
[0273] A clothing binary image acquisition module, used to acquire a binary image representing clothing as a clothing binary image;
[0274] The predicted clothing image superposition module is used to superimpose the clothing binary image on the candidate human body binary image to generate a binary image representing the two-dimensional contour of the human body when wearing the clothing as the predicted clothing binary image.
[0275] In one embodiment of the present invention, the predicted clothing image superposition module includes:
[0276] An anchor point selection module, used for selecting anchor points in the original human body binary image, wherein the anchor points are used for locating the upper body and / or lower body of the two-dimensional contour of the human body;
[0277] A superimposed binary image selection module, used for selecting the clothing binary image suitable for superimposing on the upper body and / or the lower body as the superimposed binary image;
[0278] The predicted affine transformation module is used to perform affine transformation on the superimposed binary image until the distance between the contour of the superimposed binary image and the anchor point meets the preset superposition condition, thereby obtaining a binary image representing the two-dimensional contour of the human body when dressed as the predicted dressing binary image.
[0279] In one embodiment of the present invention, the predicted clothing image superposition module further includes:
[0280] An image enhancement processing module, used for performing image enhancement processing on the predicted clothing binary image;
[0281] The image enhancement processing includes at least one of the following:
[0282] The candidate clothing binary image is randomly blocked, noise is randomly added to the candidate clothing binary image, and Gaussian blur processing is performed on the candidate clothing binary image.
[0283] In one embodiment of the present invention, the prediction heat map construction module includes:
[0284] A prediction three-dimensional model construction module, used to construct a three-dimensional model of a human body using the prediction model parameters as a prediction three-dimensional model;
[0285] A candidate camera parameter sampling module is used to sample camera parameters as candidate camera parameters;
[0286] The predicted orthogonal projection module is used to perform orthogonal projection on the skeleton points of the human body in the predicted three-dimensional model according to the original camera parameters to obtain a thermal map representing the two-dimensional skeleton of the human body as the predicted thermal map.
[0287] In one embodiment of the present invention, the candidate camera parameter sampling module includes:
[0288] A weak projection modeling module, used for determining that the projection of the candidate three-dimensional model to two dimensions is a weak projection under the condition that the candidate model parameters include rotation parameters;
[0289] The camera parameter setting module is used to set the displacement and scale of the camera as candidate camera parameters in the weak projection.
[0290] In one embodiment of the present invention, the second loss value calculation module includes:
[0291] A third difference value calculation module, used for calculating the norm distance between each pixel point in the original clothing binary image and each pixel point in the predicted human body binary image as a third difference value;
[0292] The second average value calculation module is used to calculate an average value of the third difference value as a second loss value.
[0293] In one embodiment of the present invention, the third loss value calculation module includes:
[0294] A fourth difference value calculation module, used for calculating the norm distance between each skeleton point in the original heat map and each skeleton point in the predicted heat map as a fourth difference value;
[0295] The third average value calculation module is used to calculate an average value of the fourth difference value as a third loss value.
[0296] In one embodiment of the present invention, the total loss value calculation module includes:
[0297] A first weighted adjustment loss value calculation module, used for multiplying the first loss value by a preset first coefficient to obtain a first weighted adjustment loss value;
[0298] A second weighted adjustment loss value calculation module, used for multiplying the second loss value by a preset second coefficient to obtain a second weighted adjustment loss value;
[0299] A third weight adjustment loss value calculation module, used for multiplying the third loss value by a preset third coefficient to obtain a third weight adjustment loss value;
[0300] a weighted adjustment loss value summing module, configured to calculate a sum of the first weighted adjustment loss value, the second weighted adjustment loss value, and the third weighted adjustment loss value as a total loss value;
[0301] The first coefficient is greater than the third coefficient, and the third coefficient is greater than the second coefficient.
[0302] The training device for a human body reconstruction network provided in an embodiment of the present invention can execute the training method for a human body reconstruction network provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0303] Embodiment 5
[0304] Figure 7 This is a structural block diagram of a human body reconstruction device provided in Embodiment 5 of the present invention, which may specifically include the following modules:
[0305] The target image data acquisition module 701 is used to acquire two-dimensional target image data collected from the user;
[0306] The target image data conversion module 702 is used to convert the target image data into a binary image representing the two-dimensional outline of the human body when wearing clothes, as the target clothing binary image, and a thermal image representing the two-dimensional skeleton of the human body, as the target thermal image;
[0307] A target model parameter prediction module 703 is used to input the original clothing binary image and the target heat map into a human body reconstruction network to predict parameters of a three-dimensional model of the human body as target model parameters;
[0308] The human body three-dimensional model building module 704 is used to build a three-dimensional human body model for the user using the target model parameters.
[0309] Wherein, the training method of the human body reconstruction network is as follows:
[0310] Sampling parameters in the three-dimensional human body model as original model parameters;
[0311] Constructing a binary image representing the two-dimensional contour of the human body when wearing clothes according to the original model parameters as the original clothing binary image;
[0312] Constructing a thermal map representing a two-dimensional skeleton of a human body according to the original model parameters as an original thermal map;
[0313] The human body reconstruction network is trained using the original clothing binary image and the original heat map as samples and the original model parameters as labels.
[0314] The human body reconstruction device provided in the embodiment of the present invention can execute the human body reconstruction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0315] Embodiment 6
[0316] Figure 8 This is a structural block diagram of a fitting device provided in Embodiment 6 of the present invention, which may specifically include the following modules:
[0317] Clothes determination parameters 801, used to determine the clothes selected by the user, the three-dimensional model of the clothes having clothes model parameters;
[0318] The target image data acquisition module 802 is used to collect two-dimensional target image data of the user;
[0319] The target image data conversion module 803 is used to convert the target image data into a binary image representing the two-dimensional outline of the human body when wearing clothes, as the target clothing binary image, and a thermal image representing the two-dimensional skeleton of the human body, as the target thermal image;
[0320] A target model parameter prediction module 804 is used to input the target clothing binary image and the target heat map into a human body reconstruction network to predict parameters of a three-dimensional model of the human body as target model parameters;
[0321] The dressing model building module 805 is used to build and display a dressing model using the target model parameters and the clothing model parameters, wherein the dressing model is used to represent the user wearing the clothing.
[0322] Wherein, the training method of the human body reconstruction network is as follows:
[0323] Sampling parameters in the three-dimensional human body model as original model parameters;
[0324] Constructing a binary image representing the two-dimensional contour of the human body when wearing clothes according to the original model parameters as the original clothing binary image;
[0325] Constructing a thermal map representing a two-dimensional skeleton of a human body according to the original model parameters as an original thermal map;
[0326] The human body reconstruction network is trained using the original clothing binary image and the original heat map as samples and the original model parameters as labels.
[0327] The fitting device provided in the embodiment of the present invention can execute the fitting method provided in any embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method.
[0328] Embodiment 7
[0329] Fig. 9 A schematic diagram of the structure of a computer device provided in Embodiment 7 of the present invention. Fig. 9 A block diagram of an exemplary computer device 12 suitable for use in implementing embodiments of the present invention is shown. Fig. 9 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0330] like Fig. 9 As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).
[0331] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0332] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0333] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Fig. 9 not shown, usually called a "hard drive"). Although Fig. 9 Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.
[0334] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0335] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), may communicate with one or more devices that enable a user to interact with the computer device 12, and / or may communicate with any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the computer device 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computer device 12, including, but not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0336] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the training method of the human body reconstruction network or the human body reconstruction method or the fitting method provided in the embodiment of the present invention.
[0337] Embodiment 8
[0338] Embodiment 8 of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the various processes of the above-mentioned human body reconstruction network training method or human body reconstruction method or fitting method, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0339] Among them, computer-readable storage media may include, for example, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.
[0340] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A training method for a human body reconstruction network, It is characterized in that include: Sampling parameters in the three-dimensional human body model as original model parameters; Obtaining an original human body binary image according to the original model parameters; Obtain a binary image representing clothes as a clothes binary image; Selecting anchor points in the original human body binary image, wherein the anchor points are used to locate the upper body and / or lower body of the two-dimensional contour of the human body; Selecting the clothing binary image suitable for superimposing on the upper body and / or the lower body as the superimposed binary image; Performing an affine transformation on the superimposed binary image until the distance between the outline of the superimposed binary image and the anchor point meets a preset superposition condition, thereby obtaining a binary image representing the two-dimensional outline of the human body when wearing the clothes, as the original clothing binary image; Constructing a thermal map representing a two-dimensional skeleton of a human body according to the original model parameters as an original thermal map; The human body reconstruction network is trained using the original clothing binary image and the original heat map as samples and the original model parameters as labels.
2. The method according to claim 1, It is characterized in that The step of obtaining an original human body binary image according to the original model parameters comprises: Using the original model parameters to construct a three-dimensional model of the human body; The parameters of the sampled camera are used as the original camera parameters; The surface mesh of the human body in the three-dimensional model is projected into a binary image representing the two-dimensional contour of the human body according to the original camera parameters, as the original human body binary image.
3. The method according to claim 1, It is characterized in that The step of obtaining a binary image representing a two-dimensional outline of a human body when wearing clothes, as an original clothing binary image, includes: Performing image enhancement processing on the original clothing binary image; The image enhancement processing includes at least one of the following: The original clothing binary image is randomly blocked, noise is randomly added to the original clothing binary image, and Gaussian blur processing is performed on the original clothing binary image.
4. The method according to claim 1, It is characterized in that The step of constructing a thermal map representing a two-dimensional skeleton of a human body according to the original model parameters as an original thermal map comprises: Using the original model parameters to construct a three-dimensional model of the human body; The parameters of the sampled camera are used as the original camera parameters; Orthogonal projection is performed on the skeleton points of the human body in the three-dimensional model according to the original camera parameters to obtain a thermal map representing the two-dimensional skeleton of the human body as the original thermal map.
5. The method according to claim 2 or 4, It is characterized in that The parameters of the sampled camera, as original camera parameters, include: Under the condition that the original model parameters include rotation parameters, determining that the projection of the three-dimensional model to the two-dimensional is a weak projection; In the weak projection, the displacement and scale of the camera are set as original camera parameters.
6. The method according to any one of claims 1 to 4, It is characterized in that The method of training a human body reconstruction network using the original clothing binary image and the original heat map as samples and the original model parameters as labels includes: Querying original camera parameters, where the original camera parameters are used to project the three-dimensional model into two dimensions; Inputting the original clothing binary image and the original heat map into a human body reconstruction network to predict parameters of a human body three-dimensional model and parameters of a camera as predicted model parameters and predicted camera parameters; Calculating the difference between the original model parameters, the original camera parameters and the predicted model parameters, the predicted camera parameters as a first loss value; Updating the human body reconstruction network according to the first loss value; Determine whether the number of steps of the current iteration reaches a preset threshold; if so, determine that the training of the human body reconstruction network is completed; if not, return to execute the input of the original clothing binary image and the original heat map into the human body reconstruction network to predict the parameters describing the human body three-dimensional model and the parameters of the camera as the predicted model parameters and the predicted camera parameters.
7. The method according to claim 6, It is characterized in that The difference between the original model parameters, the original camera parameters and the predicted model parameters, the predicted camera parameters, as a first loss value, comprises: Calculating a norm distance between each of the original model parameters and each of the predicted model parameters as a first difference value; Calculating a norm distance between each of the original camera parameters and each of the predicted camera parameters as a second difference value; An average value is calculated for the first difference value and the second difference value as a first loss value.
8. The method according to claim 6, It is characterized in that The updating of the human body reconstruction network according to the first loss value comprises: Constructing a binary image representing the two-dimensional contour of a human body when wearing clothes according to the prediction model parameters as a predicted human body binary image; Constructing a thermal map representing a two-dimensional skeleton of a human body according to the prediction model parameters as a prediction thermal map; Calculate the difference between the original clothing binary image and the predicted human body binary image as a second loss value; Calculate the difference between the original heat map and the predicted heat map as a third loss value; Merging the first loss value, the second loss value and the third loss value to obtain a total loss value; The weights in the human body reconstruction network are adjusted according to the total loss value.
9. The method according to claim 8, It is characterized in that The calculating the difference between the original clothing binary image and the predicted human body binary image as a second loss value includes: Calculate the norm distance between each pixel point in the original clothing binary image and each pixel point in the predicted human body binary image as a third difference value; An average value is calculated for the third difference values as the second loss value.
10. The method according to claim 8, It is characterized in that The calculating the difference between the original heat map and the predicted heat map as a third loss value includes: Calculate the norm distance between each bone point in the original heat map and each bone point in the predicted heat map as a fourth difference value; An average value is calculated for the fourth difference values as the third loss value.
11. The method according to any one of claims 8 to 10, It is characterized in that The fusing the first loss value, the second loss value and the third loss value to obtain a total loss value includes: Multiplying the first loss value by a preset first coefficient to obtain a first weighted adjustment loss value; Multiplying the second loss value by a preset second coefficient to obtain a second weighted adjustment loss value; Multiplying the third loss value by a preset third coefficient to obtain a third weight adjustment loss value; Calculating the sum of the first weighted adjustment loss value, the second weighted adjustment loss value, and the third weighted adjustment loss value as a total loss value; The first coefficient is greater than the third coefficient, and the third coefficient is greater than the second coefficient.
12. A human body reconstruction method, It is characterized in that include: Acquire two-dimensional target image data collected from the user; Convert the target image data into a binary image representing the two-dimensional outline of a human body when dressed, as a target dressing binary image, and a thermal image representing the two-dimensional skeleton of a human body, as a target thermal image; Inputting the target clothing binary image and the target heat map into a human body reconstruction network trained by the method according to any one of claims 1 to 11, and predicting parameters of a three-dimensional model of the human body as target model parameters; A three-dimensional human body model is constructed for the user using the target model parameters.
13. A method for fitting clothes, It is characterized in that Applied to a mobile terminal, the method comprises: Determining clothes selected by a user, wherein the three-dimensional model of the clothes has clothes model parameters; Collecting two-dimensional target image data from the user; Convert the target image data into a binary image representing the two-dimensional outline of a human body when dressed, as a target dressing binary image, and a thermal image representing the two-dimensional skeleton of a human body, as a target thermal image; Inputting the target clothing binary image and the target heat map into a human body reconstruction network trained by the method according to any one of claims 1 to 11, and predicting parameters of a three-dimensional model of the human body as target model parameters; A three-dimensional clothing model is constructed and displayed using the target model parameters and the clothing model parameters, wherein the clothing model is used to represent the user wearing the clothing.
14. A computer device, It is characterized in that The computer device comprises: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the human body reconstruction network as described in any one of claims 1 to 11, the human body reconstruction method as described in claim 12, or the fitting method as described in claim 13.
15. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the training method of the human body reconstruction network as described in any one of claims 1 to 11, the human body reconstruction method as described in claim 12, or the fitting method as described in claim 13.
Citation Information
Patent Citations
Training method of SMPL parameter prediction model, server and storage medium
CN109859296A
Parameterized model of 2d articulated human shape
US20130249908A1