A clothing model driving method, device and storage medium
By combining fabric simulation and skinning methods, deep neural networks and mathematical models are used to drive the posture changes of human body models, solving the problem of slow modeling speed and insufficient accuracy in virtual fittings, and achieving high-reality virtual fitting effects to meet the convenient needs of the Internet.
Patent Information
- Application Number
- CN202010876682.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-27
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2040-08-27
AI Technical Summary
When generating human body models and clothing models, existing virtual fitting technology has problems such as slow modeling speed, insufficient accuracy, large calculation volume, high hardware cost, and insufficient matching degree of clothing and human body, resulting in unreal virtual fitting effect.
The fabric simulation and skinning method are combined to generate an accurate three-dimensional model of the human body through a deep neural network analysis, and a standard human body model is constructed by combining mathematical models to drive the human body model from the initial posture to the target posture. The fabric simulation and skin method are used during the movement of the clothing to simulate the collision and deformation of the clothing and the human body, and generate a realistic virtual fitting effect.
It achieves a high-reality virtual fitting effect, the clothes naturally follow the changes in the human body state, the fabric texture is restored with a realistic texture, the user's operation is simple, the calculation amount is reasonable, and it meets the convenient needs of the Internet era.
Smart Images

Figure CN114119908B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of user virtual dressing and fitting, and specifically relates to human body modeling, clothing modeling, and the fitting of clothing models with human body models used in virtual dressing, especially a driving method, device and storage medium for a three-dimensional clothing model worn on a custom human body model. Background Art
[0002] With the development of internet technology, online shopping has become increasingly popular. Compared to in-store shopping, online shopping offers advantages such as a wider variety of products and greater convenience. However, online shopping also presents some difficult-to-solve challenges, the most prominent of which is the inability to physically view the items. This issue is particularly prominent in clothing. Unlike in-store shopping, where customers can change outfits and see how they look in real time, online clothing shopping offers no personalized images. Instead, they only provide pictures of models trying on clothes, or sometimes even no pictures at all. This prevents consumers from visually assessing the degree to which the clothing matches their body image in real time, leading to a high number of returns and exchanges.
[0003] To address this issue, businesses are attempting to utilize virtual fitting technology to provide consumers with simulated fitting experiences. Of course, there are other practical applications for virtual fitting technology, such as in online games. Consequently, this technology has seen rapid development.
[0004] Virtual fitting refers to a technology application that allows users to view the desired outfit changes in real time on a terminal screen, without having to physically put on the desired clothing. Existing fitting technologies primarily include two-dimensional fitting and three-dimensional virtual fitting. The former essentially captures a user's image and clothing images, then stretches or compresses the clothing to the same size as the person, then crops and splices the pieces to create a "dressed" image. However, these images lack realism due to crude image processing, completely ignoring the user's actual body shape and simply forcing the clothing onto the user's photo, failing to meet user needs. The latter typically uses 3D acquisition equipment to capture 3D information about the person and synthesize it with clothing features. Alternatively, users manually input body data, generate a virtual 3D human body mesh according to specific rules, and then combine it with clothing textures. Overall, this type of 3D virtual fitting requires extensive data acquisition and 3D data calculations, resulting in high hardware costs and limited adoption for the general public.
[0005] With the development of cloud computing technology, artificial intelligence technology, and intelligent terminal processing capabilities, two-dimensional virtual fitting technology has emerged. This technology mainly includes three steps: (1) processing the personal body information provided by the user to obtain a target human body model; (2) processing the clothing information to obtain a clothing model; (3) fusing the human body model and clothing model to generate a simulated image of the person wearing the clothing.
[0006] Regarding point (1), due to the accumulation of many uncertain factors such as process design, model parameter selection, and neural network training methods, the quality of the final generated clothing-changing pictures is not as good as that of traditional three-dimensional virtual fitting technology. Among them, the establishment of the human body model is its basic step, and the subsequent dressing process must also be based on the human body model generated previously. Therefore, once the human body model is generated inaccurately, it is easy to cause problems such as a large difference in body shape between the human body model and the person being fitted, loss of skin texture, loss of body parts, etc., which will affect the final generated clothing-changing picture effect.
[0007] In the field of image processing, three-dimensional reconstruction refers to the establishment of a mathematical model suitable for computer representation and processing of three-dimensional objects. It is the basis for processing, operating and analyzing their properties in a computer environment. It is also a key technology for establishing virtual reality in computers to express the objective world. It is widely used in computer animation, virtual reality, industrial inspection and other fields.
[0008] In the general field of computer vision, human body modeling can be achieved through a variety of approaches. These typically include 3D scanning of the real body using omnidirectional scanning equipment, 3D reconstruction methods based on multi-view depth-of-field photography, and methods that combine a given image with a human model to achieve 3D reconstruction. Using 3D scanning equipment to scan the real body provides the most information and is the most accurate. However, this equipment is typically expensive and requires the close cooperation of the human model. The entire processing process places high demands on the processing equipment, so it is generally used in specialized fields. Secondly, multi-view 3D reconstruction methods require providing overlapping images of the reconstructed body from multiple perspectives and establishing spatial transformations between them. Using multiple cameras to capture multiple images and stitch them together to create a 3D model is relatively simple, but computational complexity remains high. In most cases, multi-angle images can only be obtained by on-site personnel. The model created by stitching together the textures obtained from multi-angle depth camera photography lacks body scale data and cannot provide a foundation for 3D perception. Secondly, the single-image method combined with a human body model only requires a single image. A neural network-based intelligent method for generating 3D human feature curves uses neural network training to obtain weights and thresholds that can be used to describe curves of parts such as the neck, chest, waist, and hips. Then, based on dimensional parameters such as the girth, width, and thickness of the human cross-section, a 3D curve that matches the real human body shape can be directly generated, resulting in a predicted human body model. However, this method requires relatively little input information, and the solution process still requires a high level of computational effort, resulting in unsatisfactory final model results.
[0009] Regarding point (2), there are several different methods in the existing technology for generating three-dimensional clothing models. At present, the more traditional way to build a three-dimensional clothing model is based on the design and stitching method of two-dimensional clothing pieces. This method requires a certain amount of clothing expertise to design the sample, which is not a quality that all virtual fitting users have. At the same time, this method also requires manual specification of the stitching relationship between the samples, which will consume a lot of time to set up. In addition, another relatively new three-dimensional modeling method is based on hand-drawing, which can generate a simple clothing model through the line information hand-drawn by the user. However, this method requires professional personnel to hand-draw, and its reproducibility and repeatability are poor. It takes a lot of time for users to draw the details of the clothing, and it is difficult to promote it on a large scale in e-commerce. Both methods tend to innovate and design new clothing rather than perform three-dimensional modeling on existing clothing for sale. Another method is to obtain clothing picture information and use image processing technology and graphic simulation technology in combination to finally generate a virtual three-dimensional clothing model. The contour and size of the clothing are obtained through contour detection and classification in the image. The key points of the edges are found from the contour through machine learning methods. The stitching information is generated based on the correspondence between the key points. Finally, the physical stitching of the clothing is simulated in three-dimensional space to obtain the real effect of the clothing being worn on the human body.
[0010] Regarding point (3), the virtual fitting rooms commonly found in the current market mainly focus on style matching, without directly simulating the natural properties of the collision between virtual characters and clothing fabrics, and therefore still lack a great deal of realism. Currently, more and more manufacturers are increasing the cohesion between the virtual world and the real world by using virtual characters to vividly represent user postures, and by simulating the collision response between clothing fabrics and the human body in real time and rendering them in real time. This has brought more fun to virtual fitting users when changing clothes, and has also allowed more people to enjoy the convenience brought by purchasing clothes.
[0011] In summary, based on the characteristics of Internet technology and the network environment in which it is located, the method of directly outputting the final image or photo after changing clothes from a single human body image is undoubtedly the most preferred. It is the most convenient and the user does not need to be present at the scene. With just one photo, the entire virtual dressing process can be completed. Then the problem that follows is that as long as the resulting photo effect can be guaranteed to be basically equivalent to the real 3D simulated dressing process, it will become the mainstream. Among them, (1) how to obtain a human body model that is closest to the real state of the human body through a photo, and (2) how to put the three-dimensional clothing model on the target human body model in the state closest to the real state, have become the two most important and unavoidable problems in the virtual dressing method.
[0012] Regarding the first point. In the prior art, there are usually several methods for constructing human body models: (1) Regression-based methods, which use convolutional neural networks to reconstruct a voxel-represented human body model. The algorithm first estimates the positions of the main joints of the human body based on the input image, and then estimates whether each unit voxel in the voxel grid of a given size is occupied based on the key point positions, thereby using the entire shape of the occupied voxels to describe the reconstructed human body shape; (2) Human body reconstruction based on a single image, which simultaneously estimates the three-dimensional shape and posture of the human body. This method first roughly marks the simple key points of the human skeleton on the image, and then performs initial matching and fitting of the human body model based on these rough key points to obtain the approximate shape of the human body. (3) Use 23 bone nodes to represent the human skeleton, and then use the rotation of each bone node to represent the posture of the entire human body. At the same time, use 6890 vertex positions to express the human body shape. In the fitting process, the shape and posture parameters are fitted at the same time given the bone node positions, thereby performing three-dimensional human body reconstruction; or first use a CNN model to predict the key points on the image, and then use the SMPL model for fitting to obtain the initial human body model. Next, the fitted shape parameters are used to regress a bounding box for each joint. Each joint corresponds to a bounding box, represented by its axis length and radius. Finally, the initial model and the regressed bounding boxes are combined to create a 3D human reconstruction. However, these methods suffer from slow modeling speed, insufficient modeling accuracy, and a strong dependence on the established body and posture database.
[0013] Prior art 1 discloses a human body modeling method based on body measurement data, such as Figure 1 As shown, the method includes: obtaining anthropometry data; performing linear regression on a pre-created human body model using a pre-trained prediction model based on the anthropometry data to obtain a predicted human body model, wherein the pre-created human body model includes multiple pre-defined sets of marker feature points and corresponding standard shape bases, and the anthropometry data includes measurement data corresponding to each set of marker feature points; and obtaining a target human body model based on the predicted human body model, wherein the target human body model includes measurement data, a target shape base, and a target shape coefficient. However, this method requires very high anthropometry data, including body length and girth data, such as height, arm length, shoulder width, leg length, calf length, thigh length, foot length, head circumference, chest circumference, waist circumference, and thigh circumference. This requires not only measurement but also calculation. While this method does save computational effort, the user experience is very poor and the process is cumbersome. Furthermore, the training of the human body model is based on the training method of the SMPL model.
[0014] The SMPL model is a parametric human body model proposed by the Max Planck Institute for Human Modeling. This method can model and animate any human body. The key difference between this method and traditional location-based systems lies in its proposed method for surface topography of human pose images. This method simulates the convexity and concavity of human muscles during limb movement. This method avoids surface distortion during human motion and accurately depicts the topography of muscle extension and contraction. In this method, β and θ are input parameters. β represents 10 parameters related to a person's height, weight, head-to-body ratio, and other proportions, while θ represents 75 parameters representing the overall motion pose and the relative angles of 24 joints. However, this model generation method relies on the accumulation of extensive training data to obtain the relationship between body shape and shape bases. However, due to the strong interdependence between these relationships, each shape base cannot be independently controlled, making decoupling difficult. For example, the arms and legs are also related; theoretically, when the arms move, the legs will also move. This makes it difficult to improve the SMPL model for different body types.
[0015] Prior art II discloses a 3D human modeling method based on a single photo, comprising: obtaining a photo, parsing the photo, marking key points of the human body in the photo, and calculating the spatial coordinates of the key points; obtaining the distances between skeletal points in a pre-created standard human body model and key points in the photo, aligning the skeletal points with the key points to generate a basic human body model; obtaining a basic texture map from the pre-created standard human body model, performing a difference calculation between the basic texture map and the skin texture of the face in the photo, and then fusing the difference using an edge channel to generate basic texture data; and generating a 3D human body model based on the basic human body model and the basic texture data. 3D human body modeling is achieved using a single photo, and the model is supported by a skeletal and muscular system, enabling expression and movement. However, this method matches the distances between key points in the user's photo and key points in the standard human body model, then adjusts the distances to achieve the target human pose. Subsequently, a difference calculation and fusion are performed between the basic texture map and the skin texture in the photo to obtain the final human body model. This method is simple and computationally inefficient, but the accuracy and realism of the resulting human body model are not very high.
[0016] Prior art three discloses a method for generating a three-dimensional human body model, comprising: obtaining a two-dimensional human body image; inputting the two-dimensional human body image into a three-dimensional human body parameter model to obtain three-dimensional human body parameters corresponding to the two-dimensional human body image; inputting the training sample into a neural network for training to obtain a three-dimensional human body parameter model, including: inputting the standard two-dimensional human body image in the training sample into the neural network to obtain predicted three-dimensional human body parameters corresponding to the standard two-dimensional human body image; adjusting a three-dimensional flexible deformable model based on the predicted three-dimensional human body parameters to obtain a predicted three-dimensional human body model; and obtaining the predicted joint point positions in the standard two-dimensional human body image through reverse mapping based on the joint point positions in the predicted three-dimensional human body model. This modeling method uses model judgment and the parameters output by the neural network to only focus on node parameters. The parameters are then adjusted in detail to align with the target human body posture using the mature body shape of the SMPL model. Although the computational complexity is reduced, due to the small number of input parameters and the fact that the adjustments can only be completed based on the SMPL prediction model, it is difficult to output a particularly ideal human body model that is highly consistent with the target human body posture.
[0017] Regarding the second point. There is such a virtual fitting solution in the prior art, including: obtaining a dressed reference human body model and an undressed target human body model; embedding skeletons of the same hierarchical structure for the reference human body model and the target human body model respectively; performing skin binding on the skeletons of the reference human body model and the target human body model; calculating the rotation amount of the bones in the skeleton of the target human body model, and recursively adjusting all the bones in the skeleton of the target human body model so that the posture of the target human body model skeleton is consistent with that of the reference human body model skeleton; using the LBS skinning algorithm to deform the skin of the target human body model according to the rotation amount of the bones in the skeleton of the target human body model; based on the skin deformation of the target human body model, migrating the clothing model from the reference human body model to the target human body model. This invention can reduce the difficulty of migrating the clothing model from the reference human body to the target human body after the postures of the target human body model and the reference human body model are adjusted to be consistent, and convert the inefficient non-rigid registration problem into an efficient rigid registration problem, thereby realizing the migration of the clothing model from the reference human body model to the target human body model. This approach solves the technical challenge of automatically fitting clothing on different bodies and in different poses, while maintaining the same size before and after fitting. However, this pure skinning approach places too much emphasis on a fixed distance between clothing and skin. While this approach offers advantages in fitting speed, it suffers from significant disadvantages in terms of clothing matching and realism, making it suitable only for situations where clothing needs to be quickly and easily adapted to follow the movement of the skin mesh.
[0018] Regarding the third point. Prior art four discloses a virtual fitting method, which includes: obtaining a dressed reference human body model and an undressed target human body model; embedding skeletons of the same hierarchical structure into the reference human body model and the target human body model respectively; performing skin binding on the skeletons of the reference human body model and the target human body model; calculating the rotation of the bones in the skeleton of the target human body model, and recursively adjusting all bones in the skeleton of the target human body model so that the posture of the target human body model skeleton is consistent with that of the reference human body model skeleton; using the LBS skinning algorithm to deform the skin of the target human body model according to the rotation of the bones in the skeleton of the target human body model; based on the skin deformation of the target human body model, migrating the clothing model from the reference human body model to the target human body model; wherein, migrating the clothing model from the reference human body model to the target human body model specifically includes: using an iterative closest point algorithm to rigidly align the target human body model after skin deformation with the reference human body model to obtain an affine transformation; and applying the affine transformation to the clothing model, thereby achieving the migration of the clothing model from the reference human body model to the target human body model. This method improves the specific implementation of skin-driven, but in fact it still does not deviate from the idea of relying on skeletal skinning to complete movement, and the actual clothing simulation effect has not been substantially improved.
[0019] To keep pace with the development trends of the internet industry, in the niche field of virtual fitting, minimal input information, minimal computational effort, and optimal results are three fundamental goals that are constantly being pursued. This invention addresses the issues of realism and fidelity in garment model movement. While ensuring effective garment simulation, it also improves simulation speed in certain areas and processes, attempting to find an optimal balance between these three. The goal is to provide a virtual fitting method that allows for simple input, minimizes computational effort within the capabilities of the terminal device, and produces results similar to those of real clothing. Summary of the Invention
[0020] Based on the above problems, the present invention provides a clothing model driving method, device and storage medium that overcome the above problems.
[0021] The present invention provides a clothing model driving method, which includes: obtaining a two-dimensional image of clothing; making a three-dimensional model of the clothing according to the two-dimensional image of the clothing; constructing a three-dimensional standard human body model in combination with a mathematical model, wherein the three-dimensional standard model is in an initial posture; fitting the three-dimensional clothing model to the three-dimensional standard human body model in the initial posture; obtaining a two-dimensional image of a target human body; obtaining parameters of the three-dimensional target human body model through calculation of a secondary neural network model; inputting several groups of obtained posture and body shape parameters into the three-dimensional standard human body model for fitting; driving the human body model to move from the initial posture to the target posture; and obtaining a target human body model having the same posture and body shape as the target human body and wearing the changed clothing.
[0022] Preferably, the posture of the target human body is determined by the three-dimensional human body action posture parameters, and the skeleton is driven to move from the initial posture to the target posture so that the three-dimensional human body model is basically consistent with the target human body posture. The movement of the clothing model adopts the combination of cloth simulation and skinning method. The part that does not deform substantially after the clothing moves adopts the skinning method, and the part that deforms greatly during the clothing movement adopts the cloth simulation method. In the driving process, the relationship between the model mesh vertices and the skeleton is set up. After the rotation parameters of the skeleton under the target posture are calculated, the relationship between the skeleton and the skeleton is used to drive the vertices near it to complete the movement in space, and the associated vertices are driven to arrive at the position of the target posture. The moving state from the initial posture to the target posture is calculated frame by frame. When each frame is calculated, the skinning drive is completed first and the cloth simulation is calculated again. After the current frame is calculated, the state of the next frame is calculated again.
[0023] Cloth simulation utilizes a collision system, modeling the mannequin's vertex mesh as a rigid body and the clothing model's mesh as a non-rigid body. A physics engine is used to simulate the collision relationship between these two bodies. During the cloth simulation, collisions between these two bodies are calculated, while also considering the connection forces between the clothing model's own meshes. The mesh state of the clothing model is calculated frame by frame, simulating the collision motion process in the physical world. To drive the mannequin to the target pose, gravity calculations are performed on the clothing fabric for several frames. The continuity of the motion changes is determined using a neural network prediction model. The target pose is analyzed to determine whether it is a candid shot or a posed pose. If the pose is a candid shot, the clothing model will maintain its velocity during motion due to inertia, making the target pose state unstable. Therefore, the number of gravity calculations is reduced when the mannequin reaches the target pose. If the pose is posed, the number of gravity calculations is increased when the mannequin reaches the target pose to ensure the realism of the clothing fabric in the target pose.
[0024] Preferably, the fitting movement step further includes the following sub-steps: 1) obtaining the position coordinates of the initial posture and the target posture; 2) generating an animation sequence moving from the initial posture to the target posture; 3) in the process of generating the animation sequence, processing is performed in a grid interpolation manner; 4) the interpolation speed is set to be slow in the front and back distances between the initial point and the target point, and fast in the intermediate movement process; 5) when driving to the final target posture, the skeleton is stopped for several frames to obtain the entire animation sequence; 6) completing the movement of the skeleton from the initial posture to the target posture.
[0025] Preferably, the step of obtaining human body model parameters also includes: substituting the two-dimensional human body contour image into a first neural network that has undergone deep learning to perform joint point regression, and obtaining a joint point map, a semantic segmentation map, body bone points and key point information of the target human body; substituting the above-mentioned human body information generated into a second neural network that has undergone deep learning to perform human body posture and body shape parameter regression, and obtaining three-dimensional human body parameters, including three-dimensional human body motion posture parameters and three-dimensional human body shape parameters; the three-dimensional human body model has a mathematical weight relationship between bone points and human body mesh, and the determination of bone points can be associated with the determination of the human body model of the target human body posture.
[0026] Also disclosed is a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, any of the above-mentioned method steps is implemented.
[0027] An electronic device includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement any of the method steps described above when executing the program stored in the memory.
[0028] The beneficial effects of the present invention are:
[0029] 1. The virtual clothing has a high degree of realism. Here, realism actually includes two aspects: one is the high degree of realism of the clothing that naturally follows the changes in the human body state; the other is the high degree of restoration of the clothing fabric texture. We first put the 3D clothing model on a 3D standard human body model, input the body parameters of the target human body model to obtain the target human body model, and use mathematical model processing to make the 3D clothing model follow the changes in body shape from the standard human body model to the target human body model, making the changes in clothing state very realistic. In addition, we also use cloth simulation, which is actually to simulate a cloth effect that is close to reality (of course, it can also be made to be inconsistent with the physical effect in reality). What we mainly highlight here is the high degree of restoration of cloth texture simulation, including the accuracy of the simulation of cloth printing. Especially at the end of the movement, the cloth solution is performed on the clothes for several frames, and the final texture state of the clothes is reproduced more completely, and the authenticity is guaranteed.
[0030] 2. During the fabric simulation process, when the human body model reaches the target posture, a neural network is used to determine whether the input image is static or in motion. For different situations, gravity calculations are performed on the fabric of the garment for several frames to simulate the final state of the garment under real motion. This ensures that the fabric of the garment exhibits a movement trend consistent with the motion form in the photo under the target posture, achieving excellent realism and texture.
[0031] 3. Balancing quality and speed. While ensuring that the cloth simulation effect exceeds that of ordinary models, a skeleton-driven skinned mesh is used to drive the model movement in some less important areas where deformation is not very severe. The clothing model is driven by a combination of cloth simulation and skinning. Skinning is used to ensure speed in areas where the clothing barely deforms after movement, such as the upper body. Cloth simulation is used to ensure realism and texture quality in areas where the clothing deforms during movement, such as the legs and hem.
[0032] 4. Simple user operation. The present invention provides a method for analyzing full-body photos of the human body through deep neural networks to obtain accurate three-dimensional human body model parameters. Only an ordinary photo is needed to quickly model the human body. At the same time, the three-dimensional clothing model is obtained by processing the two-dimensional clothing pictures in advance. The user does not need to participate in these behind-the-scenes work. He only needs to select the clothing style he wants to try on virtually, and the system will automatically match the corresponding clothing model. It adapts very well to the characteristics and trends of the Internet era. First, it is simple, and second, it is fast. The user does not need to prepare anything. Uploading a photo is all the work the user needs to do.
[0033] 5. Reasonable skeletal driving process. To realistically fit the target human pose, we employ an optimized interpolation method to achieve the transition from the initial pose to the target pose. Compared to traditional interpolation methods, our target pose skeletal information is predicted by model regression. Simultaneously, an animation sequence is generated from the initial pose to the target pose. Interpolation methods such as linear interpolation and nearest neighbor interpolation are used to form a time series of skeletal information from the initial pose to the target pose. During the generation of the animation sequence, mesh interpolation is used. The interpolation speed is set to slow the distance between the initial and target points, while the intermediate movement is fast. Crucially, the model is paused for several frames when it reaches the final target pose, thus generating the entire animation sequence. This approach more closely reflects the laws of motion in the real world than uniform interpolation. The pose of the human model is stabilized, and the state of the garment model is also stabilized within these few frames, allowing for simultaneous cloth simulation. This results in better garment and human pose simulation and significantly reduces processing time.
[0034] 6. High-frequency use of deep neural networks. The present invention fully utilizes the advantages of deep learning networks and can restore the posture and shape of the human body with high precision in various complex scenes. Different neural networks are used for different purposes, and neural network models with different input conditions and training methods are used to achieve accurate contour separation of the human body in a complex background, semantic segmentation of the human body, determination of key points and joints, eliminating the influence of loose clothing and hairstyle, and achieving the greatest possible approximation to the real body shape and form of the human body. Neural network models are also used in the existing technology, but due to differences in input conditions, input parameters, and training methods, the functions and effects of neural network models vary greatly.
[0035] 7. The setting of the neural network model is more scientific and targeted. Some image processing methods in the existing technology are too focused on simply outputting the model directly, without spending time to polish the details of the model. They simply complete the mapping from 2D pictures to 3D body models through training with massive image data. Although the efficiency is very high, the processing flow is too simple and relies entirely on the neural network model to generate a three-dimensional human body model. The consistency and effect of the body proportions and details are unsatisfactory, and it is completely unhelpful for subsequent processing, and may become an obstacle that subsequent programs are difficult to overcome. In our first-level neural network, human body contours, human semantic segmentation, key points and joints are all used as input items. The model parameters can be generated from multiple angles, and the parameters output by the second-level neural network include two categories: posture and shape, which can control movement and body shape respectively. Combined with our reference model, the posture and body shape of the human body model can be accurately replicated.
[0036] 8. The human body model is precise and controllable. Currently, the popular single-image-based human body reconstruction methods are mainly divided into reconstructing parametric human body models. The most commonly used parametric model is the Max Planck Institute's SMPL model, which contains two sets of 72 parameters for describing human body posture and body shape. For the single-image reconstruction problem, the two-dimensional joint positions are first estimated from the image, and then the SMPL parameters are optimized by minimizing the projection distance between the three-dimensional joints and the two-dimensional plane joints to obtain the human body. However, the SMPL model is mainly learned and trained through a large number of human body model instances. The relationship between body shape and shape basis is a holistic association relationship, which is very difficult to decouple. It is impossible to control the desired body part at will, resulting in the generated model not being able to achieve a high degree of consistency with the real human body posture and body shape. In addition, if it is further applied to the subsequent dressing process, it will also lead to limited ability to express the geometric details of the human body surface, and it will not be able to reconstruct the detailed texture of the clothing on the human body surface. However, our human body model isn't trained. Instead, its parameters have mathematically defined correspondences. In other words, our individual parameter sets are independent of each other, making our model more interpretable during transformations and better able to characterize the shape changes of specific parts of the body. Simply put, human body shapes vary greatly, and many people's thigh-to-calf ratios don't meet a precise set. Our model can manipulate the input parameters to independently control and adjust the length of the thigh and calf, achieving precise leg proportions.
[0037] The present invention uses a series of methods, such as combining skin and cloth simulation, performing gravity calculation for several frames according to the situation at the end of movement, variable speed interpolation to drive model movement, etc., and in the matching of clothing models and human body models, it not only maintains the realism and restoration of the three-dimensional clothing model, but also ensures a certain processing speed, thereby achieving an excellent simulated dressing effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0039] Figure 1 is a flowchart of the entire process of one embodiment;
[0040] Figure 2 A processing flow chart of a model parameter acquisition module according to an embodiment;
[0041] Figure 3A flowchart of a human body model fitting process according to an embodiment;
[0042] Figure 4 Schematic diagram of the system of the present invention. DETAILED DESCRIPTION
[0043] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In order to make the objects, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present invention by illustrating examples of the present invention.
[0044] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0045] The method for processing a human body image provided by an embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0046] like Figure 1 As shown, the present invention provides a clothing model driving method, which includes: obtaining a two-dimensional image of clothing; making a three-dimensional model of clothing according to the two-dimensional image of clothing; constructing a three-dimensional standard human body model in combination with a mathematical model, wherein the three-dimensional standard model is in an initial posture; fitting the three-dimensional clothing model to the three-dimensional standard human body model in the initial posture; obtaining a two-dimensional image of a target human body; obtaining three-dimensional target human body model parameters through calculation of a secondary neural network model; inputting several sets of obtained posture and body shape parameters into the three-dimensional standard human body model for fitting; driving the human body model to move from the initial posture to the target posture; obtaining a target human body model with the same posture and body shape as the target human body and wearing the changed clothing.
[0047] The method generally includes several steps: first, generating a 3D clothing model; second, generating a standard human body model and fitting the 3D clothing model onto the standard human body model; third, obtaining parameters of a target human body model; and fourth, aligning the standard human body model's body shape and posture with the target human body model, while simulating the realistic changes that occur in the 3D clothing model as the human body model changes.
[0048] The first part primarily involves generating a 3D clothing model. Several different methods exist in the existing art for generating 3D clothing models. Currently, the more traditional method for creating 3D clothing models is based on the design and stitching of 2D clothing pieces. This method requires specialized clothing knowledge to design the pattern. Another relatively novel 3D modeling method is based on hand-drawing, which can generate a simple clothing model using user-drawn line information. Another method uses image processing and graphics simulation techniques based on garment image information to ultimately generate a virtual 3D clothing model. Contour detection and classification are used to obtain the garment's outline and dimensions from the image. Machine learning methods are used to identify key points between edges within the outline. Stitching information is generated based on the key point correspondences. Finally, the garment is physically stitched together in 3D space to simulate the garment's appearance on the human body. Other methods include mapping and mathematical model simulation. This method is not specifically limited to this aspect of the present invention. However, the 3D clothing model must be matched to a standard human body model. The general requirement is that a garment model that has already been adapted to a standard human body model is adapted to the human body model in the target pose through fabric physics simulation, ensuring the garment's naturalness and appropriateness.
[0049] Therefore, some basic requirements must usually be met, including but not limited to the following: a. It must completely fit the initial posture of the standard mannequin without any penetration through the mold; b. The output must be a uniform quadrilateral; c. The UV of the model must be unfolded, flattened and compactly aligned, and the texture must be manually aligned with the UV using the Photoshop tool; d. Vertex merging must have been performed; e. The output model should be uniformly reduced in face count, with the total face count not exceeding 150,000 faces / set as a reference standard; f. The material must be adjusted in mainstream fashion design software, and a 10-frame animation must be solved to observe the fabric effect, achieving the desired effect and saving the material parameters; g. The rendering material must be adjusted in mainstream design software, and a rendering preview must be taken to ensure that the material Lambert properties are reasonable.
[0050] The second part involves pre-designing and modeling some standard mannequins, or basic mannequins, based on our human modeling methodology. The 3D clothing model is then placed on the standard mannequins to achieve compatibility with our subsequent workflow. The main work involves combining mathematical models to construct a 3D standard mannequin, or basic mannequin. The Max Planck Institute's SMPL mannequin avoids surface distortion during human motion and accurately depicts the morphology of muscle extension and contraction. In this method, β and θ are the input parameters. β represents 10 parameters related to a person's height, weight, head-to-body ratio, and other proportions, while θ represents 75 parameters representing the overall motion pose and the relative angles of 24 joints. The β parameter is a ShapeBlend pose parameter that can be used to control the shape of the human body using 10 incremental templates. Specifically, the changes in the shape of each parameter can be captured through animated graphics. By studying the continuous animation of the parameter changes, we can clearly see that each continuous change in the parameter controlling the human shape leads to cascading changes in the model, both locally and globally. To reflect the movement of human muscle tissue, linear changes in each parameter of the SMPL mannequin result in large-scale mesh changes. To illustrate this, for example, when adjusting the parameters of β1, the model will directly interpret these changes as changes to the entire body. You might only want to adjust the waist proportions, but the model will force adjustments to the legs, chest, and even the weight of the hands. While this working mode greatly simplifies the workflow and improves efficiency, it is indeed very inconvenient for projects that focus on modeling quality. This is because the SMPL human body model is ultimately trained using Western body photos and measurements to conform to Western body shapes. Its body shape changes generally conform to the typical curve of Western people. Applying this model to Asian human bodies can lead to many problems, such as arm-leg proportions, waist-to-body ratios, neck proportions, and leg and arm lengths. Our research has shown that there are significant discrepancies in these aspects. If the SMPL human body model is rigidly applied, the final generated results will not meet our requirements.
[0051] To achieve this, we employed a custom-built human model approach to achieve enhanced performance. The core of this approach is the creation of a custom blendshape base to enable precise, independent manipulation of the human body. Preferably, the three-dimensional standard human model (basic mannequin) consists of 20 body base parameters and 170 skeletal parameters. These bases comprise the entire human model, with each base independently controlled by its own parameters, without interfering with each other. Precise manipulation, on the one hand, involves increasing the number of control parameters, rather than relying on the Max Planck Institute's ten β control parameters. This allows for adjustable parameters beyond the standard body shape, including arm length, leg length, waist, hip, and chest shape. This more than doubles the number of skeletal parameters, significantly expanding the range of adjustable parameters and providing a solid foundation for the refined design of standard human models. Independent manipulation means that each base, such as the waist, legs, hands, and head, can be manipulated independently, and each bone can be independently adjusted in length, independently of each other and without any physical interaction. This allows for precise adjustments to the human model. The model no longer appears "cumbersome and clumsy," unable to be adjusted to the designer's satisfaction. Our current model embodies a mathematically based correspondence. This is essentially a combination of human aesthetics and statistical data analysis, designed to produce a model that we believe is accurate for the Asian body shape according to our design rules. This significantly differs from the SMPL body model trained on big data. Therefore, our parameter transformations are more interpretable and can better represent local shape changes in the body model. Furthermore, these changes are based on mathematical principles, ensuring that no parameters influence each other, and the arms and legs remain completely independent. The reason for designing so many different parameters is to avoid the drawbacks of body models trained on big data. This allows for precise control of the body model in multiple dimensions, beyond just a few parameters like height, significantly improving modeling performance. Setting so many independent control parameters is only meaningful when building a custom body base; both are essential to achieving designer-level performance.
[0052] As for putting the three-dimensional clothing model on the standard human body model, it is a conventional technology in this field. The present invention does not make too many restrictions on this, as long as the required effect can be achieved.
[0053] The third part is to process the acquired human body images to obtain the parameter information required to generate the human body model. In the past, the selection of these skeletal key points was usually done manually, but this method is very inefficient and does not adapt to the fast-paced requirements of the Internet era. Therefore, in today's era when neural networks are popular, it has become a trend to use deep learning neural networks instead of manual key point selection. However, how to use neural networks efficiently is a problem that requires further research. Generally speaking, we adopted the idea of a secondary neural network plus data "fine-tuning" to construct our parameter acquisition system. Figure 2 As shown in the figure, we use a deep learning neural network to generate these parameters, which mainly includes the following sub-steps: 1) obtain a two-dimensional image of the target human body; 2) process and obtain a two-dimensional human body contour image of the target human body; 3) substitute the two-dimensional human body contour image into the first deep learning neural network for joint point regression; 4) obtain a joint point map of the target human body; obtain semantic segmentation maps of various parts of the human body; body key points; body bone points; 5) substitute the generated joint point map, semantic segmentation map, body bone points and key point information of the target human body into the second deep learning neural network for human posture and body shape parameter regression; 6) obtain the output three-dimensional human body parameters, including three-dimensional human body motion posture parameters and three-dimensional human body shape parameters.
[0054] The two-dimensional image of the target human body can be a two-dimensional image including a human body in any posture and any clothing. The acquisition of the two-dimensional human body contour image utilizes a target detection algorithm, which is a target region rapid generation network based on a convolutional neural network.
[0055] Before inputting the two-dimensional human image into the first neural network model, a neural network training process is also included. The training sample includes a standard two-dimensional human image with the original joint point locations manually annotated with high accuracy on the two-dimensional human image. Here, a target image is first acquired and human body detection is performed on the target image using a target detection algorithm. Human body detection does not mean using a measuring instrument to detect a real human body. In this invention, it refers to a given image, typically a two-dimensional photograph, that contains sufficient information, such as a face, limbs, and body. A specific strategy is then used to search the given image to determine whether a human body is present. If a human body is present, parameters such as the human body's location and size are determined. In this embodiment, before obtaining the human body key points in the target image, human body detection is performed on the target image to obtain a human body frame that annotates the human body's location. Because the input image can be any image, some non-human background elements, such as tables, chairs, trees, cars, and buildings, are inevitably present. This unused background is removed using a sophisticated algorithm.
[0056] At the same time, we also need to perform semantic segmentation, joint detection, skeleton detection, and edge detection. By collecting these 1D point information and 2D surface information, we can lay a good foundation for generating a 3D human body model later. A first-level neural network is used to generate a human body joint map. Optionally, a target detection algorithm can quickly generate a network for the target area based on a convolutional neural network. This first neural network requires a large amount of data training. The joints of some photos collected from the Internet are manually annotated and then input into the neural network for training. After deep learning, the neural network can basically obtain a joint map with the same accuracy and effect as manually annotated joints immediately after inputting the photo, and the efficiency is dozens or even hundreds of times that of manual annotation. Human joints usually exist as key points of the human body.
[0057] In the present invention, obtaining the joint positions of the human body in the photo is only the first step. After obtaining 1D point information, 2D surface information must be generated based on this 1D point information. All of this can be accomplished using a neural network model and mature algorithms in the prior art. By redesigning the process and timing of the neural network model's involvement and rationally designing various conditions and parameters, the present invention makes parameter generation more efficient and reduces the degree of human involvement. This makes it very suitable for Internet application scenarios. For example, in a virtual dress-up program, users can obtain the dress-up results almost instantly without waiting, which plays a vital role in increasing the program's appeal to users.
[0058] After obtaining the relevant 1D point information and 2D surface information, these parameters or results, the target human body's joint point map, semantic segmentation map, body bone points and key point information can be substituted as input items into the second neural network that has undergone deep learning to regress human body posture and body shape parameters. After the regression calculation of the second neural network, several groups of three-dimensional human body parameters can be immediately output, including three-dimensional human body action posture parameters and three-dimensional human body shape parameters. Preferably, the loss function of the neural network is designed based on the three-dimensional standard human body model (basic mannequin), the predicted three-dimensional human body model, the standard two-dimensional human body image with the original joint point positions marked, and the standard two-dimensional human body image including the predicted joint point positions.
[0059] The fourth and most critical part is to fit the parameters of the human body model to the human body model and drive it, while ensuring that the state of the clothes after moving is as realistic as possible.
[0060] like Figure 3As shown, the movement process includes the following substeps: First, the obtained 3D human pose and SHAPE parameters are matched to several bases and skeletal parameters of a 3D standard human model; then, the obtained sets of bases and skeletal parameters are input into a standard 3D human parameter model for fitting; the 3D human model has a mathematical weight relationship between skeletal points and the model mesh; the determination of skeletal points can be correlated to determine the human model of the target human pose. In this step, the two parameters generated in the previous step are substituted into the pre-designed human model to construct the 3D human model. These two types of parameters are similar in name to the parameters of the Max Planck Institute's SMPL human model, but their actual content differs significantly. This is because the two have different foundations. Specifically, the present invention uses a self-made 3D standard human model (basic mannequin), each base of which is designed based on the body shape and proportions of an Asian, including several areas not covered by the SMPL model. The Max Planck Institute's SMPL model uses a standard human model generated through big data training. The two models are generated using different calculation methods. Although both are ultimately reflected in the generated 3D human model, their content differs significantly. After this step, a preliminary 3D human body model will be obtained, including the mesh of the human body model with bone position and length information.
[0061] In this part, after fitting the 3D clothing model onto a standard human body model, we need to align the standard human body model's shape and posture with the target human body, while also simulating the realistic changes that the 3D clothing model undergoes as the human body model changes. We use several methods to ensure this is achieved.
[0062] First, the target human body posture is determined by the 3D human body motion posture parameters, and the skeleton is driven to move from the initial posture to the target posture, so that the 3D human body model is basically consistent with the target human body posture.
[0063] In the animation field, using bones or joints as the driving source is a common approach, often involving skinning. Skinning establishes relationships between model mesh vertices and bones during the driving process. After calculating the bones' rotational parameters at the target pose, the bones are then used to drive nearby vertices through spatial movement, driving these vertices to the target pose. For example, when a bone bends and rotates, it also drives the movement of nearby bound mesh vertices. Similarly, thanks to our custom-built human body model, which possesses numerous parameters representing human characteristics, we can accurately depict every detail of the human body. For example, after obtaining the rotational parameters for all 170 bones at the target pose, we can use all the skinning weights to drive the entire human body mesh vertices to the target pose. The advantage of skinning is its high speed, meaning that interpolation calculations are unnecessary for each frame, allowing the cloth model to be driven directly from its initial state to the target pose. However, its drawback is also significant: without any physics-based simulation, the transformations can be unnatural, and the state after reaching the target position differs from the actual state after the movement. For tight clothing models, the penetration is not obvious, but for loose three-dimensional clothing models, it is obviously unnatural.
[0064] To address these issues, we utilize a combination of cloth simulation and skinning for the movement of clothing models. Skinning is used for areas that experience minimal or minimal deformation after movement, while cloth simulation is used for areas that experience significant deformation during movement. Skinning primarily improves movement processing speed, while cloth simulation enhances realism. Cloth simulation essentially simulates cloth effects that are close to reality (though it can also differ from physical reality). We primarily utilize the high fidelity of cloth texture simulation, including the accuracy of fabric print simulation. Our approach balances quality and speed. While ensuring cloth simulation performance surpasses that of standard models, we utilize a skeleton-driven skinned mesh to drive model movement in less critical areas where deformation is less drastic. This combination of cloth simulation and skinning is used to drive the model. Skinning is used to ensure faster processing speed for areas that experience minimal deformation after movement, such as the upper body of a dress, yoga attire, tight-fitting clothing, and clothing where shape and posture closely match. Fabric simulation is used to ensure realism and texture fidelity for parts of clothing that deform during movement, such as legs and hems, windbreakers and coats with hems, and gauze garments. There's no clear-cut distinction between which parts use skinning and which use fabric simulation; the final garment model determines the final result. For example, if the upper body of a dress has features like water sleeves or lotus leaf sleeves that change significantly with arm movement, fabric simulation should still be used. This invention emphasizes the ability to combine these two methods to find the optimal balance between speed and quality.
[0065] Secondly, the movement state from the initial pose to the target pose is calculated frame by frame. When calculating each frame, the skin drive is completed first and then the cloth simulation is calculated. After the current frame is calculated, the state of the next frame is calculated. This actually forms a sequence of interpolated frames, maintaining the high realism brought by frame-by-frame movement and frame-by-frame calculation. Cloth simulation uses a collision body system, modeling the vertex mesh of the human model as a rigid body and the mesh of the clothing model as a non-rigid body. The physics engine is used to simulate the collision relationship between rigid and non-rigid bodies. During the cloth simulation process, the collision between rigid and non-rigid bodies is calculated, and the connection force between the clothing model's own mesh is considered. The mesh state of the clothing model is calculated frame by frame to simulate the collision movement process in the physical world. This frame-by-frame simulation method can maintain the real state of the clothing model's movement to the greatest extent. Although the simulation speed has been reduced, it has maintained a relatively good balance against the background of enhanced device computing power.
[0066] Thirdly, the driving method of the present invention has another important feature. When driving the human body model to reach the target posture, it is necessary to perform several frames of gravity solution on the fabric of the clothing. The continuity of the movement change is judged according to the neural network prediction model, and the target posture is analyzed to see whether it is a snapshot or a static pose. If it is a snapshot, the clothing model will maintain the speed during the movement due to inertia, making the target posture state an unsteady state. When the human body model reaches the target posture, the number of gravity solution is reduced; if it is a posed shot, the number of gravity solution is increased when the human body model reaches the target posture to ensure the realism of the clothing fabric in the target posture. In particular, at the end of the movement, several frames of fabric solution are performed on the clothing, which is perfectly coordinated with the variable speed interpolation drive in terms of time. This process originally exists and is originally intended to obtain a stable model state. At the same time, it is applied to the clothing model, using a parallel method, with basically no time loss, and the final texture state of the clothing is reproduced more completely, and the authenticity is improved. During the fabric simulation process, when the human body model reaches the target posture, a neural network is used to determine whether the input image is static or in motion. For different situations, the gravity of the fabric of the clothing is solved for several frames to simulate the final state of the clothing under real motion conditions. This ensures that the clothing fabric presents a movement trend consistent with the motion form in the photo under the target posture, achieving high realism and texture.
[0067] In this part, the human body model must complete the transformation from the initial pose to the target pose. Because we only input a photo, the target human body pose in the photo is usually different from the basic human body. At this time, in order to fit the target human body pose, it is necessary to complete the transformation from the initial pose to the target pose. In order to simulate more realistically, in the fitting step, when fitting several sets of basis and bone parameters in the standard 3D human body parameter model, the following steps are also included:
[0068] 1) Obtain the position coordinates of the initial pose and the target pose; the initial pose parameters are determined by the initialization parameters of the standard mannequin, and the skeletal information of the target pose is obtained by regression prediction of the neural network model.
[0069] 2) Generate an animation sequence that moves from the initial posture to the target posture; after obtaining the initial state of the skeleton information and the target posture state parameters, form a skeleton information time series from the initial posture to the target posture through linear interpolation, nearest neighbor interpolation and other interpolation methods. During the driving process, according to the number of bones driven per frame, it can be divided into two methods: global linear interpolation and driving the parent node first and then the child node. Considering the driving state in the simulated physical world, this patent adopts the latter method, driving the parent skeleton node first and then the child skeleton node. In this way, the interpolated animation sequence action is more in line with the real physical world, and the simulation effect is better.
[0070] 3) In the process of generating animation sequences, mesh interpolation is used for processing; that is, after each frame drives the skeleton movement, the current state of the human body model vertex or surface information is calculated through the weight parameters of the standard mannequin, and the current human body model mesh state is updated and recorded and saved.
[0071] 4) The interpolation speed is set to be slow before and after the initial and target points, and fast during the intermediate movement. This patent uses a non-uniform interpolation rate, meaning that the individual frame movement amplitude is small during the initial and final stages of the movement, and large during the intermediate stages. This simulates the acceleration at the start of a real-world physical action, while maintaining a high inter-frame displacement distance during the movement, and then reducing the driving speed at the end of the movement.
[0072] 5) When driving to the final target pose, pause for a few frames to obtain the entire animation sequence. This approach is closer to the real-world motion laws than uniform interpolation, and the simulation effect is better.
[0073] 6) Complete the movement of the skeleton from the initial posture to the target posture.
[0074] Combine Figures 1 to 3 The method for generating a three-dimensional human body model according to the embodiment of the present invention may be implemented by a human body image processing device. Figure 4 FIG. 3 is a schematic diagram showing a hardware structure 300 of a device for processing a human body image according to an embodiment of the present invention.
[0075] The present invention also discloses a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the aforementioned clothing model driving method and steps are implemented.
[0076] And an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement the aforementioned clothing model driving method and steps when executing the program stored in the memory.
[0077] like Figure 4 As shown, the device 300 for implementing virtual fitting in this embodiment includes: a processor 301, a memory 302, a communication interface 303 and a bus 310, wherein the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and communicate with each other.
[0078] Specifically, the processor 301 may include a central processing unit (CPU) or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.
[0079] Memory 302 can include a large capacity memory for data or instructions. For example, but not limitation, memory 302 can include HDD, floppy disk drive, flash memory, CD, magneto-optical disk, magnetic tape or universal serial bus (USB) drive or two or more of these combinations. In appropriate cases, memory 302 can include removable or non-removable (or fixed) media. In appropriate cases, memory 302 can be inside or outside the processing human body image device 300. In a specific embodiment, memory 302 is a non-volatile solid-state memory. In a specific embodiment, memory 302 includes a read-only memory (ROM). In appropriate cases, the ROM can be a mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically rewritable ROM (EAROM) or flash memory or two or more of these combinations.
[0080] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiment of the present invention.
[0081] Bus 310 comprises hardware, software or both, and the parts of the equipment 300 of processing human body image are coupled to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more above these combinations.In suitable case, bus 310 can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.
[0082] That is to say, Figure 4 The device 300 for processing human body images shown can be implemented as including: a processor 301, a memory 302, a communication interface 303 and a bus 310. The processor 301, the memory 302 and the communication interface 303 are connected via the bus 310 and communicate with each other. The memory 302 is used to store program code; the processor 301 reads the executable program code stored in the memory 302 to run the program corresponding to the executable program code, so as to execute the virtual fitting method in any embodiment of the present invention, thereby realizing the combination of Figures 1 to 3 The invention describes a method and apparatus for virtual fitting.
[0083] An embodiment of the present invention further provides a computer storage medium having computer program instructions stored thereon; when the computer program instructions are executed by a processor, the method for processing a human body image provided by an embodiment of the present invention is implemented.
[0084] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0085] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0086] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.
[0087] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.
Claims
1. A clothing model driving method, characterized in that: The method comprises: 1) Obtain a two-dimensional image of the garment; 2) Create a three-dimensional model of the garment based on the two-dimensional image of the garment; 3) constructing a three-dimensional standard human body model in combination with the mathematical model, wherein the three-dimensional standard human body model is in an initial posture, the three-dimensional standard human body model having a mathematical weight relationship between skeletal points and a human body mesh, and the determination of the skeletal points can be associated with determining a human body model of a target human body posture; 4) Fitting the 3D clothing model to the 3D standard human body model in the initial pose through fabric physics simulation; 5) Using target detection algorithm to obtain a two-dimensional image of the target human body; 6) Obtaining the three-dimensional target human body model parameters, including posture and body shape parameters, through calculation of the secondary neural network model; 7) Inputting the obtained sets of posture and body shape parameters into the 3D standard human body model for fitting, so that the posture and body shape of the 3D standard human body model are consistent with those of the 3D target human body model; 8) Using a combination of cloth simulation and skinning methods to drive the clothing model, the human body model moves from the initial pose to the target pose; 9) Obtain a target human body model having the same posture and body shape as the target human body and wearing the changed clothing; wherein, when driving the clothing model so that the posture of the human body model reaches the target posture, it is necessary to perform gravity calculations on the clothing fabric for several frames, judge the continuity of the movement change based on the neural network prediction model, and analyze whether the target posture is a snapshot or a static pose. If it is a snapshot, the clothing model maintains the speed during the motion due to inertia, making the target posture state an unsteady state. When the human body model reaches the target posture, the number of gravity calculations is reduced; if it is a posed shot, the number of gravity calculations is increased when the human body model reaches the target posture to ensure the realism of the clothing fabric in the target posture.
2. The method according to claim 1, characterized in that The target human body posture is determined by the 3D human body motion posture parameters, and the skeleton is driven to move from the initial posture to the target posture, so that the 3D human body model is basically consistent with the target human body posture.
3. The method according to claim 1, characterized in that The driving method of the clothing model adopts a combination of cloth simulation and skinning methods. The skinning method is used for the part of the clothing that basically does not deform after the human body model moves, and the cloth simulation method is used for the part that deforms greatly during the movement of the clothing.
4. The method according to claim 3, characterized in that During the driving process, the relationship between the model mesh vertices and bones is established. After calculating the rotation parameters of the bones at the target posture, the skin weight relationship is used to enable the bones to drive the nearby vertices to complete the spatial movement and drive the associated vertices to the position of the target posture.
5. The method according to claim 3, characterized in that The movement state from the initial posture to the target posture is calculated frame by frame. When calculating each frame, the skin drive is completed first and then the cloth simulation is calculated. After the current frame is calculated, the state of the next frame is calculated.
6. The method according to claim 3, characterized in that Cloth simulation uses a collision body system, modeling the vertex mesh of the human model as a rigid body and the mesh of the clothing model as a non-rigid body. The physics engine is used to simulate the collision relationship between rigid and non-rigid bodies. During the cloth simulation process, the collision between rigid and non-rigid bodies is calculated. At the same time, the connection force between the meshes of the clothing model itself is considered, and the mesh state of the clothing model is calculated frame by frame to simulate the collision motion process in the physical world.
7. The method according to claim 1, characterized in that The fitting movement step further includes the following sub-steps: 1) Obtain the position coordinates of the initial posture and target posture; 2) Generate an animation sequence that moves from the initial pose to the target pose; 3) In the process of generating animation sequences, grid interpolation is used for processing; 4) The interpolation speed is set so that the distance between the initial point and the target point is slow, and the intermediate movement process is fast; 5) When driving to the final target pose, stop for several frames to obtain the entire animation sequence; 6) Complete the movement of the skeleton from the initial posture to the target posture.
8. The method according to claim 1, characterized in that The step of obtaining the parameters of the three-dimensional target human body model also includes substituting the two-dimensional human body contour image into a first neural network that has undergone deep learning to perform joint point regression, thereby obtaining a joint point map, a semantic segmentation map, body bone points, and key point information of the target human body; substituting the generated above-mentioned human body information into a second neural network that has undergone deep learning to perform human body posture and body shape parameter regression, thereby obtaining three-dimensional human body parameters, including three-dimensional human body motion posture parameters and three-dimensional human body shape parameters.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
10. An electronic device, characterized in that: The invention comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement any method described in claims 1-8 when executing the program stored in the memory.
Citation Information
Patent Citations
Skeleton-based rapid garment fitting method
CN108537888A
Video-based attitude data capture method and system
CN109145788A
Video human body three-dimensional reconstruction method and device based on clothing modeling and simulation
CN110309554A