A segmented driving method, device and storage medium for a human body model
Through the combined with the segmented driving method and the deep learning neural network of the homemade basic human station, the problem of slow modeling speed and low accuracy of the three-dimensional human body model in the existing technology is solved, and more efficient and accurate three-dimensional human body model driving is achieved, which improves the effect of virtual fitting.
Patent Information
- Application Number
- CN202011609644.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-12-28
AI Technical Summary
When generating a three-dimensional human body model, the prior art has problems such as slow modeling speed, insufficient accuracy and large calculations. Especially in the driving process from the initial posture to the target posture, it is easy to produce mold penetration and the inability to achieve independent control of each part, resulting in poor virtual fitting effect.
The segmented driving method is adopted to obtain the secondary information of the mannequin model through self-made basic human stations and deep learning neural networks, and drive the mannequin model grid using interpolation technology and non-uniform interpolation frame rate, thereby realizing the generation of animation sequences from the initial pose to the target pose, and independently control it with the parameter relationship of mathematical principles.
The accuracy and speed of model fitting is improved, the probability of wearing the model is reduced, and the generated three-dimensional human body model is closer to the actual human body state, improving the authenticity and efficiency of virtual fittings.
Smart Images

Figure CN114758039B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of human body three-dimensional model modeling, and particularly relates to a method for segmentally driving a human body model, in particular to a method, device and storage medium for separately driving different parts of a basic mannequin. Background Art
[0002] With the development of Internet technology, online shopping has become increasingly popular. Compared with shopping in physical stores, online shopping has advantages such as a wide variety of goods and convenient shopping. However, when purchasing goods online, there are also some problems that are not easy to solve. The most important one is that it is impossible to view the goods to be purchased on the spot. Among all types of goods, this problem is most prominent in clothing. Compared with being able to try on clothes in real time in physical store shopping to view the clothing effect, online clothing shopping cannot provide an effect picture for the consumer himself / herself, and can only provide pictures of models trying on clothes. Some don't even have try-on pictures at all, and consumers cannot intuitively obtain the matching degree between the clothing and their own body shape in real time. This has caused a large number of returns and exchanges.
[0003] To solve this problem, operators have tried to use virtual fitting technology to provide simulated fitting effects for consumers. Of course, there are also other occasions in reality where virtual clothing changing and fitting technology can be used, such as in online games. Therefore, this technology has developed relatively rapidly.
[0004] Virtual fitting refers to a technical application in which users can view the "fitting" effect in real time on the terminal screen without actually putting on the clothes they want to see the wearing effect of. Existing fitting technology applications mainly include two-dimensional fitting and three-dimensional virtual fitting technology. The former basically collects pictures of users, collects pictures of clothes, and then performs cutting and splicing to form an image after "putting on clothes". However, such images have poor authenticity due to simple and crude image processing methods, and completely do not consider the actual body shape of users. It just mechanically applies the clothing to the user's photo and cannot meet the needs of users. The latter usually collects three-dimensional information of a person through a three-dimensional acquisition device and combines it with the characteristics of the clothing, or manually inputs the body data information provided by the user, and virtually generates a three-dimensional human body model according to certain rules, and then combines it with the clothing texture. Generally speaking, such three-dimensional virtual fitting requires a large amount of data collection or three-dimensional data calculation, and the hardware cost is high, making it not easy to promote among ordinary users.
[0005] With the development of cloud computing technology, artificial intelligence technology, and the processing power of intelligent terminals, technologies have emerged to generate a three-dimensional human model from a two-dimensional virtual human image and achieve virtual fitting after dressing the three-dimensional clothing model. Such technologies mainly include three steps: (1) processing the personal body information provided by the user to obtain a target human model; (2) processing the clothing information to obtain a clothing model; (3) fusing the human model and the clothing model together to generate a simulation diagram of a person wearing the clothing.
[0006] However, due to the accumulation of many uncertain factors such as process design, model parameter selection, and the training method of neural networks, the quality of the final generated fitting pictures is not as good as that of traditional three-dimensional virtual fitting technologies. Among them, the fitting of the human model is the basic step, and the subsequent dressing process must also be based on the previously generated human model. Therefore, once the generated human model is inaccurate, problems such as a large gap between the human model and the fitter's body shape, loss of skin texture, and loss of body parts are likely to occur; in addition, since the human model needs to be driven from the standard pose to the target pose, this involves the dual effects of processing speed and generated quality, thus affecting the effect of the final generated fitting model.
[0007] In the general field of computer vision, there are many initial steps for human body modeling, usually including three categories: using a 3D scanning device to perform a full-range scan of a real human body, a three-dimensional reconstruction method based on multi-view depth-of-field photography, and a method of achieving three-dimensional reconstruction by combining a single or multiple images with a neural network model and a standard human body model.
[0008] Starting from the characteristics of the Internet ecosystem, directly outputting the final completed dressed human model from a single image is undoubtedly the best choice. Users do not need to be present on-site. With just one photo, they can complete the entire fitting process. Among them, how to obtain a human model that is closest to the real human pose from a photo with various poses becomes the top priority.
[0009] In the prior art, there are usually several categories of methods for constructing a human body model: (1) Regression-based methods, which reconstruct a voxel-based human body model through a convolutional neural network. The algorithm first estimates the positions of the main body joints based on the input image, and then estimates the positions within a specified-size voxel grid according to the key point positions. According to whether each unit voxel inside is occupied, the entire shape of the occupied voxels inside is used to describe the reconstructed human body shape. (2) Single-image-based human body reconstruction, which simultaneously estimates the three-dimensional shape and pose of the human body. This method first roughly annotates simple human body bone key points on the image, and then performs initial matching and fitting of the human body model based on these rough key points to obtain a general human body shape. (3) Represent the human body skeleton with 23 bone nodes, and then use the rotation of each bone node to represent the pose of the entire human body. At the same time, express the human body shape with 6890 vertex positions. During the fitting process, given the bone node positions, the parameters of the shape and pose are simultaneously fitted to perform three-dimensional human body reconstruction; or first use a CNN model to predict the key points on the image, and then use the SMPL model for fitting to obtain an initial human body model. Then, use the shape parameters obtained by fitting to regress a human body joint bounding box. Each joint corresponds to a bounding box, and its bounding box is represented by the axis length and radius. Finally, combine the initial model and the bounding box obtained by regression to obtain three-dimensional human body reconstruction. The above methods have the problems of slow modeling speed, insufficient modeling accuracy, and strong dependence of the reconstruction effect on the created body and pose database.
[0010] The prior art one discloses a human body modeling method based on physical fitness test data, including: obtaining physical fitness test data; according to the physical fitness test data, performing linear regression on a pre-created human body model through a pre-trained prediction model to fit and obtain a predicted human body model. The pre-created human body model includes multiple groups of predefined marker feature points and corresponding standard shape bases, and the physical fitness test data includes measurement data corresponding to each group of marker feature points; according to the predicted human body model, obtain a target human body model, and the target human body model includes measurement data, a target shape base, and target shape coefficients. However, this method has very high requirements for physical fitness test data, including body length data and girth data, such as height, arm length, shoulder width, leg length, calf length, thigh length, foot length, head circumference, chest circumference, waist circumference, thigh circumference, etc. It not only requires measurement but also calculation. Although it saves the amount of calculation, the user experience is very poor and the program is very cumbersome. And in the training of the human body model, the training method of the SMPL model is referred to.
[0011] The SMPL model is a parametric human body model, a human body modeling method proposed by the Max Planck Society. This method can perform arbitrary human body modeling and animation driving. The biggest difference between this method and traditional LBS lies in the method it proposed for the body surface morphology of human body postures. This method can simulate the protrusions and depressions of human muscles during limb movement. Therefore, it can avoid surface distortion during human movement and accurately depict the morphology of human muscle stretching and contraction movements. In this method, β and θ are the input parameters, where β represents 10 parameters such as the proportions of a person's height, weight, body ratio, and head-to-body ratio, and θ represents 75 parameters of the overall movement pose of the human body and the relative angles of 24 joints. However, in this model generation method, the core is the large accumulation of training data to obtain the relationship between body shapes and shape bases. However, due to the extremely strong correlation between them, it is impossible to independently control each shape base, and it is not easy to perform decoupling operations. For example, there is also a certain correlation between the arms and legs. When the arms move, the legs will theoretically move accordingly. It is very difficult to implement improvements for different body types on the SMPL model. During the model driving process, the characteristics of this model still seriously affect the final driving effect of the model. To move, it moves as a whole, and it is impossible to achieve individual frame-by-frame movement of each part, and it cannot reflect some detailed improvements in the driving method.
[0012] The second prior art discloses a 3D human body modeling method based on a single photo, including: obtaining a photo, parsing the photo, marking the key points of the human body in the photo, and calculating the spatial coordinates of the key points; after obtaining the distances between the bone points in the pre-created basic mannequin and the key points in the photo, aligning the bone points and the key points to generate a basic human body model; obtaining the basic texture map in the pre-created basic mannequin, performing difference calculation between the basic texture map and the skin texture of the face in the photo, and then using the edge channel for fusion to generate basic texture data; generating a 3D human body model based on the basic human body model and the basic texture data. 3D human body modeling is achieved through a single photo, and there is a bone and muscle system support in the model, which can produce expressions and movements. However, in this method, after matching the distances between the key points of the user photo and the key points of the basic mannequin, the distances are adjusted to achieve the pose of the target human body, and subsequently, the final human body model can only be obtained through difference calculation and fusion between the basic texture map and the skin texture in the photo. This method has a simple process and small computational amount, but the accuracy and realism of the generated human models are not very high.
[0013] The prior art three discloses a single-image human body three-dimensional reconstruction system based on grid deformation, including: collecting a single picture and a human body model, as well as the corresponding initial human body three-dimensional model of the human body model, and constructing the human body model database with the initial human body three-dimensional model as the initial data of the convolutional neural network; rendering the initial data to obtain the input picture during the training of the convolutional neural network; using a data flow programming deep learning platform to construct a convolutional neural network to extract the human joint position probability distribution map according to the input picture; performing human body segmentation annotation on the single picture to obtain a human body segmentation annotation map; training the convolutional neural network according to the single picture, the human body segmentation annotation map and the human joint position probability distribution map to obtain a final human body three-dimensional model. This modeling method also inputs a single photo, and at the same time uses model judgment and the parameters output by the neural network at the end. However, there are only joint point parameters, and then the mature body shape of the SMPL model is used to perform detailed adjustment consistent with the target human body posture. Although the calculation amount is reduced, due to fewer input parameters and only being able to complete the adjustment based on the SMPL prediction model, it is very difficult to output a particularly ideal human body model that is highly consistent with the target human body posture. And the driving of the model also relies on the SMPL model, and a more refined driving method and effect cannot be provided.
[0014] It can be seen that during the process of the human body model changing from the initial position to the target posture, the least input information, the least calculation amount, and the best effect will be three basic goals that have always been pursued. We need to find an optimal balance among the three to provide a human body model driving method that can achieve simple, direct, and easily obtainable input information, a calculation amount that does not exceed the bearing capacity of the terminal device, and an effect close to that of professional devices. Summary of the Invention
[0015] Based on the above problems, the present invention provides a segmented driving method, device, and storage medium for a human body model, which can significantly reduce the probability of penetration during model driving, making the finally generated target human body model closer to the actual human body state in terms of posture and authenticity.
[0016] The present invention provides a segmented driving method for a human body model. The method includes: calculating and obtaining secondary information of the human body model using an algorithm; establishing a basic mannequin in an initial T-pose posture, and the acquisition of the initial posture parameters is determined by the initialization parameters of the basic mannequin model; obtaining the posture and body shape parameters of the target human body model by regression prediction of the neural network model according to the secondary information, and the three-dimensional human body posture and body shape parameters correspond to the bones and several basic parameters of the three-dimensional basic mannequin; inputting the obtained several groups of basic and bone parameters into the basic mannequin model for fitting to obtain the target posture and target body shape; generating an animation sequence from the initial posture to the target posture; using a segmented driving method to drive the human body model mesh to move from the initial posture to the target posture; obtaining a three-dimensional target human body model mesh with the same posture as the target human body two-dimensional image.
[0017] Preferably, during the process of generating the animation sequence, the movement of the human body model mesh is carried out in an interpolation manner. After driving the bones to move in each frame, the vertex and face information of the human body model in the current state are calculated through the weight parameters of the standard human model, and the current state of the human body model mesh is updated and recorded.
[0018] Preferably, during the segmented driving process, first move the mesh parts with relatively small position changes during the driving process that are not likely to cause penetration, and then move the mesh parts with relatively large position changes that are likely to cause penetration.
[0019] Preferably, the parts that are not likely to cause penetration include the head, neck, chest, abdomen, and buttocks; the parts that are likely to cause penetration include the limbs and hands.
[0020] Preferably, after obtaining the initial state of the bone information and the target posture state parameters, drive the bones to move from the initial posture to the target posture, and form a time series of bone information from the initial posture to the target posture through an interpolation method such as linear interpolation or nearest neighbor interpolation. During the driving process, first drive the parent bone nodes, and then drive the child bone nodes.
[0021] Preferably, during the process of generating the animation sequence, the interpolation speed adopts a non-uniform interpolation rate, which is set to be slow at the positions far from the initial point and the target point and fast in the middle movement process, that is, the single-frame movement amplitude is small during the initial action and the end action process, and the movement amplitude is large during the middle movement process. When the interpolation reaches near the end of the movement, the driving speed is reduced, and it is stationary for several frames when moving to the final target posture to obtain the entire animation sequence.
[0022] Preferably, the basic mannequin is a three-dimensional basic mannequin constructed in combination with a mathematical model; the three-dimensional basic mannequin has a mathematical weight relationship between skeletal points and model meshes, and the determination of skeletal points can be associated with a human body model for determining the target human body posture; the three-dimensional basic mannequin is defined by a number of body base parameters and a number of skeletal parameters, and the number of body bases constitutes the entire human body model mesh, and each body base is separately controlled and changed by the base parameters without affecting each other.
[0023] Preferably, the steps of obtaining the parameters of the target human body model further include: 1) obtaining a two-dimensional image of the target human body; 2) processing to obtain a two-dimensional human body contour image of the target human body; 3) substituting the two-dimensional human body contour image into a first neural network that has undergone deep learning for joint point regression; 4) obtaining a joint point map of the target human body; obtaining a semantic segmentation map of each part of the human body; body key points; body skeletal points; 5) substituting the generated joint point map, semantic segmentation map, body skeletal points, and key point information of the target human body into a second neural network that has undergone deep learning for human body posture and body type parameter regression; 6) obtaining the output three-dimensional human body parameters, including three-dimensional human body motion posture parameters and three-dimensional human body type parameters.
[0024] In addition, the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps described in any one of the foregoing are implemented.
[0025] An electronic device includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; the processor is used to implement the method steps described in any one of the foregoing when executing the program stored on the memory.
[0026] The beneficial effects of the present invention are:
[0027] 1. The model has a good fitting effect and less penetration. In order to drive the model to completely fit the pose of the target human body, the traditional method usually adopts the way of overall movement during the change from the initial pose to the target pose. This method has a large amount of calculation and is prone to penetration in the final stage. We noticed that during the movement of the human body model, in fact, the grid movement rules of each part of the limb are different. Some grids change significantly, while some grids change little. Correspondingly, some grids move violently, while some grids hardly move. In view of this characteristic, in order to reduce the probability of penetration, we adopted the method of segmentally driving the grids of the human body model, making full use of this objectively existing rule. First, move some model parts with a very small probability of penetration, and place the parts prone to penetration at the back for movement. In this way, different parts can adopt different driving strategies and methods to achieve the purpose of improving the fitting effect. For example, the parts that are not prone to penetration, such as the head, neck, and chest, can be quickly driven to the target position first, and then other parts prone to penetration can be moved at a slower speed. And in the final stage, a collision body test is introduced. Although it increases the amount of calculation to a certain extent and slows down the movement speed, when offset by the previous situation, the speed does not slow down. However, due to the additional verification and inspection, the fitting effect will be improved and the probability of penetration will decrease.
[0028] 2. The human body model is precisely controllable. The segmental driving method involved in the present invention is actually based on our self-made basic mannequin. Because the current popular single-image-based human body reconstruction methods mainly involve reconstructing parameterized human body models, such as the SMPL model. It contains two sets of 72 parameters for describing human poses and body shapes. However, for the SMPL model, it mainly conducts deep learning and training through a large number of human body model instances. The relationship between the body type and the shape basis is an overall correlation relationship, and it is very difficult to decouple, making it impossible to freely control the body parts that one wants to control, resulting in the generated model not being able to achieve a high degree of consistency with the real human pose and body type. In addition, if further applied to the subsequent clothing process, it will also lead to limited ability to represent the geometric details of the human body surface and unable to well reconstruct the detailed texture of the clothing on the human body surface. However, our human body model is not obtained through training, and there is a corresponding relationship based on mathematical principles between the parameters. That is to say, there is no interrelated relationship between our groups of parameters, and they are independent of each other. Therefore, our model is more interpretable during the transformation process and can better represent the shape and position changes of a certain part or specific parts of the body. That is to say, the partial movement of the human limb will not affect the states of other parts.
[0029] 3. Fast model fitting speed. We adopted an optimized frame interpolation method to complete this. Compared with the traditional frame interpolation method, the skeletal information of the target pose is predicted by a neural network model. At the same time, an animation sequence from the initial pose to the target pose is generated. Through frame interpolation methods such as linear interpolation and nearest neighbor interpolation, a time series of skeletal information from the initial pose to the target pose is formed. During the process of generating the animation sequence, the grid mesh interpolation method is used for processing. The frame interpolation speed is set to be slow at the positions of the initial point and the target point, and fast during the intermediate movement process. Especially importantly, when the model reaches the final target pose, it pauses for several frames, so that the model can get an effective buffer before stopping after high-speed movement, and then the entire complete animation sequence is obtained, and the pose fitting accuracy of the human model is higher. This approach is closer to the laws of motion in the real physical world than uniform frame interpolation, and the simulation effect of clothing and human poses is better, and a considerable amount of processing time can be reduced. During the fitting process, there are various different methods to improve the effect of a certain aspect of the fitting process. Sometimes we will use these methods comprehensively, and sometimes we will only use one of them. This is because according to the current statistical situation, in the human body dressing project, in most cases, the gap between the original human pose and the target human pose is not very large. If multiple intervention methods are used, the amount of calculation and processing will become larger, and it is easy to cause disharmony between various methods. Therefore, the positive and negative contribution values brought by various methods can be comprehensively considered to determine the final solution.
[0030] 4. High-frequency creative use of hierarchical deep neural networks. In the prior art, neural network models are also used, but due to different input conditions, input parameters, and training methods, the functions and roles played by the neural network models are quite different. In this invention, different neural networks are respectively used for different purposes in obtaining the secondary information of the human model and the model body data. By using neural network models with different input conditions and training methods, the precise contour separation of the human body in a complex background, the semantic segmentation of the human body, and the determination of key points and joint points are realized, excluding the influence of loose clothing and hairstyles, and approaching the real body shape and form of the human body to the greatest extent. The advantages of the deep learning network are fully utilized, and the pose and body shape of the human body can be restored with high precision in various complex scenarios. Moreover, the parameters output by the latter-level neural network include two categories: pose and shape, which can control the movement and body shape respectively. Combined with our reference model, the pose and body shape of the human model can be accurately replicated.
[0031] The present invention optimizes the fitting method of the basic mannequin and the target human model, forms an animation series from the initial pose to the target pose, and completes the whole process of bone movement by using an independent and more controllable self-owned basic mannequin and a segmented driving method, so that the whole fitting process achieves a good balance in terms of time and effect. Brief Description of the Drawings
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0033] Figure 1 Schematic diagram of the processing flow for driving a complete human body model in an embodiment;
[0034] Figure 2 Schematic diagram of the processing flow of the model parameter acquisition module in an embodiment;
[0035] Figure 3 Schematic diagram of the processing flow for modeling a human body model in an embodiment;
[0036] Figure 4 Schematic diagram of the system of the present invention. Detailed Embodiments
[0037] The following will describe in detail the features and exemplary embodiments of various aspects of the present invention. In order to make the purpose, technical solutions and advantages of the present invention clearer, the following will further describe the present invention in detail in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present invention by showing examples of the present invention.
[0038] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.
[0039] The following will describe in detail the segmented driving method of the human body model described in the embodiments of the present invention in conjunction with the drawings.
[0040] As shown Figures 1-3 in the figure, the present invention provides a method for segmentally driving a human body model. The method includes the following sub-steps: calculating and obtaining secondary information of the human body model using an algorithm; establishing a basic mannequin in the initial T-pose posture, and the acquisition of the initial posture parameters is determined by the initialization parameters of the basic mannequin model; obtaining the posture and body shape parameters of the target human body model by regression prediction from the neural network model according to the secondary information, and the three-dimensional human body posture and body shape parameters correspond to the bones and several basic parameters of the three-dimensional basic mannequin; inputting the obtained several groups of basic and bone parameters into the standard three-dimensional human body model for fitting to obtain the target posture and target body shape; generating an animation sequence from the initial posture to the target posture; driving the human body model mesh to move from the initial posture to the target posture by a segmental driving method; obtaining a three-dimensional target human body model mesh with the same posture as the target human body two-dimensional image. Of course, it is not necessary to strictly follow the strict front-to-back order between these steps, because some steps are independent preparation steps in themselves, and the placement order does not have a decisive impact on the final result.
[0041] Thus, it can be seen that the fitting method of the present invention generally involves three parts of steps. One is to generate several standard human body models; the second is to obtain the parameters of the human body model in the target posture; the third is to make the body shape and posture of the standard human body model fit to be consistent with the target human body model.
[0042] The first part is to pre-design and model some basic mannequins. As Figure 3As shown in the figure, the main work content is as follows: Construct a three-dimensional basic mannequin in combination with a mathematical model, that is, the basic mannequin or the basic mannequin. The SMPL human body model of the Max Planck Institute can avoid surface distortion during human movement and can accurately depict the morphology of human muscle stretching and contraction movements. In this method, β and θ are the input parameters, where β represents 10 parameters such as the proportion of a person's height, weight, body ratio, and head-to-body ratio, and θ is 75 parameters representing the overall movement pose of the human body and the relative angles of 24 joints. The β parameter is the shape Blend pose parameter, and the change of the human body shape can be controlled by 10 incremental templates. Specifically, the change of each parameter controlling the human body morphology can be depicted by a moving picture. By studying the continuous animation of parameter changes, we can clearly see that the continuous change of each parameter controlling the human body morphology will cause local or even overall chain changes in the human body model. In order to reflect the movement of human muscle tissue, the linear change of each parameter of the SMPL human body model will cause large-area mesh changes. Figuratively speaking, for example, when adjusting the parameter of β1, the model will directly understand the change of the parameter of β1 as the overall change of the body. You may only want to adjust the proportion of the waist, but the model will forcefully adjust the fatness and thinness of the legs, chest, and even hands together. Although this working mode can greatly simplify the work process and improve efficiency, it is really very inconvenient for projects pursuing modeling effects. Because the SMPL human body model is ultimately a model that conforms to the Western body type trained from Western human body photos and measurement data, and the law of its body shape change basically conforms to the general change curve of Westerners. When applied to the modeling of Asian human body models, many problems will occur, such as the proportion of arms and legs, the proportion of the waist, the proportion of the neck, the length of the legs and the length of the arms, etc. Through our research, if the SMPL human body model is rigidly applied on the basis of the driving method of the present invention, the minimum requirement of segmented movement cannot be achieved, resulting in the inability to fully achieve the technical purpose.
[0043] To this end, we adopted a self-made human model to improve the technical feasibility. The core is to build a human blend body base to achieve precise independent control of the human body. The three-dimensional basic mannequin has a mathematical weight relationship between the bone point and the model grid. The determination of the bone point can be associated with the human model of the target human posture; the three-dimensional basic mannequin is defined by a number of body base parameters and a number of root bone parameters. The several body bases constitute the entire human model grid, and each body base is controlled by the base parameters separately, without affecting each other. Preferably, the three-dimensional basic mannequin (basic mannequin) is composed of 20 body base parameters and 170 bone parameters. The so-called precise control, on the one hand, increases the control parameters, and does not follow the ten β control parameters of the Max Planck Institute. In this way, the adjustable parameters, in addition to the usual fat and thin, also add the length of the arm, the length of the leg, the fat and thin waist, hips and chest, etc., in terms of bone parameters, the parameters are more than doubled, which greatly enriches the range of adjustable parameters and provides a good foundation for the refined design of the basic mannequin. The so-called independent control can be understood as each base is controlled separately, such as the waist, legs, hands, head, etc. Each bone can also be adjusted in length separately, which is independent of each other and will not produce physical linkage, so that the human body model can be finely adjusted. The model no longer appears to be "stupid, big and clumsy", and it can never be adjusted to the shape that satisfies the designer. Our existing model reflects a correspondence based on mathematical principles. In fact, it is equivalent to redesigning this model from two parts: artificial aesthetics and data statistical analysis, so that it can generate the correct model that we think is suitable for the Asian body shape according to our design rules, which is significantly different from the big data training model of the SMPL human body model. Therefore, our parameter transformation is more interpretable and can better represent the local shape changes of the human body model. Moreover, this change is based on mathematical principles. There is no influence between the parameters, and the arms and legs remain completely independent. In fact, so many different parameters are designed to avoid the defects of the human body model trained by big data, and to accurately control the human body model in more dimensions, not limited to a few indicators such as height, which greatly improves the modeling effect.
[0044] The second part is to process the acquired human body images to obtain the parameter information required to generate the human body model. In the past, the selection of these skeletal key points was usually done manually, but this method is very inefficient and does not meet the fast-paced requirements of the Internet era. Therefore, in today's era of neural networks, the use of deep learning neural networks instead of manual key point selection has become a trend. However, how to use neural networks efficiently is an issue that requires further research. Generally speaking, we adopted the idea of a secondary neural network plus data "fine-tuning" to construct our parameter acquisition system. Figure 2As shown in the figure, we use a neural network with deep learning to generate these parameters, mainly including the following sub-steps: 1) Obtain a two-dimensional image of the target human body; 2) Process to obtain a two-dimensional human body contour image of the target human body; 3) Substitute the two-dimensional human body contour image into the first neural network with deep learning for the regression of joint points; 4) Obtain a joint point map of the target human body; Obtain semantic segmentation maps of various parts of the human body; Body key points; Body bone points; 5) Substitute the generated joint point map, semantic segmentation map, body bone points and key point information of the target human body into the second neural network with deep learning for the regression of human body posture and body shape parameters; 6) Obtain the output three-dimensional human body parameters, including three-dimensional human body motion posture parameters and three-dimensional human body shape parameters.
[0045] The two-dimensional image of the target human body can be a two-dimensional image including a human body image in any pose and any clothing. The acquisition of the two-dimensional human body contour image uses an object detection algorithm, and the object detection algorithm is a region proposal network based on a convolutional neural network.
[0046] Before inputting the two-dimensional human body image into the first neural network model, it also includes the process of training the neural network. The training samples include standard two-dimensional human body images with the original joint point positions marked, and the original joint point positions are marked on the two-dimensional human body images by humans with high accuracy. Here, first obtain the target image, and use the object detection algorithm to perform human body detection on the target image. Human body detection does not use measuring instruments to detect real human bodies. In the present invention, it actually refers to any given image, usually a two-dimensional photo containing sufficient information, such as a face, and the limbs and body of a person are required to be included in the picture. Then, use a certain strategy to search the given image to determine whether the given image contains a human body. If the given image contains a human body, parameters such as the position and size of the human body are given. In this embodiment, before obtaining the human body key points in the target image, it is necessary to perform human body detection on the target image to obtain a human body box marking the human body position in the target image. Since the pictures we input can be any pictures, there will inevitably be some backgrounds of non-human body images, such as tables, chairs, big trees, cars, buildings, etc. We need to remove these useless backgrounds through some mature algorithms.
[0047] Meanwhile, we also need to perform semantic segmentation, joint point detection, bone detection, and edge detection. After collecting this 1D point information and 2D surface information, it can lay a good foundation for generating a 3D human body model later. Use the first-level neural network to generate the joint point map of the human body. Optionally, the object detection algorithm can be a fast generation network for the target area based on a convolutional neural network. This first neural network requires a large amount of data training. Some photos collected from the network are manually annotated with joint points and then input into the neural network for training. After going through the deep learning neural network, it can basically obtain a joint point map with the same accuracy and effect as the manually annotated joint points immediately after inputting the photo, and the efficiency is dozens or even hundreds of times that of manual annotation.
[0048] In the present invention, obtaining the joint point positions of the human body in the photo only completes the first step and obtains the 1D point information. It is also necessary to generate 2D surface information based on this 1D point information. All these tasks can be completed through the neural network model and mature algorithms in the prior art. Since the present invention has redesigned the working process and intervention timing of the neural network model to reasonably design various conditions and parameters, the generation work of the parameters is more efficient and the degree of manual participation is reduced, which is very suitable for Internet application scenarios. For example, in a virtual clothing-changing program, users do not need to wait and can basically obtain the clothing-changing result instantaneously, which plays a crucial role in improving the attraction of the program to users.
[0049] After obtaining the relevant 1D point information and 2D surface information, these parameters or results, such as the joint point map of the target human body, semantic segmentation map, body bone points, and / or key point information, can be used as input items and substituted into the second neural network that has undergone deep learning for regression of human body posture and body shape parameters. Through the regression calculation of the second neural network, several groups of three-dimensional human body parameters can be immediately output, including three-dimensional human body motion posture parameters and three-dimensional human body shape parameters. Preferably, the loss function of the neural network is designed according to the three-dimensional basic mannequin (basic mannequin), the predicted three-dimensional human body model, the standard two-dimensional human body image with the original joint point positions marked, and the standard two-dimensional human body image including the predicted joint point positions.
[0050] The third part is to fit the parameters of the human body model with the human body model, which is also the innovation point of the present invention.
[0051] The segmented driving method further includes the following sub-steps: inputting the obtained several groups of basis and bone parameters into the standard three-dimensional human body model for fitting to obtain the target posture and target body shape; generating an animation sequence from the initial posture to the target posture; using the segmented driving method to drive the human body model mesh to move from the initial posture to the target posture; obtaining the three-dimensional target human body model mesh with the same posture as the target human two-dimensional image.
[0052] The three-dimensional human body model has a mathematical weight relationship between skeletal points and model meshes. The determination of skeletal points can be associated with the determination of the human body model for the target human body posture. In this part, by using these two types of parameters generated in the previous part, they can be substituted into the pre-designed human body model for the construction of the 3D human body model. The names of these two types of parameters are similar to those of the human SMPL model parameters of the Max Planck Institute, but the actual content is quite different. Because their bases are different, that is to say, the present invention uses a self-made three-dimensional basic mannequin (basic mannequin), and the SMPL model of the Max Planck Institute uses a basic mannequin generated by big data training. The generation calculation methods of the two models are different. Although finally both are embodied as the generated 3D human body models, their connotations are quite different. After this step, a preliminary 3D human body model will be obtained, including the mesh of the human body model with skeletal position and length information.
[0053] In this part, the human body model mesh has to complete the change from the initial posture to the target posture. Because what we input is only a photo, and the target human body posture in the photo is usually different from the basic mannequin. At this time, in order to fit the target human body posture, the change from the initial posture to the target posture has to be completed. In order to more realistically simulate the motion state of the model, when several groups of bases and skeletal parameters are fitted and driven in the standard three-dimensional human body parameter model, the following links are also included:
[0054] Preferably, during the process of generating the animation sequence, the movement of the human body model mesh is carried out in an interpolation manner. After each frame drives the skeletal movement, the vertex and face information of the human body model in the current state is calculated through the weight parameters of the standard human body model, and the current state of the human body model mesh is updated and recorded. By using the interpolation method, the characteristics of the self-made basic mannequin can be fully utilized to give play to the advantage of independent control; at the same time, the interpolation method can also ensure the complete implementation of other innovative technical solutions associated with the present invention. Because in the project of virtual clothing change, we have adopted a lot of innovative methods to ensure the processing speed and generation effect of the human body model, clothing model and their cooperation. Among them, using an adaptive method to drive the movement of the human body model is also an important means, and the present invention is one of several optimization methods parallel to this method. The main effect is to decompose the movement process and reprocess and re-optimize the decomposed results by multiple means to adjust the balance among the processing speed, data volume and processing effect.
[0055] Preferably, during the segmented driving process, first move the mesh parts with relatively small position changes and less likely to cause penetration during the driving process, and then move the mesh parts with relatively large position changes and more likely to cause penetration. The parts less likely to penetrate include the head, neck, chest, abdomen, and hips; the parts more likely to penetrate include the limbs and hands. In the traditional overall driving from the initial pose to the target pose, the computational load is relatively large. When the algorithm does not accurately represent the motion trajectory, mutual interference between meshes, that is, penetration, is likely to occur. During the movement of the human body model, in fact, the mesh movement laws of various parts of the limbs are different. Some meshes change significantly, while some meshes change little. Correspondingly, some meshes move violently, while some meshes hardly move at all. To better utilize this feature to achieve the purpose of reducing the probability of penetration, a method of segmentally driving the meshes of the human body model is adopted. First, move some model parts with a very small probability of penetration, and place the parts prone to penetration at the back for movement. In this way, different parts can adopt different driving strategies and methods, treat the parts conforming to different movement laws differently, and achieve the purpose of improving the overall fitting effect. For example, the parts less likely to penetrate, such as the head, neck, and chest, can be quickly driven to the target position first, and without excessive verification and inspection, basically meet our requirements. Then move other parts prone to penetration, and the speed can be slowed down. And in the final stage, introduce collision body inspection. Although it increases the computational load to a certain extent and slows down the movement speed, in the case of offsetting the time saved previously, the speed does not slow down significantly. However, due to additional verification and inspection, the fitting effect will be improved and the probability of penetration will decrease.
[0056] Preferably, the process of frame-by-frame driving further includes the following sub-steps: 1) Obtain the position coordinates of the initial pose and the target pose; the acquisition of the initial pose parameters is determined by the initialization parameters of the basic mannequin model, and the bone information of the target pose is predicted by regression of the neural network model.
[0057] 2) Generate an animation sequence moving from the initial pose to the target pose; after having the initial state of the bone information and the target pose state parameters, through interpolation methods such as linear interpolation and nearest neighbor interpolation, form a time series of bone information from the initial pose to the target pose. During the driving process, according to the number of bones driven per frame, it can be divided into two methods: global linear interpolation and driving the parent node first and then the child node. Considering the driving state in the simulated physical world, this patent adopts the latter method of driving the parent bone node first and then the child bone node. In this way, the animation sequence interpolated has actions that are more in line with the real physical world, and the simulated effect is better.
[0058] 3) During the process of generating the animation sequence, interpolation of the mesh is used for processing. That is, after each frame drives the movement of the bones, the vertex and face information of the human body model in the current state is calculated through the weight parameters of the basic mannequin, and the mesh state of the current human body model is updated and recorded. This is also a key step to ensure the model reduction degree, which can reduce the unwanted deformation and distortion of the human body model to an acceptable level.
[0059] 4) The interpolation speed is set to be slow at the initial and target positions of the front and back distances, and fast during the intermediate movement process. This patent adopts a non-uniform interpolation rate, that is, the single-frame movement amplitude is smaller during the initial and end actions, and larger during the intermediate movement process. It simulates that the starting state of physical actions in the real world has a certain acceleration process, while a higher inter-frame displacement distance is maintained during the movement process, and the driving speed is reduced at the end of the movement.
[0060] 5) When driving to the final target posture, it stays still for several frames to obtain the entire animation sequence. This approach is closer to the movement law of the real physical world than uniform interpolation. It enables the model to obtain an effective buffer before stopping after high-speed movement, thereby obtaining the entire complete animation sequence, forming a region with denser interpolation near the final posture, with higher accuracy in posture fitting of the human body model and better simulation effect.
[0061] 6) Complete driving the bones to move from the initial posture to the target posture. At this time, the basic mannequin has become a human body model that is basically consistent with the posture and body shape of the target human body model, with a natural posture and no penetration.
[0062] Combined Figures 1 to 3 The three-dimensional human body model segmentation fitting method according to the embodiments of the present invention described above can be implemented by a human body fitting processing device. Figure 4 FIG. 300 is a schematic diagram of the hardware structure of a device for processing human body fitting according to an embodiment of the invention.
[0063] The present invention also discloses a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the driving methods and steps described above are implemented.
[0064] And an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; when the processor executes the program stored in the memory, the driving methods and steps described above are implemented.
[0065] Such as Figure 4As shown in the figure, the device 300 for realizing human body fitting in this embodiment includes: a processor 301, a memory 302, a communication interface 303, and a bus 310. Among them, the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication with each other.
[0066] Specifically, the above-mentioned processor 301 may include a central processing unit (CPU), or a specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present invention.
[0067] The memory 302 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include an HDD, a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 302 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 302 may be inside or outside the device 300 for processing human body images. In a specific embodiment, the memory 302 is a non-volatile solid-state memory. In a specific embodiment, the memory 302 includes a read-only memory (ROM). In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0068] The communication interface 303 is mainly used to realize communication between various modules, devices, units, and / or devices in the embodiments of the present invention.
[0069] The bus 310 includes hardware, software, or both, and couples the components of the device 300 for processing human body images to each other. By way of example and not limitation, the bus may include an accelerated graphics port (AGP) or other graphics buses, an enhanced industry standard architecture (EISA) bus, a front-side bus (FSB), a hypertransport (HT) interconnect, an industry standard architecture (ISA) bus, an infinite bandwidth interconnect, a low-pin count (LPC) bus, a memory bus, a microchannel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standards association local (VLB) bus, or other suitable buses, or a combination of two or more of these. In a suitable case, the bus 310 may include one or more buses. Although the embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.
[0070] That is to say, Figure 4The processing device 300 shown can be implemented to include: a processor 301, a memory 302, a communication interface 303, and a bus 310. The processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication with each other. The memory 302 is used to store program codes; the processor 301 runs a program corresponding to the executable program code by reading the executable program code stored in the memory 302, so as to execute the method for three-dimensional human model fitting in any embodiment of the present invention, thereby implementing the combination Figures 1 to 3 the described three-dimensional human model fitting method and device.
[0071] An embodiment of the present invention further provides a computer storage medium, on which computer program instructions are stored; when the computer program instructions are executed by a processor, the method for three-dimensional human model fitting provided by the embodiment of the present invention is implemented.
[0072] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0073] It should also be noted that the functional blocks shown in the above structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0074] It also needs to be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0075] As described above, this is only the specific implementation manner of the present invention. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, modules, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A segmented driving method for a human body model, characterized in that The method includes: 1) Using an algorithm to calculate and obtain secondary information of a human body model, where the secondary information includes a joint point map, a semantic segmentation map, body bone points, and key point information of a target human body; 2) Establishing a basic mannequin in an initial T-pose posture, and obtaining the initial posture parameters determined by the initialization parameters of the basic mannequin model; 3) According to the secondary information, using a neural network model to regress and predict to obtain the posture and body shape parameters of the target human body model, where the posture and body shape parameters of the target human body model correspond to the bones and several base parameters of the three-dimensional basic mannequin; 4) Inputting the obtained several groups of base and bone parameters into the basic mannequin model for fitting to obtain the target posture and target body shape; 5) Generating an animation sequence from the initial posture to the target posture; 6) Adopting a segmented driving method to drive the human body model mesh to move from the initial posture to the target posture; 7) Obtaining a three-dimensional target human body model mesh with the same posture as the target human body two-dimensional image; Among them, during the process of generating the animation sequence, the movement of the human body model mesh is carried out in an interpolation manner. After each frame drives the bone movement, the vertex and face information of the human body model in the current state are calculated through the weight parameters of the standard human model, and the current state of the human body model mesh is updated and recorded and saved; Among them, during the segmented driving process, first move the mesh part with small position changes and not easy to cause penetration during the driving process, and then move the mesh part with large position changes and easy to cause penetration. Among them, different parts adopt different driving strategies and methods. Quickly move the mesh part with small position changes and not easy to cause penetration to the target position, and then slow down the speed to move the mesh part with large position changes and easy to cause penetration, and introduce a collision body check in the final stage.
2. The method according to claim 1, wherein The parts not easy to penetrate include the head, neck, chest, abdomen, and buttocks; the parts easy to penetrate include the limbs and hands.
3. The method according to claim 1, characterized in that, After obtaining the initial state of the bone information and the target posture state parameters, drive the bone to move from the initial posture to the target posture. Through the interpolation method of linear interpolation or nearest neighbor interpolation, a time series of bone information from the initial posture to the target posture is formed. During the driving process, first drive the parent bone node, and then drive the child bone node.
4. The method according to claim 1, wherein During the process of generating the animation sequence, the interpolation speed adopts a non-uniform interpolation rate, which is set to be slow at the positions far from the initial point and the target point, and fast in the middle movement process, that is, the single-frame movement amplitude is small in the initial action and the end action processes, and large in the middle movement process. When the interpolation reaches near the end of the movement, the driving speed is reduced, and it stops for several frames when moving to the final target posture to obtain the entire animation sequence.
5. The method according to claim 1, wherein The basic mannequin is a three-dimensional basic mannequin constructed by combining a mathematical model; the three-dimensional basic mannequin has a mathematical weight relationship between bone points and model meshes, and the determination of bone points can be associated with the human body model for determining the target human body posture; the three-dimensional basic mannequin is defined by several body shape base parameters and several bone parameters, and the several body shapes form the entire human body model mesh, and each body shape is separately controlled and changed by the base parameters without affecting each other.
6. The method according to claim 1, wherein The steps of obtaining the parameters of the target human body model further include: 1) obtaining a two-dimensional image of the target human body; 2) processing to obtain a two-dimensional human body contour image of the target human body; 3) substituting the two-dimensional human body contour image into a first neural network that has undergone deep learning for joint point regression; 4) obtaining a joint point map of the target human body; obtaining a semantic segmentation map of each part of the human body; body key points; body bone points; 5) substituting the generated joint point map, semantic segmentation map, body bone points, and key point information of the target human body into a second neural network that has undergone deep learning for human pose and body shape parameter regression; 6) obtaining the output three-dimensional human body parameters, including three-dimensional human body motion pose parameters and three-dimensional human body shape parameters.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.
8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Three-dimensional human body model reconstruction method, storage equipment and control equipment
CN110827342A