A method, device and storage medium for fast fitting of a human model

By combining multiple initial pose bases and hierarchical deep neural networks, the problems of slow speed and insufficient accuracy in human body modeling in existing technologies are solved, generating efficient and accurate human body models suitable for Asian body types, thus improving the effect of virtual try-on.

CN114119912BActive Publication Date: 2026-02-06BEIJING MOMO INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010876707.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-27
Publication Date
2026-02-06
Estimated Expiration
2040-08-27

AI Technical Summary

Technical Problem

Existing technologies for human body modeling based on single photographs suffer from slow modeling speed, insufficient accuracy, and an inability to independently control body shape and posture. They are particularly ineffective in virtual try-on, failing to meet the needs of internet users.

Method used

A standard human body model with multiple initial pose bases is used, combined with a hierarchical deep neural network, to generate an accurate human body model through clustering calculation and optimization fitting process.

Benefits of technology

It enables rapid fitting of human body models, improving modeling speed and accuracy, is suitable for Asian body types, allows independent control of individual body parts, and enhances the virtual try-on experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114119912B_ABST
    Figure CN114119912B_ABST
Patent Text Reader

Abstract

The application discloses a quick fitting method of a human body model, and the method comprises the following steps: establishing a standard human body model in an initial posture; acquiring a posture and a body shape parameter of a target human body model, wherein the three-dimensional human body posture and the body shape parameter correspond to a skeleton and a plurality of base parameters of a three-dimensional standard human body model; selecting an initial posture base of the standard three-dimensional human body model according to the obtained posture parameter of the target human body model; starting fitting from the selected initial posture base according to a plurality of groups of base and skeleton parameters of the target human body model; obtaining a three-dimensional human body model with the same body shape as the target human body; and obtaining a three-dimensional target human body model grid with the same posture as a two-dimensional image of the target human body. The initial posture base of the three-dimensional standard human body model can be a plurality of preset initial posture bases including a T-pose. The fitting process of the standard human body model and the target human body model is optimized in the application, a plurality of initial posture bases are used instead of a single T-pose initial posture in the prior art, and a variable-speed frame interpolation method is used to optimize the process of driving the skeleton to move, so that a good balance is achieved in time and effect in the whole fitting process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of human three-dimensional model modeling, and particularly relates to a human model rapid fitting method, in particular a method, device and storage medium for fitting a standard human model and a target human model based on multiple initial posture bases. BACKGROUND

[0002] With the development of Internet technology, online shopping is becoming more and more popular. Compared with shopping in a physical store, online shopping has the advantages of a wide variety of goods and convenience. However, there are still some problems that cannot be easily solved when buying goods online, the most important of which is that the goods to be purchased cannot be viewed in person. Among all the types of goods, the problem of clothing goods is the most prominent. Compared with the real-time dressing and viewing of clothing effects in physical stores, online clothing shopping cannot provide effect pictures for consumers themselves, but only provides model fitting pictures, and some do not even have fitting pictures. Consumers cannot obtain the matching degree of clothing and their own body shape in real time and intuitively. This has caused a large number of returns and exchanges.

[0003] In view of this problem, operators try to use virtual fitting technology to provide simulated fitting effects for consumers to solve this problem. Of course, there are other occasions where virtual dressing and fitting technology can be used in reality, such as in online games. Therefore, this technology has developed rapidly.

[0004] Virtual fitting refers to a technology application in which users do not need to actually change into clothes that they want to view the wearing effect, but can view the "dressing" effect in real time on the terminal screen. Existing dressing technology applications mainly include flat fitting and three-dimensional virtual fitting technology. The former basically collects pictures of users and pictures of clothes, and then performs cutting and splicing to form an image after "dressing", but such images have poor authenticity due to simple and crude image processing methods, and do not take into account the actual body shape of the user, but only fit the clothes to the user's photo, which cannot meet the needs of users. The latter usually collects three-dimensional information of a person through three-dimensional collection devices and combines the characteristics of clothes for synthesis, or manually inputs body data information provided by the user and virtually generates a three-dimensional human model according to certain rules, and then combines with clothes mapping. Overall, this kind of three-dimensional virtual fitting needs a large amount of data collection or three-dimensional data calculation, and the hardware cost is high, which is not easy to popularize among ordinary users.

[0005] With the development of cloud computing technology, artificial intelligence technology and intelligent terminal processing capability, two-dimensional virtual fitting technology is generated. Such technology mainly includes three steps: (1) processing the personal body information provided by the user to obtain a target human model; (2) processing the clothing information to obtain a clothing model; and (3) fusing the human model and the clothing model together to generate a simulation diagram of a person wearing the clothing.

[0006] For the first point, due to the accumulation of many uncertain factors such as process design, model parameter selection, neural network training method, etc., the quality of the generated clothes-changing picture is not as good as that of the traditional three-dimensional virtual fitting technology. Among them, the fitting of the human model is the basic step, and the subsequent dressing process must also be based on the previously generated human model. Therefore, if the human model is not accurate, it is easy to cause the gap between the human model and the body type of the person trying on clothes to be too large, the skin texture to be lost, the body parts to be lost, etc., which affects the effect of the finally generated clothes-changing picture.

[0007] In the general field of computer vision, there are many initial starting points for human modeling, usually including using a 3D scanning device to scan a real human body in all directions, a multi-view depth photograph-based three-dimensional reconstruction method, and a method of combining a given image with a human model to achieve three-dimensional reconstruction. Among them, using a 3D scanning device to scan a real human body in all directions obtains the most information and is the most accurate, but such devices are usually expensive and require the high cooperation of a human model, and the entire processing process has very high requirements for the processing device, so it is generally applied to some professional fields; secondly, the multi-view three-dimensional reconstruction method needs to provide multiple overlapping images of the reconstructed human body from different angles and establish the spatial conversion relationship between the images. A 3D model is obtained by using multiple cameras to take multiple pictures, which simplifies the operation to some extent, but the computational complexity is still large, and in most cases, only the person on site can obtain multi-angle pictures. The model obtained by the multi-angle photographing method of the depth camera is a map, which cannot provide a basis for 3D perception because it does not have body scale data. Thirdly, the method of combining a single image with a human model only needs to provide one image. A three-dimensional human feature curve intelligent generation method based on a neural network is used to train the neural network to obtain weights and thresholds that can be used to describe the curves of the neck, chest, waist and hips of the human body. Then, according to the size parameter information of the circumference, width and thickness of the human cross section, a human three-dimensional curve that agrees with the real human body shape can be directly generated to obtain the predicted human model. However, this method still consumes a large amount of calculation due to the small amount of input information, resulting in an unsatisfactory final model effect.

[0008] Based on the Internet technology and the network environment characteristics, the way of directly outputting the final human model from a single image is undoubtedly the preferred one, which is the most convenient, and the user does not need to be on site, only a photo can complete the whole dressing process. Then the problem that follows is that as long as the result photo effect obtained is basically equivalent to the real 3D simulation dressing, it will become the mainstream. Among them, how to quickly and well obtain the human model closest to the real state of the human body through a photo has become the top priority.

[0009] In the prior art, there are several methods for constructing a human model: (1) a regression-based method, which reconstructs a voxel-represented human model through a convolutional neural network. The algorithm first estimates the positions of the main joints of the human body according to the input picture, then estimates the given specified size voxel grid according to the key point position, and according to whether each unit voxel inside is occupied, the entire shape of the internal occupied voxel is used to describe the reconstructed human shape; (2) a single picture-based human body reconstruction method, which estimates the three-dimensional shape and posture of the human body at the same time. This method first roughly labels the simple human skeleton key points on the image, then performs initial matching and fitting of the human model according to these rough key points to obtain the general shape of the human body. (3) A human skeleton is represented by 23 skeletal nodes, and the rotation of each skeletal node is used to represent the posture of the entire human body. At the same time, the shape of the human body is expressed by 6890 vertex positions. In the fitting process, the skeletal node position is given, and the shape and posture parameters are fitted at the same time to perform three-dimensional human body reconstruction. Or first use a CNN model to predict the key points on the image, then use the SMPL model to fit to obtain the initial human model. Then, the shape parameters obtained by fitting are used to regress a human joint bounding box, each joint corresponds to a bounding box, and the axis length and radius are used to represent the bounding box. Finally, the initial model and the bounding box obtained by regression are combined to obtain three-dimensional human body reconstruction. The above methods have the problems of slow modeling speed, insufficient modeling accuracy, and strong dependence of the reconstruction effect on the created body and posture database.

[0010] The prior art discloses a human body modeling method based on body measurement data, like Figure 1As shown, the method comprises: acquiring body measurement data; performing linear regression on a pre-created human body model by a pre-trained prediction model according to the body measurement data, to fit a predicted human body model, wherein the pre-created human body model comprises a plurality of pre-defined sets of marker feature points and corresponding standard shape bases, and the body measurement data comprises measurement data corresponding to each set of marker feature points; and obtaining a target human body model according to the predicted human body model, wherein the target human body model comprises measurement data, a target shape base and a target shape coefficient. However, this method requires very high body measurement data, including body length data and girth data, such as height, arm length, shoulder width, leg length, calf length, thigh length, foot length, head circumference, chest circumference, waist circumference, thigh circumference, etc., which not only needs to be measured but also needs to be calculated. Indeed, it saves calculation, but the user experience is very poor, and the program is very tedious. Moreover, the training of the human body model refers to the training method of the SMPL model.

[0011] The SMPL model is a parameterized human body model, which is a human body modeling method proposed by the Max Planck Institute. This method can perform arbitrary human body modeling and animation driving. The biggest difference between this method and the traditional LBS is the method of human posture image surface topography proposed by it. This method can simulate the bulging and concave of human muscles during limb movement. Therefore, it can avoid the surface distortion of the human body during movement, and can accurately depict the muscle stretch and contraction movement topography. In this method, β and θ are input parameters, wherein β represents 10 parameters of human body height, fat, head-to-body ratio, etc., and θ represents 75 parameters of overall human body movement posture and 24 joint relative angles. However, the core of this model generation method is the accumulation of a large amount of training data to obtain the relationship between body shape and shape base. However, due to the strong correlation between each other, it is not easy to decouple the operation of each shape base, such as the correlation between the arm and the leg. Theoretically, the leg will also move when the arm moves, and it is difficult to improve the model for different characteristics of the body shape on the SMPL model.

[0012] The prior art two discloses a 3D human modeling method based on a single photo, comprising: acquiring a photo, analyzing the photo, marking the key points of the human body in the photo, and calculating the spatial coordinates of the key points; obtaining the distance between the key points in the photo and the skeleton points in the pre-created standard human model, aligning the skeleton points with the key points to generate a basic human model; obtaining the basic map in the pre-created standard human model, performing difference calculation on the basic map and the skin texture of the face in the photo, and then using the edge channel to fuse to generate basic texture data; and generating a 3D human model according to the basic human model and the basic texture data. 3D human modeling is realized through a photo, and the model has a skeleton and muscle system support, which can produce expressions and actions. However, this method matches the distance between the key points of the user photo and the key points of the standard human platform, then adjusts the distance to achieve the posture of the target human body, and then performs difference calculation and fusion on the basic map and the skin texture in the photo to obtain the final human model. This method has a simple process and small calculation amount, but the precision and reality of the generated human model are not high.

[0013] The prior art three discloses a three-dimensional human model generation method, comprising: acquiring a two-dimensional human image; inputting the two-dimensional human image into a three-dimensional human parameter model to obtain a three-dimensional human parameter corresponding to the two-dimensional human image; inputting the training sample into a neural network for training to obtain a three-dimensional human parameter model, comprising: inputting the standard two-dimensional human image in the training sample into the neural network to obtain a predicted three-dimensional human parameter corresponding to the standard two-dimensional human image; adjusting a three-dimensional flexible deformable model according to the predicted three-dimensional human parameter to obtain a predicted three-dimensional human model; and obtaining a predicted joint position in the standard two-dimensional human image through reverse mapping according to the joint position in the predicted three-dimensional human model. This modeling method uses model judgment and the parameters output by the neural network at the end to only have joint parameters, and then uses the mature body shape of the SMPL model to make detailed adjustments consistent with the target human posture. Although the calculation amount is reduced, since the input parameters are few and the adjustment can only be completed on the basis of the SMPL prediction model, it is difficult to output a human model highly consistent with the target human posture.

[0014] Therefore, in order to adapt to the development trend of the Internet industry, in the virtual fitting sub-field, the least input information, the least calculation amount and the best effect will be the three basic goals that have been pursued. We need to find a best balance point among the three, and provide a fast human model fitting method that can achieve fast fitting speed, calculation amount not exceeding the bearing capacity of terminal equipment, and effect close to professional equipment. SUMMARY

[0015] Based on the above problems, the application provides a human model fast fitting method, device and storage medium for overcoming the above problems.

[0016] The application provides a human model fast fitting method, which comprises the following steps: establishing a standard human model in an initial posture; obtaining a posture and a body shape parameter of a target human model, the three-dimensional human posture and the body shape parameter corresponding to the skeleton of a three-dimensional standard human model and a plurality of base parameters; selecting an initial posture base of the standard three-dimensional human model according to the obtained posture parameter of the target human model; starting fitting from the selected initial posture base according to the obtained base and skeleton parameters of the target human model; obtaining a three-dimensional human model with the same body shape as the target human model; and obtaining a three-dimensional target human model grid with the same posture as the two-dimensional image of the target human model.

[0017] Preferably, the initial posture base of the three-dimensional standard human model can be a plurality of preset initial posture bases including a T-pose.

[0018] Preferably, in the process of three-dimensional human model movement fitting, first, the posture parameter representing the skeleton position of the target human model is obtained; second, the distance from which initial posture to the target posture is the shortest is calculated according to the position coordinates of a plurality of common preset initial posture bases; and finally, the fitting is started from the initial posture base of the preset posture to improve the processing speed.

[0019] Preferably, the obtaining of the plurality of preset initial posture bases comprises the following steps: 1) obtaining a certain number of posture parameters representing the skeleton position of the target human model; 2) performing clustering calculation on the obtained plurality of sets of posture parameters according to a set rule; and 3) calculating the initial posture base required for achieving the target according to the clustering calculation target.

[0020] Preferably, the clustering rule adopts a distance criterion. The target of the clustering calculation is that the movement distance of each initial posture to the target posture is less than a given limit value or the interpolation frame number during fitting is less than a given limit value.

[0021] Preferably, the step of obtaining the pose parameters of the target human model bone position further comprises: 1) obtaining a two-dimensional image of a target human body; 2) processing the obtained two-dimensional human body contour image; 3) substituting the two-dimensional human body contour image into a first neural network trained by deep learning to regress the joint nodes; 4) obtaining a joint node map of the target human body; obtaining a human body part semantic segmentation map; body key points; body skeleton points; 5) substituting the generated joint node map, semantic segmentation map, body skeleton points and key point information of the target human body into a second neural network trained by deep learning to regress the human body pose and body shape parameters; 6) obtaining the output three-dimensional human body parameters, including three-dimensional target human body bone position pose parameters and three-dimensional human body shape parameters. Further comprising, the two-dimensional human body contour image is obtained by using a target detection algorithm, and the target detection algorithm is a target region fast generation network based on a convolutional neural network; before the two-dimensional human body image is input into the first neural network model, the process of training the neural network is further included, and the training sample includes a standard two-dimensional human body image labeled with original joint node positions, and the original joint node positions are manually labeled on the two-dimensional human body image with high accuracy.

[0022] In addition, a computer readable storage medium is also provided, characterized in that the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method steps of any of the preceding methods.

[0023] An electronic device, characterized in that it comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; the processor is used to execute the program stored on the memory to implement the method steps of any of the preceding methods.

[0024] The beneficial effects of the present application are:

[0025] 1. Fast Model Fitting Speed. As is well known, the fitting of human body poses in virtual try-on typically uses a standard human body model in a T-pose as the initial pose. This standard human body model then moves frame by frame from the initial pose to the target pose. This process involves mesh and clothing calculations, making each frame's calculation extremely time-consuming and resource-intensive. To quickly fit the target human body's pose, we use an optimized initial pose base as the starting point for fitting, replacing the traditional T-pose. The skeletal information of the target pose is obtained through model regression prediction. Compared to the traditional standard human body model starting from the T-pose, our standard human body model starts fitting from a position "closer" to the target pose, reducing the movement distance by more than ten times and reducing the required number of interpolation frames by more than 80%. This significantly saves the system's computational processing time, demonstrating a substantial speed advantage in internet-based consumer entertainment applications.

[0026] 2. Multiple standard human models were designed in advance, encompassing initial pose bases for several common human postures obtained through statistical data processing. These include basic postures such as sitting, raising hands, raising legs, and folding hands, offering a much richer range of postures than a single T-pose standard human model. With this design, the 3D standard human model can be any of the predetermined poses other than the initial T-pose. During the movement of the 3D human model, the shortest distance to the target pose is calculated from which common pose. Then, the standard human model and the target human model are fitted starting from the initial pose base of this common pose. This significantly shortens the system's computation time and saves computing power, allowing the standard human model to be driven to the target pose quickly and accurately. This is highly advantageous given that mobile devices may become the most commonly used terminal for virtual clothing changing.

[0027] 3. Creative use of hierarchical deep neural networks to obtain clustering data. This invention fully utilizes the advantages of deep learning networks, enabling high-precision reconstruction of human posture and body shape in various complex scenarios. Different neural networks are used for different purposes, employing neural network models with different input conditions and training methods to achieve accurate contour separation, semantic segmentation, and keypoint and joint point determination of the human body in complex backgrounds. This eliminates the influence of loose clothing and hairstyles, achieving the closest approximation to the real human body shape and form. Our first-level neural network uses human contour, semantic segmentation, keypoints, and joint points as input items, generating model parameters from multiple perspectives. Furthermore, the output parameters of the subsequent-level neural network include both posture and body shape categories, obtaining accurate posture data for the human model. With this accurate human posture data, clustering methods can be used to group these data into several groups that meet certain conditions; these groups represent the most "common" postures in these data samples.

[0028] 4. Precise and Controllable Human Body Model. Currently popular single-image-based human body reconstruction methods mainly fall into two categories: reconstructing parametric human body models. The most commonly used parametric model is the Max Planck Institute's SMPL model, which contains two sets of 72 parameters describing human posture and body shape. For single-image reconstruction, the SMPL parameters are first estimated from the image, and then optimized by minimizing the projection distance between the 3D joints and the 2D planar joints, thus obtaining the human body. However, the SMPL model is mainly trained through a large number of human body model instances using deep learning. The relationship between body shape and shape basis is a holistic one, making decoupling very difficult. It is impossible to control the desired body parts at will, resulting in a model that cannot achieve a high degree of consistency with the real human posture and body shape. Furthermore, if it is further applied to the subsequent dressing process, it will also have limited ability to represent the geometric details of the human body surface, and cannot reconstruct the detailed texture of clothing on the human body surface well. However, our human body model is not obtained through training. The parameters have a correspondence based on mathematical principles. In other words, our different sets of parameters are independent of each other and have no interrelationship. Therefore, our model is more interpretable during transformation and can better represent the shape changes of a certain part of the body. In layman's terms, everyone's body shape is different, and many people's thigh and calf ratios do not meet a certain precise proportion. Our model can control the input parameters to achieve separate control and length adjustment of the thigh and calf, thus achieving a precise determination of the leg proportions.

[0029] 5、More suitable for Asian body shape. Human modeling usually designs some standard human models, that is, we call them standard human models or basic human models. Through the self-built standard human model, the human body can be controlled, that is, from the starting pose to the target pose process control, which is the basis for the change of the clothes with the human body pose after the work is completed. Only the human body can accurately reach the target pose, the specific process of the clothes following the human body to reach the target pose can be calculated. In this process, we abandon the use of the SMPL model of the Max Planck Institute which relies on the basic human models trained by the European body shape data, and instead establish a set of standard human models (standard human models) that are more suitable for the body shape of Asians. This set of human models can include 170 bones and 20 parameters of body shape bases, greatly enriching the details of the human model, and the expression of details exceeds the SMPL model. And combined with the characteristics of each base being independently controlled as described above, each part of the human model can be independently and accurately controlled and modified according to the requirements, achieving a more beautiful effect of each human model. In addition, we will also manually adjust the local part of the human model, such as the number of vertices and the number of faces, which is a function that other models represented by the SMPL model cannot complete. In addition to the height of the model which can be accurately adjusted, other parameters such as body shape, arm length, leg ratio, waist length and waist circumference can be accurately controlled, so that the human model is more suitable for the user's body shape.

[0030] The fitting process of the standard human model and the target human model is optimized, a plurality of initial pose bases are used instead of the traditional single T-pose initial pose, and the variable speed frame interpolation method is used to optimize the process of driving the bones to move, so that the whole fitting process is balanced in time and effect. In addition, through the self-owned standard human model, the body shape bases suitable for the characteristics of Asian body shape are selected, and a three-dimensional human model more close to the Asian body shape and better independent operation and control than the SMPL model of the Max Planck Institute human body is generated. BRIEF DESCRIPTION OF DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments described in the present application, and those skilled in the art can also obtain other drawings according to these drawings without creating any creative labor.

[0032] Figure 1 A human model fitting process flowchart for an embodiment;

[0033] Figure 2 A model parameter obtaining module process flowchart for an embodiment;

[0034] Figure 3 A flowchart of a complete modeling process for an embodiment;

[0035] Figure 4 A flowchart of a fitting process method for an embodiment;

[0036] Figure 5 A schematic diagram of a system of the present application. DETAILED DESCRIPTION

[0037] Features and exemplary embodiments of various aspects of the present application will be described below in detail, in order to make the purposes, technical solutions and advantages of the present application more clear and apparent, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are configured only to explain the present application, and are not configured to limit the present application. The present application can be implemented without some of these specific details by those skilled in the art. The following description of the embodiments is merely to provide a better understanding of the present application by showing examples of the present application.

[0038] It should be noted that, in this document, relational terms such as first and second, and the like, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the phrase "comprising a... " does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0039] The method for processing human body images provided by the embodiments of the present application will be described in detail below in combination with the drawings.

[0040] As Figure 1 shown, the embodiments of the present application provide a fast generation method, device and storage medium of a three-dimensional human body model, by pre-setting the initial posture base of a plurality of standard human body models, the calculation amount and complexity during the generation of the three-dimensional human body model can be greatly reduced, and the effect of exceeding the real degree of the existing 2D picture generated human body model can be achieved.

[0041] First, this invention discloses a rapid fitting method for human body models. The method includes: establishing a standard human body model in an initial posture; obtaining the posture and body shape parameters of the target human body model, wherein the three-dimensional human body posture and body shape parameters correspond to the skeleton and several basis parameters of the three-dimensional standard human body model; selecting an initial posture basis of the standard three-dimensional human body model based on the obtained posture parameters of the target human body model; performing fitting from the selected initial posture basis based on the obtained several sets of basis and skeleton parameters of the target human body model; obtaining a three-dimensional human body model with the same body shape as the target human body; and obtaining a three-dimensional target human body model mesh with the same posture as the two-dimensional image of the target human body.

[0042] The method generally includes three steps: first, generating several standard human body models; second, obtaining the parameters of the target human body model; and third, fitting the body shape and posture of the standard human body models to match those of the target human body model.

[0043] The first part involves processing the acquired human images to obtain the parameter information needed to generate the human model. Previously, the selection of these skeletal key points was usually done manually, but this method is inefficient and unsuitable for the fast-paced demands of the internet age. Therefore, with the widespread use of neural networks, using deep learning-based neural networks to replace manual key point selection has become a trend. However, how to efficiently utilize neural networks is a problem that requires further research. Overall, we adopted a two-level neural network plus data "refinement" approach to construct our parameter acquisition system. For example... Figure 2 As shown, we use a deep learning neural network to generate these parameters, which mainly includes the following sub-steps: 1) Obtain a two-dimensional image of the target human body; 2) Process the image to obtain a two-dimensional human body contour image; 3) Substitute the two-dimensional human body contour image into the first deep learning neural network to perform joint point regression; 4) Obtain the joint point map of the target human body; obtain semantic segmentation maps of various parts of the human body; body key points; body skeleton points; 5) Substitute the generated joint point map, semantic segmentation map, body skeleton points and key point information of the target human body into the second deep learning neural network to perform regression of human posture and body shape parameters; 6) Obtain the output three-dimensional human body parameters, including three-dimensional human body action posture parameters and three-dimensional human body shape parameters.

[0044] The two-dimensional image of the target human body can be a two-dimensional image including a human figure in any pose and with any clothing. The acquisition of the two-dimensional human body contour image utilizes a target detection algorithm, which is a target region fast generation network based on a convolutional neural network.

[0045] Before inputting the two-dimensional human body image into the first neural network model, the process of training the neural network is further included, and the training sample includes a standard two-dimensional human body image with labeled original joint position, which is labeled by artificial high-accuracy labeling on the two-dimensional human body image. Here, first, the target image is acquired, and the target image is detected by using a target detection algorithm. The human body detection is not detected by using a measuring instrument to detect a real human body, and in the present application, it actually refers to searching a given image, usually a two-dimensional photo containing sufficient information, such as a human face, human limbs and a body required to be included in the picture. Then, a certain strategy is used to search the given image to determine whether the given image contains a human body, and if the given image contains a human body, the position, size and other parameters of the human body are given. In the embodiment, before acquiring the human body key points in the target image, the target image needs to be detected to acquire the human body frame labeled with the human body position, because the input picture can be any picture, so there are inevitably some non-human body image backgrounds, such as tables, chairs, trees, cars and buildings, which need to be removed by using some mature algorithms.

[0046] At the same time, semantic segmentation, joint detection, skeleton detection and edge detection are also performed, and the collected 1D point information and 2D surface information can lay a good foundation for generating a 3D human body model. The first neural network is used to generate a human body joint graph, and the target detection algorithm can be a target region fast generation network based on a convolutional neural network. The first neural network needs to be trained with a large amount of data, and some photos collected from the network are labeled with joint points by artificial labeling, and then input into the neural network for training. After the deep learning neural network, the joint point graph with the same accuracy and effect as the artificial joint point labeling can be obtained immediately after the input photo, and the efficiency is dozens of times or even hundreds of times of the artificial labeling.

[0047] In the present application, obtaining the joint position of the human body in the photo only completes the first step, and 1D point information is obtained, and 2D surface information needs to be generated according to the 1D point information, and these work can be completed by using the neural network model and the mature algorithm in the prior art. In the present application, the work flow and intervention time of the neural network model are redesigned, various conditions and parameters are reasonably designed, the parameter generation work is more efficient, the degree of human participation is reduced, and it is very suitable for internet application scenarios, such as a virtual dressing program, so that the user does not need to wait, and basically the dressing result can be obtained instantly, which plays a crucial role in improving the attraction of the program to the user.

[0048] After obtaining the relevant 1D point information and 2D surface information, these parameters or results, the target human body joint point graph, the semantic segmentation graph, the body skeleton point and / or key point information can be taken as input items into the second neural network trained by deep learning to regress the human body posture and body shape parameters. Through the regression calculation of the second neural network, several groups of three-dimensional human body parameters, including three-dimensional human body action posture parameters and three-dimensional human body shape parameters, can be immediately output. Preferably, the loss function of the neural network is designed according to the three-dimensional standard human body model (basic mannequin), the predicted three-dimensional human body model, the standard two-dimensional human body image with labeled original joint point positions and the standard two-dimensional human body image including predicted joint point positions.

[0049] The second part is to pre-design and model some basic mannequins. The main work content is to construct a three-dimensional standard human body model, i.e. a standard human body model or a basic mannequin, in combination with a mathematical model. The SMPL human body model of Max Planck Institute can avoid surface distortion of the human body during movement, and can accurately depict the appearance of muscle stretching and contraction movement. In this method, β and θ are input parameters, in which β represents 10 parameters of personal body height, fatness, head-to-body ratio, etc., and θ represents 75 parameters of overall human body movement posture and 24 joint relative angles. The β parameter is a body shape Blend pose parameter, which can control the shape change of the human body through 10 incremental templates. Specifically, the change of each parameter controlling the human body shape can be depicted by a motion picture. By studying the continuous animation of parameter changes, we can clearly see that the continuous change of each parameter controlling the human body shape will cause local or even overall chain changes of the human body model. In order to reflect the movement of human muscle tissue, the linear change of each parameter of the SMPL human body model will cause large-area grid changes. In other words, for example, when adjusting the β1 parameter, the model will directly understand the β1 parameter change as the entire change of the body. You may only want to adjust the waist ratio, but the model will force you to adjust the fatness of the legs, chest and even hands together. Although this working mode can greatly simplify the workflow and improve efficiency, it is indeed very inconvenient for projects that pursue modeling effects. Because the SMPL human body model is ultimately a model trained by Western human body photos and measurement data, which conforms to the Western body shape, and the shape change rule basically conforms to the usual change curve of Westerners. When applied to human body modeling in Asia, there will be many problems, such as the proportion of arms and legs, the proportion of waist, the proportion of neck, the length of legs and arms, etc. Through our research, there are large differences in these aspects. If the SMPL human body model is used rigidly, the final generated effect cannot meet our requirements.

[0050] To this end, we use the self-made human model to improve the effect. The core is to build a human body blend body base to realize the accurate independent control of the human body. The three-dimensional standard human model has a mathematical weight relationship between the bone points and the model grid, and the determination of the bone points can be associated with the determination of the human body model of the target human body posture; the three-dimensional standard human body model is defined by a plurality of body base parameters and a plurality of bone parameters, and the plurality of body bases constitute the whole human body model grid, and each body base is individually controlled and changed by the base parameter, and does not affect each other. Preferably, the three-dimensional standard human model (basic human platform) is composed of 20 body base parameters and 170 bone parameters. The so-called accurate control is to increase the control parameters, and the ten beta control parameters of the Max Planck Institute are not followed, so that the adjustable parameters are added to the length of the arm, the length of the leg, the fatness of the waist, the hip and the chest, etc. in addition to the usual fatness, and the number of bone parameters is increased by more than one time, greatly enriching the range of adjustable parameters and providing a good foundation for fine design of the standard human model. The so-called independent control can be understood as each base being individually controlled, such as waist, leg, hand, head, etc., and each bone can be individually adjusted in length, independent of each other, and without linkage in shape. This can better adjust the human body model. The model is no longer "stupid and clumsy", and it is difficult to adjust to the designer's satisfied shape. Our existing model embodies a corresponding relationship based on mathematical principles. In fact, it is equivalent to redesigning the model from two parts of artificial aesthetics and data statistical analysis to generate the correct model that we think conforms to the Asian body type according to our design rules. It is significantly different from the big data training model of the SMPL human body model, so the parameter transformation has better interpretability and can better represent the local shape change of the human body model. Moreover, this change is based on mathematical principles, and there is no influence between parameters. The arm and the leg remain completely independent. In fact, designing so many different parameters is to avoid the defects of the human body model trained by big data, to accurately control the human body model in more dimensions, not limited to a few indicators such as height, and to greatly improve the modeling effect. Only under the premise of self-built body base, setting so many independent control parameters has practical significance. To achieve the requirements of the designer, both are indispensable.

[0051] In the present application, in order to improve the fitting speed of the standard human model, a plurality of self-made standard human models (standard human platforms) in different postures are creatively designed, and the initial posture base of the three-dimensional standard human model can be a plurality of preset initial posture bases including T-pose. In the process of moving fitting of the three-dimensional human model, we can fully utilize these standard human models to greatly shorten the solving and fitting time. The fitting process starts from the following steps, see Figure 1:

[0052] First, the pose parameters representing the target human model bone position are obtained, and the determination of these parameters is equivalent to determining the pose of the target human body. These parameters do not need to be calculated and extracted separately. In fact, these data have been calculated during the processing of the two-dimensional photo input by the user, for generating the target human model. Here, we make a process innovation, and reuse these data that can be obtained without calculation and time.

[0053] Second, according to the position coordinates of several common pre-designed initial pose bases, the shortest distance from which initial pose to the target pose is calculated. In this step, we calculate and superimpose the bone position coordinates of each pre-made basic human model with the bone position coordinates of the target human model one by one, and the basic human model and the target human model with the smallest integrated value are taken as the objects of our model fitting.

[0054] Finally, the initial pose base from which the fitting starts is selected, and the basic process of the fitting method is basically consistent with the prior art. In some links, we add our unique design, and after this processing, the processing speed of the model fitting is significantly improved.

[0055] In the previous steps, an important link is to determine which pre-set initial pose bases are. Only when the number of initial pose bases is sufficient and the positions are reasonable, can the role of multiple initial pose bases be maximized. The initial parameters of the initial pose and the grid include the following steps:

[0056] 1) Obtain a certain number of pose parameters representing the target human model bone position; as described above, these massive model parameters do not need to be calculated and extracted separately. In fact, during the processing of the two-dimensional photo input by the user, these data have been calculated for generating the target human model. We reuse these data that can be obtained without calculation and time for the subsequent clustering calculation, so as to extract the standard human model we need from these massive models.

[0057] The step of obtaining the pose parameters of the target human model bone position includes obtaining two-dimensional images of several target human bodies; processing the obtained two-dimensional human body contour images; substituting the two-dimensional human body contour images into a first neural network trained by deep learning to regress the joint nodes; obtaining the joint node graph of the target human body; obtaining the semantic segmentation graph of each part of the human body; the body key points; the body skeleton points; substituting the generated joint node graph, semantic segmentation graph, body skeleton points and key point information of the target human body into a second neural network trained by deep learning to regress the human body pose and body shape parameters; and obtaining the output three-dimensional human body parameters, including three-dimensional target human body bone position pose parameters and three-dimensional human body body shape parameters. The two-dimensional human body contour image is obtained by using a target detection algorithm based on a convolutional neural network target region fast generation network. Before the two-dimensional human body image is input into the first neural network model, a process of training the neural network is further included. The training samples include standard two-dimensional human body images labeled with original joint node positions, and the original joint node positions are manually labeled on the two-dimensional human body images with high accuracy. In summary, if the number of collected human body photos is sufficient, the calculated “common” poses will be more representative, and the model fitting speed can be improved.

[0058] 2) According to the set rules, the obtained several groups of pose parameters are calculated by clustering, and the clustering rules adopt distance criteria. Clustering is to divide a data set into different classes or clusters according to a certain standard (such as distance criteria), so that the similarity of data objects in the same cluster is as large as possible, and the difference of data objects not in the same cluster is as large as possible. That is, the data in the same class is as close as possible after clustering, and simply speaking, it is to divide similar things into a group. When clustering, we do not care what a class is, and the goal we need to achieve is to gather similar things together. Therefore, a clustering algorithm usually only needs to know how to calculate the similarity to start working. Clustering usually does not need to learn using training data. If the distance criteria is used as the judgment standard, and the clustering algorithm matching the target is used, several groups of data meeting certain distance requirements can be divided into a limited number of groups.

[0059] 3) According to the target calculated by clustering, the initial pose base required to achieve the target is calculated. In the present application, the target of clustering calculation is that the moving distance of each initial pose to the target pose is less than a given limit value or the interpolation number when fitting is less than a given limit value.

[0060] If according to our self-made standard human model standard, each target human model actually includes 20 body base parameters and 170 bone position parameters, if 3000 human body photos are substituted into the neural network model at one time, 3000 groups of data containing 170 bone position parameters will be obtained, among the 3000 groups of data, they can be divided into 10 large groups, and each group of data in the 10 large groups should meet certain conditions, that is, have certain similar characteristics, specifically to the present application, a large group may be a standard model that tends to meet the following conditions: the left hand is raised relatively high, the legs are slightly opened and straight, the right hand is generally lowered, and the waist and chest are kept naturally straight. The standard human model bone position parameters of this large group are determined by all the bone parameters in this large group. Of course, according to different clustering algorithms, the number of standard models will also change, and the number can be determined by our actual needs. At present, we usually adopt the target standard that the moving distance from the initial pose to the target pose is less than a given limit, that is, the total distance moved between the bone position coordinates of all models in the large group and the virtual standard human model calculated according to the data of the large group is less than a given threshold, for example, less than 0.5 meters. Of course, the number of frames inserted when fitting from the initial pose to the target pose can also be used as a limit, for example, less than 8 frames. As a comparison, if fitting is performed from a conventional single T-pose initial position, the total moving distance will generally increase by 5-10 times, and the moving time and cloth simulation processing time will generally increase by more than 10 times. In terms of processing effect, too long fitting time and too large moving distance will greatly increase the probability of model wearing, and seriously affect the virtual fitting effect. Of course, we sometimes also use the second group of parameters to calculate and determine which initial pose base to select, that is, introduce the grid vertex distance parameter, which is theoretically no different, except that one is the bone moving distance and the other is the grid vertex moving distance. These parameters are actually obtained in a calculation process, so they are essentially the same, except that the vertex distance calculation is more complex than the bone parameter, which is more direct. However, the vertex distance will have a greater impact on the fitting time in some cases. Therefore, in some cases, combining the two parameters will obtain a more accurate comprehensive calculation result, that is, it can help us select a more optimal initial pose base.

[0061] In order to balance the time and cost of pre-made standard human body model, we usually select an appropriate initial pose base by setting a reasonable target, such as the total moving distance being less than 0.5 meters or the total number of inserted frames being less than 10 frames. When determining the initial pose base, it is possible that the total moving distance is greater than 0.5 meters, which is also normal. There are thousands of human body poses, and the number of standard human body models used for clustering calculation is limited. Therefore, the standard human body model with the minimum total moving distance between the target human body model and the initial pose of all standard human body models can be selected as the initial pose for fitting.

[0062] The third part is to fit the parameters of the human body model and the human body model. As shown in Figures 3-4 the figure, the obtained three-dimensional human body pose and body type parameters are corresponded to several bases and bone parameters of the three-dimensional standard human body model; the obtained several groups of bases and bone parameters are input into the standard three-dimensional human body parameter model for fitting; the three-dimensional human body model with the same body type as the target human body is obtained; the movement of the driven bone from the initial pose to the target pose is completed; and the three-dimensional target human body model mesh with the same pose as the target human body two-dimensional image is obtained.

[0063] The three-dimensional human body model has a mathematical weight relationship between the bone points and the model mesh, and the determination of the bone points can be associated with the determination of the human body model of the target human body pose. In this part, the two types of parameters generated in the previous part can be substituted into the pre-designed human body model to construct the 3D human body model. The two types of parameters are similar to the parameter names of the human body SMPL model of Max Planck Institute, but the actual contents are quite different. Because the foundations are different, that is, the self-made three-dimensional standard human body model (basic human platform) is used in the present application, and the standard human body model generated by big data training is used in the SMPL model of Max Planck Institute. The two models are generated in different ways, although they finally represent the generated 3D human body model, but the connotation is quite different. After this step, a preliminary 3D human body model is obtained, including the mesh of the human body model with bone position and length information.

[0064] In this part, the human body model also needs to complete the change from the initial pose to the target pose. Because the input is only a photo, the target human body pose on the photo is usually different from the basic human platform. At this time, in order to fit the target human body pose, the change from the initial pose to the target pose needs to be completed. In order to more realistically simulate the fitting of several groups of bases and bone parameters in the standard three-dimensional human body parameter model, the following steps are also included.

[0065] 1) Obtain the position coordinates of the initial pose and the target pose; the initial pose parameters are determined by the initial parameters of the standard human platform model, and the bone information of the target pose is obtained by neural network model regression prediction.

[0066] 2) Generate animation sequence from initial pose to target pose; after obtaining the initial state of the skeleton information and the target state parameters, the initial pose to the target pose of the skeleton information time sequence is formed by linear interpolation, nearest neighbor interpolation and other interpolation methods. In the driving process, according to the number of bones driven by each frame, it can be divided into two ways: global linear interpolation and driving parent node first and then driving child node. Considering the driving state in the simulation of the physical world, the latter is adopted in this patent, that is, driving the parent bone node first and then driving the child bone node. The animation sequence motion interpolated in this way is more consistent with the real physical world, and the simulation effect is better.

[0067] 3) In the process of generating animation sequence, mesh interpolation method is used for processing; that is, after driving the bone movement of each frame, the current state of the human body model vertex, that is, the surface information, is calculated by the weight parameters of the standard mannequin, the current human body model mesh state is updated and saved.

[0068] 4) The setting of the interpolation speed is slow at the initial point and the target point, and fast in the middle movement process. This patent adopts a non-uniform interpolation rate, that is, the moving amplitude of each frame is small in the initial motion and the end motion, and the moving amplitude is large in the middle moving process. The starting state of the physical motion in the real world has a certain acceleration process, and the inter-frame displacement distance is kept high in the movement process, and the driving speed is reduced at the end of the movement.

[0069] 5) When driving to the final target pose, it is static for several frames to obtain the entire animation sequence. This method is closer to the motion law of the real physical world than uniform interpolation, so that the model gets an effective buffer before being static after high-speed motion, and then obtains the entire complete animation sequence, forms a dense interpolation area near the final pose, and the pose fitting accuracy of the human body model is higher, and the simulation effect is better.

[0070] 6) Complete the driving of the skeleton from the initial pose to the target pose; at this time, the standard human body model has become a human body model basically consistent with the target human body model in pose and body shape, and the pose is natural and will not be worn.

[0071] Since we have obtained the data of the skeleton and the information and data of the mesh, the driving of the skeleton will be easier in this case, and LBS algorithm (skeleton skinning), DQS algorithm can be used, of course, the collision body is also considered, because the model of the standard mannequin is in a three-dimensional standard pose, and the change from the initial pose to the target pose may cause unreasonable penetration between the meshes of the human body model. Only by combining the collision body can the mutual penetration defect between the meshes be avoided.

[0072] Combination Figures 1 to 4The method for generating a three-dimensional human body model according to an embodiment of the present invention can be implemented by a human body image processing device. Figure 5 This is a schematic diagram illustrating the hardware structure 300 of a device for processing human body images according to an embodiment of the invention.

[0073] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the fitting method and steps described above.

[0074] And an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement the fitting method and steps described above when executing the program stored in the memory.

[0075] like Figure 5 As shown, the human body fitting device 300 in this embodiment includes: a processor 301, a memory 302, a communication interface 303, and a bus 310, wherein the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication between them.

[0076] Specifically, the processor 301 may include a central processing unit (CPU), an ASIC, or one or more integrated circuits that can be configured to implement embodiments of the present invention.

[0077] Memory 302 may include a large-capacity memory for data or instructions. For example, and not limitingly, memory 302 may include an HDD, floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where suitable, memory 302 may include removable or non-removable (or fixed) media. Where suitable, memory 302 may be internal or external to the human image processing device 300. In a particular embodiment, memory 302 is a non-volatile solid-state memory. In a particular embodiment, memory 302 includes read-only memory (ROM). Where suitable, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0078] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of the present invention.

[0079] Bus 310 includes hardware, software, or both, to couple components of device 300 for processing human images to each other. For example, but not limited to, bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or other suitable bus or combination of two or more of these. Where appropriate, bus 310 can include one or more buses. Although the present embodiments describe and show a particular bus, the present embodiments contemplate any suitable bus or interconnect.

[0080] That is, Figure 5 The device 300 for processing human images shown can be implemented to include: a processor 301, a memory 302, a communication interface 303 and a bus 310. The processor 301, the memory 302 and the communication interface 303 are connected through the bus 310 and complete communication among each other. The memory 302 is used to store program codes; the processor 301 runs programs corresponding to the executable program codes by reading the executable program codes stored in the memory 302, to execute the method for fitting a three-dimensional human model in any of the embodiments of the present application, so as to realize the method for fitting a three-dimensional human model and the device described in combination with the above. Figures 1 to 4 The method and device for fitting a three-dimensional human model are described.

[0081] The embodiments of the present application further provide a computer storage medium, which has computer program instructions stored thereon; the computer program instructions are executed by a processor to implement the method for processing human images provided by the embodiments of the present application.

[0082] It needs to be made clear that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of simplicity, detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the present application.

[0083] The functional blocks shown in the structural block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, functional cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. A "machine-readable medium" includes any medium that can store or transport information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, and the like. The code segments can be downloaded via computer networks such as the Internet, intranets, and the like.

[0084] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in an order different from that in the embodiments, or several steps can be performed simultaneously.

[0085] The above description is merely a specific implementation of the present application. Those skilled in the art can clearly understand the specific working processes of the above-described system, modules and units for the convenience and brevity of description, which can refer to the corresponding processes in the foregoing method embodiments, which will not be described here. It should be understood that the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.

Claims

1. A method for fast fitting of a human model, characterized in that, The method comprises: 1) establishing a three-dimensional standard human body model in an initial pose, the three-dimensional standard human body model comprising 20 independently controlled body base parameters and 170 bone parameters, and the initial pose being selected from a plurality of preset initial pose bases generated based on distance criteria clustering; 2) obtaining three-dimensional human body pose and body shape parameters of a target human body model, the three-dimensional human body pose and body shape parameters corresponding one-to-one to the 170 bone parameters and 20 body base parameters of the three-dimensional standard human body model; the obtaining of the three-dimensional human body pose and body shape parameters of the target human body model comprising: obtaining a two-dimensional human body contour image based on a target area fast generation network based on a convolutional neural network; substituting the two-dimensional human body contour image into a first neural network subjected to deep learning to regress a joint node, and then substituting the two-dimensional human body contour image into a second neural network subjected to deep learning to regress human body pose and body shape parameters; 3) calculating a bone movement distance of each preset initial pose base to a target pose, and selecting an initial pose base of the standard three-dimensional human body model with the shortest movement distance; 4) starting from the selected initial pose base, fitting according to the obtained 20 body base parameters and 170 bone parameters of the target human body model; 5) obtaining a three-dimensional human body model with the same body shape as the target human body; 6) obtaining a three-dimensional target human body model mesh with the same pose as the target human body two-dimensional image; The fitting starting from the selected initial pose base specifically comprises: (1) obtaining pose parameters representing a bone position of the target human body model; and the two-dimensional human body contour image of the target human body is obtained through a target area fast generation network based on a convolutional neural network; (2) clustering and calculating a plurality of sets of obtained pose parameters according to a set rule, and calculating an initial pose base required to achieve the target according to the clustering calculation target, wherein the clustering calculation target is that a movement distance of each initial pose to the target pose is less than a given limit value or an interpolation frame number during fitting is less than a given limit value; (3) calculating a shortest distance from an initial pose to a target pose according to position coordinates of a plurality of preset initial pose bases; and selecting the initial pose base from the preset pose to start fitting; (4) obtaining position coordinates of the initial pose and the target pose, the initial pose parameters being determined by initial parameters of the standard human body model, and the bone information of the target pose being obtained by a neural network model regression prediction; (5) generating an animation sequence from the initial pose to the target pose in a variable-speed mesh interpolation manner; after driving the bone movement of each frame in the generation of the animation sequence, the vertex and surface information of the human body model in the current state are calculated through weight parameters of the three-dimensional standard human body model, and the mesh state of the human body model is updated and recorded; (6) stopping for a plurality of frames when driving to the final target pose to obtain the entire animation sequence; and the initial pose is driven to the target pose, and the pose and body shape of the final standard human body model are basically consistent with those of the target human body model.

2. The method of claim 1, wherein, The initial pose base of the three-dimensional standard human body model is a plurality of preset initial pose bases including a T-pose.

3. The method of claim 2, wherein, The step of obtaining the pose parameters representing the position of the target human model skeleton further comprises: 1) obtaining a two-dimensional image of a target human body; 2) processing the obtained two-dimensional human body contour image; 3) substituting the two-dimensional human body contour image into a first neural network trained by deep learning to regress the joint nodes; 4) obtaining a joint node map of the target human body; obtaining a human body part semantic segmentation map; body key points; body skeleton points; 5) substituting the generated joint node map, semantic segmentation map, body skeleton points and key point information of the target human body into a second neural network trained by deep learning to regress the human body pose and body shape parameters; and 6) obtaining output three-dimensional human body parameters, including three-dimensional target human body skeleton position pose parameters and three-dimensional human body shape parameters. Further comprising, before inputting the two-dimensional human body image into the first neural network model, a process of training the neural network, and the training samples include standard two-dimensional human body images with original joint node positions labeled by manual high-accuracy labeling on the two-dimensional human body images.

4. The method of claim 3, wherein, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-4.

5. A computer readable storage medium, characterized in that, The device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store a computer program; and the processor is used to execute the program stored in the memory to implement the method of any one of claims 1-4.

6. An electronic device, comprising: The device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store a computer program; and the processor is used to execute the program stored in the memory to implement the method of any one of claims 1-4.

Citation Information

Patent Citations

  • Universal posture recognition method based on joint vectors

    CN110084140A

  • Three-dimensional human body model reconstruction method, storage equipment and control equipment

    CN110827342A