A fusion processing method, device and storage medium for human body model images
Through the combination of self-made three-dimensional basic human station and deep learning neural network, the distortion problem of image fusion in virtual clothes changing technology is solved, and an efficient and real clothes changing effect is achieved, which is suitable for ordinary user terminal devices.
Patent Information
- Application Number
- CN202011609643.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-12-28
AI Technical Summary
When the existing virtual dressing technology generates the mannequin after the dressing and fuses with the original image, it is easy to have problems such as head aberration, loss of skin texture or inconsistent skin tone, and loss of body parts, resulting in distortion of synthetic images, and high calculation volume and high hardware cost, making it difficult to promote among ordinary users.
The homemade three-dimensional basic human platform model is adopted, combined with deep learning neural network to obtain the human model parameters, drive bone movement through the interpolation method, fabric calculation is performed frame by frame, and the limbs and skin colors are integrated and adjusted in image processing to generate a realistic dressing effect.
It realizes efficient generation of realistic dressing changes on ordinary terminal devices. The user's head image, hand image and exposed skin color are consistent with the original image. The overall effect is real and natural, avoiding obvious bugs, and improving the accuracy and computing efficiency of model fitting.
Smart Images

Figure CN114693570B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of virtual dressing, and specifically relates to an image fusion processing method after matching of a human body model and a clothing model, especially a method, device and storage medium for fusing splicing parts and exposed skin color. Background Art
[0002] With the development of internet technology, online shopping has become increasingly popular. Compared to in-store shopping, online shopping offers advantages such as a wider variety of products and greater convenience. However, online shopping also presents some difficult-to-solve challenges, the most prominent of which is the inability to physically view the items. This issue is particularly prominent in clothing. Unlike in-store shopping, where customers can change outfits and see how they look in real time, online clothing shopping offers no personalized images. Instead, they only provide pictures of models trying on clothes, or sometimes even no pictures at all. This prevents consumers from visually assessing the degree to which the clothing matches their body image in real time, leading to a high number of returns and exchanges.
[0003] To address this issue, businesses are attempting to utilize virtual fitting technology to provide consumers with simulated fitting experiences. Of course, there are other practical applications for virtual fitting technology, such as in online games. Consequently, this technology has seen rapid development.
[0004] Virtual fitting refers to a technology application that allows users to view the desired outfit in real time on a terminal screen, without having to physically put on the clothes they want to see. Existing fitting technologies primarily include two-dimensional fitting and three-dimensional virtual fitting. The former essentially captures a user's photo and clothing images, then crops and splices them to create an image of the "dressed" person. However, these images lack realism due to crude image processing methods. They completely ignore the user's actual body shape and simply force the garment onto the user's photo, failing to meet user needs. The latter typically uses 3D acquisition equipment to capture 3D information about the person and synthesize it with clothing features. Alternatively, users manually input body data, create a virtual 3D human model based on specific rules, and then combine it with clothing textures. Overall, these types of 3D virtual fitting require extensive data acquisition and 3D data calculations, resulting in high hardware costs and limited adoption for the general public.
[0005] With the development of cloud computing technology, artificial intelligence technology, and intelligent terminal processing capabilities, technologies have emerged that generate three-dimensional human models from two-dimensional human virtual images, and then implement virtual fitting by changing into three-dimensional clothing models. This type of technology mainly includes several steps: (1) processing the personal body information provided by the user to obtain the target human model; (2) processing the clothing information to obtain the clothing model; (3) matching the human model and the clothing model; (4) fusing the generated image of the person wearing the clothing with the original image.
[0006] The first three steps primarily complete the dressing process. However, due to limitations in the dressing principle, the resulting human model does not possess the head and facial features of the person being dressed. Therefore, after generating the dressing model, a two-dimensional image of the dressing model must be superimposed onto the original two-dimensional image. This allows the synthesized image to display the person's head and body in virtual clothing. Consequently, any defects in the superposition will lead to distortion and irrationality in the image. This can easily lead to problems such as misalignment of the human model's head with the person being dressed in the original image, loss of skin texture or skin color, loss of body parts, and loss of background, which in turn affect the final effect of the generated dressing image.
[0007] In the general field of computer vision, there are many initial steps in human body modeling, which usually include using 3D scanning equipment to perform all-round scanning of the real human body, 3D reconstruction methods based on multi-view depth-of-field photography, and 3D reconstruction methods based on single or multiple images combined with neural network models and standard human body models.
[0008] In the prior art, there are several common methods for constructing a human body model from a two-dimensional image: (1) regression-based methods, which reconstruct a voxel-based human body model using a convolutional neural network; (2) single-image-based human body reconstruction, which first roughly annotates simple human skeletal key points on the image, and then performs initial matching and fitting of the human body model based on these rough key points to obtain the approximate shape of the human body; (3) CNN-based methods predict key points on the image, and then use the SMPL model for fitting to obtain an initial human body model. Finally, the initial model and the bounding box obtained by regression are combined to obtain a three-dimensional human body reconstruction.
[0009] Prior art 1 discloses a human body modeling method based on anthropometry data, comprising: obtaining anthropometry data; performing linear regression on a pre-created human body model using a pre-trained prediction model based on the anthropometry data to obtain a predicted human body model, wherein the pre-created human body model includes multiple pre-defined groups of marked feature points and corresponding standard shape bases, and the anthropometry data includes measurement data corresponding to each group of marked feature points; and obtaining a target human body model based on the predicted human body model, wherein the target human body model includes the measurement data, a target shape base, and a target shape coefficient. However, this method has very high requirements for anthropometry data. Although it saves computational effort, it provides a poor user experience and is very cumbersome.
[0010] It should be noted that the SMPL model is a parametric human body model, a human body modeling method proposed by the Max Planck Institute in Germany. This method can model and animate any human body. The key difference between this method and traditional location-based body-building (LBS) lies in its proposed method for analyzing the surface topography of human pose images, which can simulate the convexity and concavity of human muscles during limb movement. This avoids surface distortion during human motion and accurately depicts the topography of muscle extension and contraction. In this method, β and θ are input parameters. β represents 10 parameters related to a person's height, weight, head-to-body ratio, and other proportions, while θ represents 75 parameters representing the overall motion pose and the relative angles of 24 joints. However, this model generation method relies on accumulating a large amount of training data to obtain the relationship between body shape and shape bases. However, due to the strong interdependence between these two shapes, independent control of each shape base is difficult, making decoupling difficult. For example, the arms and legs are also related to each other; theoretically, when the arms move, the legs will also move. This makes it difficult to improve the SMPL model for different body types. During the driving process of the model, the characteristics of this model still seriously affect the final driving effect of the model. If it is to move, it moves as a whole, which is significantly different from the human body model used in the present invention in which each part can be independently controlled. There is still a lot of room for improvement in the model driving effect.
[0011] Prior art 2 discloses a method for processing a virtual fitting model image, comprising: determining a set of pixels in a facial region whose colors fall within a skin color range; using the color average of all pixels in the set as the average skin color value of a preselected region in a reference image, and calculating a ratio obtained by dividing the average skin color value by the average skin color value of the virtual fitting model image; multiplying each pixel value in the body region of the virtual fitting model image by the ratio, and using the multiplication result as each pixel value in the body region of the virtual fitting model image; and calculating the average skin color value of the virtual fitting model image before calculating the ratio obtained by dividing the average skin color value by the average skin color value of the virtual fitting model image. Before determining the average skin color value of the preselected region in the reference image, the method further comprises receiving information used to determine the preselected region in the reference image. This method can simplify complex calculations, but it only considers the average skin color value of the facial region and uses this value instead of the skin color values of all other parts, which can easily cause the new skin color to appear unnatural.
[0012] Prior art three discloses a virtual object synthesis method, including: obtaining a target user image and a virtual object image containing clothing features; extracting user head features from the target user image; performing skin color processing on the user head features and the virtual object image according to the skin color features of a reference object image, respectively, to obtain user head features and a virtual object image that match the skin color features; obtaining neck features of the reference object image, and fusing them into the virtual object image after skin color processing; integrating the user head features after skin color processing into the virtual object image fused with the neck features, to synthesize a virtual object containing the clothing features, the neck features and the user head features after skin color processing. Based on the skin color features of the reference object image, the user's head features and the virtual object image are processed for skin color to obtain the user's head features and virtual object image that match the skin color features. The method includes: obtaining pixel values for each pixel representing skin color from the user's head features and the virtual object image, and obtaining a pixel matrix based on the obtained pixel values; processing the skin color features of the reference object image to obtain a skin color mapping matrix; performing operations on the skin color mapping matrix and the pixel matrix corresponding to the user's head features to obtain user's head features that match the skin color features; and performing operations on the skin color mapping matrix and the pixel matrix corresponding to the virtual object image to obtain a virtual object image that matches the skin color features. This method is actually a head replacement operation, virtually replacing the model's head. Therefore, it only extracts the skin color features of the head and cannot fully reflect the conditions of other parts of the body.
[0013] As can be seen, in the field of virtual clothing changing, when a human model is transformed from an initial position to a target pose and already dressed, the resulting synthetic 2D image still has many visually visible defects. Currently, most methods focus on improving processing speed to meet the requirements of animation or games. However, in some cases, the quality of the final clothing simulation is more important. Therefore, it is urgent to find a method for fusion processing of human model and clothing model images that is computationally intensive and can produce excellent results without exceeding the end-user's tolerance. Summary of the Invention
[0014] To address these challenges, the present invention provides a method, device, and storage medium for fusion processing of human model images that overcome these issues. By repairing and fusing the joints of the superimposed images and extracting the skin color of each part of the original 2D image, the original image can be effectively fused with the 2D image output after the human model is dressed. This allows the resulting composite image to more closely resemble the original image in both form and realism.
[0015] The present invention provides a fusion processing method for a human body model image, the method comprising: producing a three-dimensional basic human form in an initial posture, wherein the initial posture parameters are determined by initialization parameters of the basic human form model; obtaining a three-dimensional clothing model; fitting the three-dimensional clothing model to the three-dimensional basic human form model in the initial posture; using a neural network model to obtain secondary information of the human form model based on a two-dimensional image of a target human body; regressively predicting the neural network model based on the secondary information to obtain posture and body shape parameters of the target human form model, wherein the three-dimensional human form and body shape parameters correspond to the skeleton and a plurality of basis parameters of the three-dimensional basic human form; inputting the obtained plurality of groups of basis and skeleton parameters into the basic human form model for fitting to obtain a target posture and a target body shape; driving the skeleton of the human form model to move from the initial posture to the target posture; obtaining a three-dimensional target human form model having the same posture as the two-dimensional image of the target human body and having completed the three-dimensional clothing model change; fusing a head portrait with the image of the target human form model, processing parts of the limbs that are not fully aligned and aligned; restoring the skin color of exposed parts of the human body; and obtaining a 2D image after the clothing change, which consists of the head portrait and limbs of the target human form, the background of the target image, and the clothes worn by the target human form model.
[0016] Preferably, the fusion process further includes: (1) performing a portrait cutout process on the two-dimensional image of the target human body, but retaining the head image; (2) outputting the three-dimensional human body model after the change of clothes into a two-dimensional image having the same size as the two-dimensional human body image without any background and without the head image; (3) superimposing the two-dimensional image output by the human body model on the target two-dimensional human body image, and checking whether the head image of the target human body is consistent with the neck of the two-dimensional image output by the human body model; (4) if not consistent, performing a splicing process on the joint parts by using pixel fusion.
[0017] Preferably, the part where the human body model outputs the two-dimensional image and is connected to the head of the original two-dimensional image is fused and adjusted according to the detected skin color of the face; the other exposed parts of the human body model output the two-dimensional image are fused and adjusted according to the skin color of the exposed skin in the similar parts in the original two-dimensional image.
[0018] Preferably, if the superimposed image contains missing pixels after the cutout, the missing pixels are filled according to the background pixel color in a certain area around the missing pixels.
[0019] Preferably, the driving process also includes: driving the human body model skeleton to move from the initial posture to the target posture by using interpolation; the movement of the skeleton drives the follow-up movement of the human body model mesh; driving the clothing model to follow the human body model to move to the target posture frame by frame; the clothing model cloth solution process is synchronized with the human body model mesh movement process, and the physical simulation calculation of the cloth is performed after all bones complete the movement of each frame.
[0020] Preferably, after obtaining the initial state of the skeleton information and the target posture state parameters, the skeleton is driven to move from the initial posture to the target posture, and a skeleton information time series from the initial posture to the target posture is formed by interpolation using linear interpolation or nearest neighbor interpolation.
[0021] Preferably, in the process of generating the animation sequence, the movement of the human body model mesh is also performed in an interpolation manner. After each frame drives the skeleton movement, the current state of the human body model vertex, i.e., face information, is calculated using the weight parameters of the standard human body model, and the current human body model mesh state is updated and recorded and saved.
[0022] Preferably, the step of obtaining the parameters of the target human body model also includes: 1) obtaining a two-dimensional image of the target human body; 2) processing to obtain a two-dimensional human body contour image of the target human body; 3) substituting the two-dimensional human body contour image into a first neural network that has undergone deep learning to perform joint point regression; 4) obtaining a joint point map of the target human body; obtaining semantic segmentation maps of various parts of the human body; body key points; body bone points; 5) substituting the generated joint point map, semantic segmentation map, body bone points and key point information of the target human body into a second neural network that has undergone deep learning to perform human body posture and body shape parameter regression; 6) obtaining output three-dimensional human body parameters, including three-dimensional human body motion posture parameters and three-dimensional human body shape parameters.
[0023] In addition, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned method steps is implemented.
[0024] An electronic device includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement any of the method steps described above when executing the program stored in the memory.
[0025] The beneficial effects of the present invention are:
[0026] 1. The final image is excellent. Our clothing-changing program essentially converts a 2D image back to a 2D image. That is, the user enters an original 2D image of themselves, selects the clothing they wish to change into, and without any additional steps, they receive a photo of the user after changing their clothing. The user's head image, hands, exposed skin color, and background are all consistent with the original 2D image; the only difference is the selected clothing. Therefore, to achieve this fully restored and highly natural image quality, we added a superimposition and fusion process after the mannequin changes clothes. First, we adjust any misaligned limbs, then unify the skin tone of the head and neck, and finally, unify the skin tone of the hands and feet. This results in a more realistic and natural-looking composite image, without noticeable bugs.
[0027] 2. Good model fitting effect and less model penetration. Traditionally, in order to drive the model to fully fit the posture of the target human body, in the process of changing from the initial pose to the target pose, a one-step overall movement method is usually adopted. This method has a small amount of calculation, but the model boundary conditions and position conditions change greatly, the realism is greatly reduced, and it is easy to cause model penetration. We noticed that when the human body model moves, the mesh movement patterns of various parts of the limbs are actually different. Some meshes change more obviously, while others do not change obviously. Correspondingly, some meshes move violently, while others move almost nothing. If a one-step method is used to reach the target position from the initial position, some parts will have obvious distortion or deformation, while some parts will have a reasonable degree of restoration. The overall visual effect is relatively poor, and it is easy for people to see various unreasonable details. To address this characteristic and realistically fit the target human's motion posture, better aligning with fabric simulation, we employed an optimized interpolation method to complete the skeleton's transition from initial to target pose. Compared to traditional interpolation, our target pose's skeletal information is predicted by model regression, while simultaneously generating an animation sequence that moves from the initial to target pose. Through interpolation methods like linear interpolation and nearest neighbor interpolation, a time series of skeletal information from the initial to target pose is formed. We creatively utilize the accumulation of several frames of animation sequences, avoiding a one-step driving approach. While driving the human model slightly at each frame, we perform fabric simulation frame by frame. The overall garment model solution is far superior to the one-step method. While speed is compromised, the simulation is more realistic and accurate.
[0028] 3. The human body model is precise and controllable. Our custom-made basic mannequin boasts greater control over detail during frame-by-frame processing, resulting in superior detail fidelity and fidelity compared to traditional human models. Currently popular single-image human body reconstruction methods primarily reconstruct parametric human body models, such as the SMPL model. These methods rely on deep learning and training using a large number of human body model instances. The relationship between body shape and shape basis is a holistic association, making decoupling difficult. This prevents arbitrary control of desired body parts, resulting in the generated model failing to achieve high consistency with the real human pose and shape. Furthermore, if applied to the subsequent dressing process, the ability to represent surface geometric details is limited, making it difficult to reconstruct the detailed texture of clothing on the human body. However, our human body model is not trained, and its parameters have a mathematically based correspondence. In other words, our parameter groups are independent of each other and independent of each other. This makes our model more interpretable during transformations and better characterizes the shape and position changes of specific body parts or regions. In other words, during the frame-by-frame movement of the human skeleton and mesh, our self-made basic mannequin can better reproduce the limb movement in the real world and the state of clothing following the movement.
[0029] 4. High-frequency and creative use of hierarchical deep neural networks. Neural network models are also used in the prior art, but due to differences in input conditions, input parameters, and training methods, the functions and effects of neural network models vary greatly. In terms of obtaining secondary information and model body data of the human body model, the present invention uses different neural networks for different purposes. By using neural network models with different input conditions and training methods, it achieves accurate contour separation of the human body in a complex background, semantic segmentation of the human body, determination of key points and joints, and eliminates the influence of loose clothing and hairstyle, so as to approximate the real body shape and form of the human body to the greatest extent. It fully utilizes the advantages of deep learning networks and can restore the posture and shape of the human body with high precision in various complex scenes. In addition, the parameters output by the latter level neural network include two categories: posture pose and shape shape, which can control the action and shape respectively, and combined with our reference model, the posture and shape of the human body model can be accurately replicated.
[0030] The present invention forms an animation series from an initial posture to a target posture, and uses a frame-by-frame driving method to complete the entire process of the skeleton driving the human body's epidermal mesh to move. After the model is completed, the image processing method is used to optimize the two-dimensional image generated, so that the entire dressing effect reaches a high level. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0032] Figure 1 A schematic diagram of the process of fusing a complete human body model and a clothing model according to an embodiment;
[0033] Figure 2 A schematic diagram of a process for obtaining a target human body model and a clothing model according to an embodiment;
[0034] Figure 3 A schematic diagram of a processing flow of a model parameter acquisition module according to an embodiment;
[0035] Figure 4 Schematic diagram of the system of the present invention. DETAILED DESCRIPTION
[0036] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In order to make the objects, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present invention by illustrating examples of the present invention.
[0037] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0038] The following is combined with Figure 1-3 The method described in the embodiment of the present invention is described in detail.
[0039] like Figure 1-3 As shown, the present invention provides a method for fusion processing of a human body model image, the method comprising: producing a three-dimensional basic mannequin in an initial posture, wherein the initial posture parameters are determined by initialization parameters of the basic mannequin model; obtaining a three-dimensional clothing model; fitting the three-dimensional clothing model to the three-dimensional basic mannequin model in the initial posture; using a neural network model to obtain secondary information of the mannequin based on a two-dimensional image of a target human body; regressively predicting the target mannequin's posture and body shape parameters based on the secondary information using the neural network model, wherein the three-dimensional human body posture and body shape parameters correspond to the skeleton and a plurality of basis parameters of the three-dimensional basic mannequin; inputting the obtained plurality of groups of basis and skeleton parameters into the basic mannequin model for fitting to obtain a target posture and a target body shape; driving the skeleton of the mannequin to move from the initial posture to the target posture; obtaining a three-dimensional target mannequin having the same posture as the two-dimensional image of the target human body and having completed the three-dimensional clothing model change; fusing a head portrait with the target mannequin, processing parts of the limbs that are not fully aligned and aligned; restoring the skin color of exposed parts of the human body; and obtaining a 2D image of the target mannequin after the change of clothing, which consists of the head portrait and limbs of the target human body, the background of the target image, and the clothes worn by the target mannequin. Of course, these steps do not necessarily follow a strict order, because some steps are independent preparatory steps themselves, and the order of placement does not have a decisive impact on the final result.
[0040] As can be seen, the process of the present invention generally involves five steps: first, generating a 3D base mannequin (standard human body model); second, fitting a 3D clothing model onto the base mannequin; third, obtaining the parameters of the target pose human body model; fourth, fitting the standard human body model to match the shape and pose of the target human body model, and moving the clothing model to the position of the target human body model; and fifth, fusing the 2D image output by the human body model with the original 2D image of the target human body.
[0041] The first step is to pre-design and model some basic mannequins. The main work involves constructing a three-dimensional basic mannequin, also known as a basic mannequin, by combining mathematical models. The Max Planck Institute's SMPL mannequin avoids surface distortion during human motion and accurately depicts the morphology of muscle extension and contraction. In this method, β and θ are the input parameters. β represents 10 parameters related to a person's height, weight, head-to-body ratio, and other proportions, while θ represents 75 parameters representing the overall motion pose and the relative angles of 24 joints. The β parameter is a shape blend pose parameter that can be used to control the shape changes of the human body using 10 incremental templates. Because the SMPL mannequin is ultimately trained to conform to Western body shapes using Western photographs and measurement data, its shape changes generally conform to the typical curves of Westerners. Applying it to Asian mannequins can lead to many problems, such as arm-leg proportions, waist-to-body ratios, neck ratios, and leg and arm lengths.
[0042] To this end, we employed a custom human model to enhance technical feasibility. The core of this approach is to construct a self-built blend body base to achieve precise, independent manipulation of the human body. The three-dimensional basic mannequin features a mathematically weighted relationship between skeletal points and a model mesh. The determination of skeletal points can be correlated with the human model to determine the target human pose. The three-dimensional basic mannequin is defined by several body base parameters and several skeletal parameters. These body bases comprise the entire human model mesh, with each body base independently controlled by its own base parameters, without interfering with each other. Preferably, the three-dimensional basic mannequin (basic mannequin) consists of 20 body base parameters and 170 skeletal parameters. Precise manipulation, on the one hand, involves increasing the number of control parameters, rather than relying on the Max Planck Institute's ten β control parameters. This allows for adjustable parameters, in addition to the usual body shape, to include arm length, leg length, waist, hip, and chest shape. This more than doubles the number of skeletal parameters, significantly expanding the range of adjustable parameters and providing a solid foundation for refined basic mannequin design. Independent manipulation means that each element, such as the waist, legs, hands, and head, can be manipulated independently. Each bone can also be adjusted in length independently of the others, without creating any physical interaction. This allows for precise adjustments to the human body model. Our existing model reflects a mathematical correspondence, significantly different from the SMPL human body model trained on big data. Therefore, our parameter transformations are more interpretable and can better characterize local changes in the human body model. Furthermore, these changes are based on mathematical principles, with no interaction between parameters, and the arms and legs remain completely independent.
[0043] The second part primarily involves generating a 3D clothing model. Several different methods exist in the existing art for generating 3D clothing models. Currently, the more traditional method for creating 3D clothing models is based on the design and stitching of 2D clothing pieces. This method requires specialized clothing knowledge to design the pattern. Another relatively novel 3D modeling method is based on hand-drawing, which can generate a simple clothing model from user-drawn line information. Yet another method uses image processing and graphics simulation techniques based on garment image information to ultimately generate a virtual 3D clothing model. Contour detection and classification are used to obtain the garment's outline and dimensions from the image. Machine learning methods are used to identify key points between edges within the outline. Stitching information is generated based on the key point correspondences. Finally, the garment is physically stitched together in 3D space to simulate the garment's realistic appearance on the human body. Other methods include mapping and mathematical model simulation. This aspect is not specifically limited in the present invention. However, the 3D clothing model must be matched to a standard human body model. The general requirement is that a garment model that has already been fitted to a standard human body model is adapted to the human body model in the target pose through fabric physics simulation, ensuring the garment's naturalness and appropriateness.
[0044] The third part is to process the acquired human body images to obtain the parameter information required to generate the human body model. In the past, the selection of these skeletal key points was usually done manually, but this method is very inefficient and does not adapt to the fast-paced requirements of the Internet era. Therefore, in today's era when neural networks are popular, it has become a trend to use deep learning neural networks instead of manual key point selection. However, how to use neural networks efficiently is a problem that requires further research. Generally speaking, we adopted the idea of a secondary neural network plus data "fine-tuning" to construct our parameter acquisition system. Figure 2-3 As shown in the figure, we use a deep learning neural network to generate these parameters, which mainly includes the following sub-steps: 1) obtain a two-dimensional image of the target human body; 2) process and obtain a two-dimensional human body contour image of the target human body; 3) substitute the two-dimensional human body contour image into the first deep learning neural network for joint point regression; 4) obtain a joint point map of the target human body; obtain semantic segmentation maps of various parts of the human body; body key points; body bone points; 5) substitute the generated joint point map, semantic segmentation map, body bone points and key point information of the target human body into the second deep learning neural network for human posture and body shape parameter regression; 6) obtain the output three-dimensional human body parameters, including three-dimensional human body motion posture parameters and three-dimensional human body shape parameters.
[0045] Before inputting the two-dimensional human image into the first neural network model, a neural network training process is also included. The training sample includes a standard two-dimensional human image with the original joint point locations manually annotated with high accuracy on the two-dimensional human image. Here, a target image is first acquired and human body detection is performed on the target image using a target detection algorithm. Human body detection does not mean using a measuring instrument to detect a real human body. In this invention, it refers to a given image, typically a two-dimensional photograph, that contains sufficient information, such as a face, limbs, and body. A specific strategy is then used to search the given image to determine whether a human body is present. If a human body is present, parameters such as the human body's location and size are determined. In this embodiment, before obtaining the human body key points in the target image, human body detection is performed on the target image to obtain a human body frame that annotates the human body's location. Because the input image can be any image, some non-human background elements, such as tables, chairs, trees, cars, and buildings, are inevitably present. This unused background is removed using a sophisticated algorithm.
[0046] At the same time, we also need to perform semantic segmentation, joint detection, skeleton detection, and edge detection. By collecting these 1D point information and 2D surface information, we can lay a good foundation for generating a 3D human body model later. Use the first-level neural network to generate a human body joint map. Optionally, the target detection algorithm can quickly generate a network for the target area based on the convolutional neural network. This first neural network requires a large amount of data training. The joints of some photos collected from the Internet are manually annotated and then input into the neural network for training. After deep learning, the neural network can basically obtain a joint map with the same accuracy and effect as manually annotated joints immediately after inputting the photo, and the efficiency is dozens or even hundreds of times that of manual annotation.
[0047] After obtaining the relevant 1D point information and 2D surface information, these parameters or results, the target human body's joint point map, semantic segmentation map, body bone points and / or key point information can be substituted as input items into the second neural network that has undergone deep learning to regress human body posture and body shape parameters. After the regression calculation of the second neural network, several groups of three-dimensional human body parameters can be immediately output, including three-dimensional human body action posture parameters and three-dimensional human body shape parameters. Preferably, the loss function of the neural network is designed based on the three-dimensional basic mannequin (basic mannequin), the predicted three-dimensional human body model, the standard two-dimensional human body image with the original joint point positions marked, and the standard two-dimensional human body image including the predicted joint point positions.
[0048] The fourth part is to fit the parameters of the human body model with the human body model and match the clothing model with the human body model.
[0049] Preferably, the driving process also includes: driving the human body model skeleton to move from the initial posture to the target posture by using interpolation; the movement of the skeleton drives the follow-up movement of the human body model mesh; driving the clothing model to follow the human body model to move to the target posture frame by frame; the clothing model cloth solution process is synchronized with the human body model mesh movement process, and the physical simulation calculation of the cloth is performed after all bones complete the movement of each frame.
[0050] To realistically capture the target human's motion and better complement fabric simulation, we devised an interpolation method to achieve high fidelity and fidelity in the clothing model simulation mechanism during the skeleton's transition from the initial pose to the target pose. Because the human body model's actuation involves repeated calculations and verifications, a smaller amount of computation leads to closer simulation results. Conversely, a larger computational span leads to rapidly increasing model distortion. Our invention breaks down the movement from the initial to the target position into several small movements. An animation sequence from the initial to the target pose is generated by interpolating frames in chronological order. Interpolation methods can include linear interpolation and nearest neighbor interpolation, forming a time series of skeletal information from the initial to the target pose. This creates a series of chronologically ordered animations. By decomposing the entire motion into several frames, the model's movement in each frame is very small, ultimately accumulating to form a large-scale motion, avoiding the need for a one-step actuation method. By interpolating frames between the initial and target poses, the human body model is slightly driven in each frame while fabric solution, collision body calculation, and verification are performed frame by frame. As a result, by the time the target pose is reached, the overall solution of the garment model is far superior to that of a one-step approach, resulting in a more realistic and faithfully rendered garment model.
[0051] Preferably, after obtaining the initial state of the skeleton information and the target pose state parameters, the skeleton is driven to move from the initial pose to the target pose, and a time series of skeleton information from the initial pose to the target pose is formed through linear interpolation or nearest neighbor interpolation. This method ensures that the movement of all skeletal joints meets the requirement of a small motion amplitude, making the changes in movement more consistent with the actual situation and facilitating subsequent human body mesh tracking and small changes in the clothing model.
[0052] Preferably, in the process of generating the animation sequence, the movement of the human body model mesh is also performed in an interpolation manner. After each frame drives the skeleton movement, the current state of the human body model vertex, i.e., face information, is calculated using the weight parameters of the standard human body model, and the current human body model mesh state is updated and recorded and saved.
[0053] The fifth part is to fuse the two-dimensional image output by the human body model and the original target human body two-dimensional image, which is also the core part of the present invention.
[0054] The entire process is essentially a process of going from a 2D image to a 2D image. In other words, a user enters an original 2D image of themselves, selects the clothing they want to change into, and instantly receives a composite photo of the user in their new clothes. The user's head, hands, exposed skin color, and background are all identical to the original 2D image; the only difference is the clothing they've selected.
[0055] Preferably, the fusion process further includes: (1) performing a portrait cutout process on the two-dimensional image of the target human body, but retaining the head image; (2) outputting the three-dimensional human body model after the change of clothes into a two-dimensional image having the same size as the two-dimensional human body image without any background and without the head image; (3) superimposing the two-dimensional image output by the human body model on the target two-dimensional human body image, and checking whether the head image of the target human body is consistent with the neck of the two-dimensional image output by the human body model; (4) if not consistent, performing a splicing process on the joint parts by using pixel fusion.
[0056] Here, image stitching can be simply implemented in the following steps: (1) extract feature points from each image; (2) match the feature points; (3) perform image registration; (4) copy the image to a specific location in another image; and (5) perform special processing on the overlapping boundaries. Image stitching is achieved through image registration and image fusion. It can usually effectively resolve the slight overlap of two images in the same frame and ultimately stitch them together into a multi-image, high-resolution image. Among them, image registration mainly solves the problem of converting two images at their respective coordinates into a single image at the same coordinates, and image fusion mainly solves the grayscale value problem of the pixels in the stitched image.
[0057] Due to defects in the hardware design itself, many different noises make the captured images fail to meet the image quality requirements. Therefore, it is necessary to perform image preprocessing such as denoising and correction on the original images. The accuracy of the image preprocessing stage has a great impact on the quality of the final stitched image. The main purpose of preprocessing is to enhance the image details, suppress noise, and improve image quality. Common preprocessing methods include the following:
[0058] (1) Image smoothing and edge sharpening. Due to the different shooting angles, folding transformations and random noise of the images, the details of the overlapping parts of the image sequence with overlapping areas are not exactly the same. Therefore, only the contour or other main edges can be selected as vertical edges for feature matching. (2) Phase correlation algorithm. If the image has translation, the translation can be converted to the frequency domain and the phase difference can be calculated. The pulse on the translation motion coordinate is the Fourier inverse transform of this phase difference. After the displacement position of the two images, the alignment point of the two images can be obtained by searching for the maximum value. (3) Grayscale image projection algorithm. If the vertical translation can be ignored and the horizontal translation is small, the grayscale image projection algorithm can be used to roughly locate the two adjacent images so as to reduce the error and narrow the search range when performing precise registration. First, a color image is converted to grayscale, and then converted into a grayscale image of a binary image. The grayscale values of all pixels are calculated and then projected to the vertical direction. Accumulation is expected. By comparing the adjacent curves, the projection of the position image can be roughly matched. (4) Screening of video sequence subsets. When performing video-based image stitching, the video sequence images need to be screened first. Since video sequence images have ample overlapping information that can be used because the displacement between them is very small. Therefore, in order to reduce the registration error and discontinuity of the stitched image, as well as to reduce the amount of calculation, only a subset of it can be selected instead of using all the video sequence images. (5) Algorithm based on template matching. The process based on template matching is to use a piece in the overlapping area of one image as a template, and search for a corresponding block with the same or similar value as the template in the other image, so that the overlapping range of the two images can be determined. Generally speaking, the larger the template area, the higher the accuracy of this algorithm, but its computational complexity will also be high.
[0059] The key to image stitching is to accurately locate the overlapping portions of two adjacent images and then determine the transformation relationship between the two images, a process known as image registration. Feature-based image registration methods are generally used because they are insensitive to brightness and noise and can handle large misalignments between images.
[0060] The image registration scheme consists of two stages: pre-registration and registration. First, in the pre-registration stage, a training set is generated by performing a perspective transformation on the reference image. Feature coefficients are extracted from the training set using the SIFT method, and these coefficients are then fed into the SLFN for training. Second, the output of the trained SLFN is the perspective transformation parameters. Because the SLFN has already been trained, in the registration stage, we simply use the same feature extraction method to extract feature coefficients from the registered image. These coefficients are then fed into the trained SLFN to obtain estimated perspective parameters. In other words, the reference image can be considered a "basis." A training set is generated by performing a perspective transformation on the reference image. At this point, the perspective transformation parameters of the training set relative to the reference image are already known. For example, an image in the training set is translated by 3 pixels on the x-axis, 2 pixels on the y-axis, and rotated by 20 degrees, resulting in a horizontal distortion of 0.001 and a vertical distortion of -0.003.
[0061] The training set consists of 200 images, and each image after transforming the reference image a by perceptual parameters within a predefined range is shifted to the left to make it most similar to image b, so that the perspective transformation parameters can be obtained.
[0062] Perspective transformation is one of the most common and complex transformations between 2D images. It accounts for all possible motion patterns between images. It can describe translation, scaling, rotation, horizontal and vertical deformations, and more. Because the reference and registered images are in their own pixel coordinate systems, we need to convert them to the same pixel coordinate system.
[0063]
[0064] Where (x1, y1) is transformed into (x2, y2) through perspective, where (x1, y1) are the coordinates of the reference image and (x2, y2) are the coordinates of the transformed image in the registered image. H is the perspective transformation matrix with eight non-zero parameters, which basically covers all possible motion modes between images, such as translation, rotation, and scaling.
[0065] Generally speaking, feature-based registration methods use features such as distinct blocks, lines, and points in the images to estimate the transformation matrix between the images. The general steps for image registration using this method are: (a) extracting features from the images to be registered; (b) matching the image features; (c) estimating the transformation matrix between the images using the matched features; and (d) aligning the images using the transformation matrix.
[0066] Feature detection is the foundation of image registration. Different feature detection methods should be selected based on the scene characteristics of the images to be stitched. Common edge detection methods include the Roberts operator, Sobel operator, Prewitt operator, and Canny operator. The most commonly used corner detection method is the Harris corner detection algorithm.
[0067] Image fusion is a technique that combines useful information from two or more registered images into a single image and displays it visually. Due to differences in resolution, viewing angle, and lighting, the registered images, and sometimes even the stitching of multispectral images, can produce blur, ghosting, or noise in the overlapping areas of the stitched images, and noticeable seams may form at the edges. Fusion of the stitched images is necessary to improve their visual quality and objective quality.
[0068] At present, fusion algorithms can be roughly divided into: fusion algorithms based on image grayscale, fusion algorithms based on color space changes, and fusion algorithms based on transform domains. (1) Fusion algorithms based on image grayscale. The weighted average method is the simplest image fusion algorithm. The corresponding pixels of the two images are multiplied by a weighting coefficient and then added to obtain a fused image; (2) The image fusion algorithm based on the region of interest can be regarded as an adaptive weighted average method. The region of interest of one image is segmented, its weighting coefficient is set to 1, and the weighting coefficient of the corresponding region of the other image is set to 0, that is, the region of interest of one image is embedded in the other image; (3) The contrast modulation method uses the image detail information contained in one image to extract its contrast, modulate the grayscale distribution of the other image, and achieve image fusion. These methods can achieve good image fusion and alignment.
[0069] Pixel-level image fusion is the most basic fusion method. The image obtained after pixel-level image fusion has more detailed information, such as edge and texture extraction, which is conducive to further analysis, processing and understanding of the image. It can also expose potential targets and facilitate the judgment and identification of potential target pixels. This method can preserve as much information as possible in the source image, so that the fused image has increased content and details. This advantage is unique and only exists in pixel-level fusion.
[0070] Preferably, the 2D image output by the human model is blended and adjusted based on the detected facial skin color of the head. The 2D image output by the human model of other exposed limbs is blended and adjusted based on the skin color of the exposed skin in the adjacent areas in the original 2D image. This is primarily to ensure that the skin color of the synthesized image matches the original image, as hands and some exposed limbs may remain exposed in the synthesized image, particularly the hands and neck. However, due to varying clothing styles, it is possible that newly generated human models will have additional exposed skin after changing into new clothing. Therefore, the skin color of the newly generated human model must be consistent with the exposed skin color in the original image; otherwise, significant color difference will result, significantly reducing the overall image fidelity. Preferably, if the superimposed image contains missing pixels after cropping, the missing pixels are filled in based on the background pixel color within a certain area surrounding the missing pixels. This is because changing clothing can leave pixels in the background previously covered by clothing unfilled. Therefore, to maintain a natural-looking background, the filling should be based on the surrounding color.
[0071] Combine Figures 1 to 3 The described fusion method according to the embodiment of the present invention may be implemented by a device for processing human body model fitting and fusion. Figure 4 FIG. 3 is a schematic diagram showing a hardware structure 300 of a device for processing human body model fitting and fusion according to an embodiment of the present invention.
[0072] The present invention also discloses a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the aforementioned methods and steps are implemented.
[0073] And an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement the above-mentioned method and steps when executing the program stored in the memory.
[0074] like Figure 4 As shown, the device 300 in this embodiment includes: a processor 301, a memory 302, a communication interface 303 and a bus 310, wherein the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication between them.
[0075] Specifically, the processor 301 may include a central processing unit (CPU) or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.
[0076] Memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, memory 302 may include an HDD, a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 302 may include removable or non-removable (or fixed) media. Where appropriate, memory 302 may be inside or outside of processing device 300. In a specific embodiment, memory 302 is a non-volatile solid-state memory. In a specific embodiment, memory 302 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0077] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiment of the present invention.
[0078] The bus 310 includes hardware, software, or both that couples the components of the processing device 300 to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industrial Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-x) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 310 may include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.
[0079] That is to say, Figure 4 The device 300 shown can be implemented as including: a processor 301, a memory 302, a communication interface 303 and a bus 310. The processor 301, the memory 302 and the communication interface 303 are connected via the bus 310 and communicate with each other. The memory 302 is used to store program code; the processor 301 reads the executable program code stored in the memory 302 to run the program corresponding to the executable program code to execute the fusion method in any embodiment of the present invention, thereby realizing the fusion method. Figures 1 to 3 Methods and apparatus are described.
[0080] An embodiment of the present invention further provides a computer storage medium having computer program instructions stored thereon; when the computer program instructions are executed by a processor, the fusion method provided by the embodiment of the present invention is implemented.
[0081] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0082] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0083] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.
[0084] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.
Claims
1. A method for fusion processing of a human body model image, the method comprising: 1) Producing a three-dimensional basic avatar with multiple independently controlled body base parameters and multiple skeletal parameters, wherein the three-dimensional basic avatar is constructed based on a mathematical model, each body base is independently controlled by a base parameter, and the skeletal parameters are independently adjusted; 2) Obtaining a three-dimensional clothing model; 3) fitting the 3D clothing model to the 3D base mannequin model in the initial pose; 4) using the first neural network model to extract secondary information including a 2D human body contour, a semantic segmentation map, skeleton points, and body key points based on the target human body 2D image; 5) inputting the secondary information into a second neural network model that has undergone deep learning, and using the second neural network model to regress and predict to obtain three-dimensional human body posture and body shape parameters, wherein the three-dimensional human body posture and body shape parameters correspond to the skeleton and a plurality of basic parameters of the three-dimensional basic mannequin; 6) Inputting the obtained basic parameters and bone parameters into the basic mannequin model for fitting to obtain the target posture and target body shape; 7) Using a frame-by-frame interpolation method, the skeleton of the human body model is driven to transform from the initial posture to the target posture frame by frame, and the clothing model is driven to move to the target posture frame by frame along with the human body model. The clothing model fabric solution process is synchronized with the human body model mesh movement process. After all the skeletons complete the movement of each frame, the physical simulation calculation of the fabric is performed; 8) obtaining a three-dimensional target human body model having the same posture as the two-dimensional image of the target human body and having completed the three-dimensional clothing model change; 9) Fuse the 2D head portrait output by the human body model with the target human body image to process parts where the limbs are not fully aligned; The fusion process includes: (1) performing a portrait cutout process on the target human body two-dimensional image, but retaining the head image; (2) completing the three-dimensional human body model after changing clothes, and outputting it as a two-dimensional image with the same size as the two-dimensional human body image without any background and without the head image; (3) superimposing the two-dimensional image output by the human body model on the target two-dimensional human body image, and checking whether the head image of the target human body is consistent with the neck of the two-dimensional image output by the human body model; (4) if not consistent, using pixel fusion to perform splicing processing on the joint parts; the other exposed parts of the limbs output by the human body model are fused and adjusted according to the skin color of the exposed skin of the similar parts in the original two-dimensional image; the superimposed image contains missing pixels caused by the cutout, and the missing pixels are filled according to the background pixel color in a certain area around the missing parts; 10) Restore the skin color of exposed parts of the human body; 11) Obtain a 2D image after changing clothes, which consists of the target human head and limbs, the target image background, and the clothes worn by the target human model.
2. The method according to claim 1, characterized in that The part where the human body model outputs a two-dimensional image and the head of the original two-dimensional image is connected is fused and adjusted according to the detected facial skin color.
3. The method according to claim 1, characterized in that The driving process further includes: driving the skeleton of the human body model to move from the initial posture to the target posture by using an interpolation method; and driving the follow-up movement of the human body model mesh by the movement of the skeleton.
4. The method according to claim 3, characterized in that After obtaining the initial state of the skeleton information and the target posture state parameters, the skeleton is driven to move from the initial posture to the target posture, and a time series of skeleton information from the initial posture to the target posture is formed through linear interpolation or nearest neighbor interpolation.
5. The method according to claim 1, wherein In the process of generating animation sequences, the movement of the human body model mesh is also performed in an interpolation manner. After each frame drives the skeleton movement, the current state of the human body model vertex or face information is calculated using the weight parameters of the standard human body model, and the current human body model mesh state is updated and recorded.
6. The method according to claim 1, characterized in that The steps of obtaining parameters of the target human body model include: 1) obtaining a two-dimensional image of the target human body; 2) processing to obtain a two-dimensional human body contour image of the target human body; 3) substituting the two-dimensional human body contour image into a first neural network that has undergone deep learning to perform joint point regression; 4) obtaining a joint point map of the target human body; obtaining a semantic segmentation map of each part of the human body; and body key points. body skeleton points; 5) the generated target human body joint point map, semantic segmentation map, body skeleton points and key point information are substituted into the second neural network after deep learning to regress the human body posture and body shape parameters; 6) the output 3D human body parameters are obtained, including 3D human body motion posture parameters and 3D human body shape parameters.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1 to 6 are implemented.
8. An electronic device, characterized in that: The invention comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement the method steps described in any one of claims 1 to 6 when executing the program stored in the memory.
Citation Information
Patent Citations
Skeleton-based rapid garment fitting method
CN108537888A
A virtual wear method and system with image deformation
CN109035413A
Video-based attitude data capture method and system
CN109145788A