A virtual fitting method, device and storage medium for illumination matching

By obtaining semantic segmentation and lighting matching of the target human body's two-dimensional image, combined with deep learning and mathematical models, a highly realistic three-dimensional clothing model is generated, which solves the problems of slow modeling speed and insufficient accuracy in virtual fitting, and achieves efficient and realistic clothing matching effects.

CN114202630BActive Publication Date: 2025-09-12BEIJING MOMO INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010876706.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-27
Publication Date
2025-09-12
Estimated Expiration
2040-08-27

AI Technical Summary

Technical Problem

Existing virtual fitting technology has problems such as slow modeling speed, insufficient accuracy, large computational complexity, and significant influence of lighting conditions when generating human body models and clothing models. This results in unrealistic virtual dressing effects and makes it difficult to meet the convenience and high-quality needs of Internet users.

Method used

By obtaining the semantic segmentation map of the target human body's two-dimensional image, matching the lighting intensity and lighting angle, using deep learning neural networks to generate three-dimensional clothing models, and combining mathematical models to construct standard human body models, efficient matching of clothing models and human body models is achieved. A variety of prefabricated model libraries considering lighting conditions are used, and cloth simulation and skinning methods are used to improve the realism of clothing.

Benefits of technology

It achieves the rapid generation of highly realistic virtual fitting effects based on a photo, with clothing naturally matching the human body. The operation is simple and adaptable to Internet scenarios, which improves the user experience and stickiness of virtual fitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202630B_ABST
    Figure CN114202630B_ABST
Patent Text Reader

Abstract

The present invention discloses a virtual fitting method for illumination matching, comprising: obtaining a two-dimensional image of a target human body; obtaining a semantic segmentation map of the target human body image; obtaining a two-dimensional image of clothing; matching the illumination intensity and illumination angle of prefabricated clothing with the target human body image; and selecting an appropriate clothing texture map to produce a three-dimensional model of clothing. The present invention maintains the realism and restoration of the three-dimensional clothing model through a series of methods. The three-dimensional clothing model is obtained by processing the two-dimensional clothing image in advance. The user does not need to participate in these behind-the-scenes tasks. The system will automatically match the clothing model with the corresponding illumination intensity and illumination angle, maintaining the realism and restoration. It adapts very well to the simple and fast characteristics and trends of the Internet era. Users only need to upload a photo, which is all the virtual clothing changing work they need to complete.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of user virtual dressing and fitting, and specifically relates to human body modeling, clothing modeling, and the fitting of clothing models with human body models used in virtual dressing, especially a virtual fitting method, device, and storage medium for calculating and extracting relevant lighting information from a target human body photo and matching the lighting information with the target human body according to the lighting information. Background Art

[0002] With the development of internet technology, online shopping has become increasingly popular. Compared to in-store shopping, online shopping offers advantages such as a wider variety of products and greater convenience. However, online shopping also presents some difficult-to-solve challenges, the most prominent of which is the inability to physically view the items. This issue is particularly prominent in clothing. Unlike in-store shopping, where customers can change outfits and see how they look in real time, online clothing shopping offers no personalized images. Instead, they only provide pictures of models trying on clothes, or sometimes even no pictures at all. This prevents consumers from visually assessing the degree to which the clothing matches their body image in real time, leading to a high number of returns and exchanges.

[0003] To address this issue, businesses are attempting to utilize virtual fitting technology to provide consumers with simulated fitting experiences. Of course, there are other practical applications for virtual fitting technology, such as in online games. Consequently, this technology has seen rapid development.

[0004] Virtual fitting refers to a technology application that allows users to view the desired outfit changes in real time on a terminal screen, without having to physically put on the desired clothing. Existing fitting technologies primarily include two-dimensional fitting and three-dimensional virtual fitting. The former essentially captures a user's image and clothing images, then stretches or compresses the clothing to the same size as the person, then crops and splices the pieces to create a "dressed" image. However, these images lack realism due to crude image processing, completely ignoring the user's actual body shape and simply forcing the clothing onto the user's photo, failing to meet user needs. The latter typically uses 3D acquisition equipment to capture 3D information about the person and synthesize it with clothing features. Alternatively, users manually input body data, generate a virtual 3D human body mesh according to specific rules, and then combine it with clothing textures. Overall, this type of 3D virtual fitting requires extensive data acquisition and 3D data calculations, resulting in high hardware costs and limited adoption for the general public.

[0005] With the development of cloud computing technology, artificial intelligence technology, and intelligent terminal processing capabilities, two-dimensional virtual fitting technology has emerged. This technology mainly includes three steps: (1) processing the personal body information provided by the user to obtain a target human body model; (2) processing the clothing information to obtain a clothing model; (3) fusing the human body model and clothing model to generate a simulated image of the person wearing the clothing.

[0006] Regarding point (1), due to the accumulation of many uncertain factors such as process design, model parameter selection, and neural network training methods, the quality of the final generated clothing-changing pictures is not as good as that of traditional three-dimensional virtual fitting technology. Among them, the establishment of the human body model is its basic step, and the subsequent dressing process must also be based on the human body model generated previously. Therefore, once the human body model is generated inaccurately, it is easy to cause problems such as a large difference in body shape between the human body model and the person being fitted, loss of skin texture, loss of body parts, etc., which will affect the final generated clothing-changing picture effect.

[0007] In the general field of computer vision, human body modeling can be achieved through a variety of approaches. These typically include 3D scanning of the real body using omnidirectional scanning equipment, 3D reconstruction methods based on multi-view depth-of-field photography, and methods that combine a given image with a human model to achieve 3D reconstruction. Using 3D scanning equipment to scan the real body provides the most information and is the most accurate. However, this equipment is typically expensive and requires the close cooperation of the human model. The entire processing process places high demands on the processing equipment, so it is generally used in specialized fields. Secondly, multi-view 3D reconstruction methods require providing overlapping images of the reconstructed body from multiple perspectives and establishing spatial transformations between them. Using multiple cameras to capture multiple images and stitch them together to create a 3D model is relatively simple, but computational complexity remains high. In most cases, multi-angle images can only be obtained by on-site personnel. The model created by stitching together the textures obtained from multi-angle depth camera photography lacks body scale data and cannot provide a foundation for 3D perception. Secondly, the single-image method combined with a human body model only requires a single image. A neural network-based intelligent method for generating 3D human feature curves uses neural network training to obtain weights and thresholds that can be used to describe curves of parts such as the neck, chest, waist, and hips. Then, based on dimensional parameters such as the girth, width, and thickness of the human cross-section, a 3D curve that matches the real human body shape can be directly generated, resulting in a predicted human body model. However, this method requires relatively little input information, and the solution process still requires a high level of computational effort, resulting in unsatisfactory final model results.

[0008] Regarding point (2), there are several different methods in the existing technology for generating three-dimensional clothing models. At present, the more traditional way to build a three-dimensional clothing model is based on the design and stitching method of two-dimensional clothing pieces. This method requires a certain amount of clothing expertise to design the sample, which is not a quality that all virtual fitting users have. At the same time, this method also requires manual specification of the stitching relationship between the samples, which will consume a lot of time to set up. In addition, another relatively new three-dimensional modeling method is based on hand-drawing, which can generate a simple clothing model through the line information hand-drawn by the user. However, this method requires professional personnel to hand-draw, and its reproducibility and repeatability are poor. It takes a lot of time for users to draw the details of the clothing, and it is difficult to promote it on a large scale in e-commerce. Both methods tend to innovate and design new clothing rather than perform three-dimensional modeling on existing clothing for sale. Another method is to obtain clothing picture information and use image processing technology and graphic simulation technology in combination to finally generate a virtual three-dimensional clothing model. The contour and size of the clothing are obtained through contour detection and classification in the image. The key points of the edges are found from the contour through machine learning methods. The stitching information is generated based on the correspondence between the key points. Finally, the physical stitching of the clothing is simulated in three-dimensional space to obtain the real effect of the clothing being worn on the human body.

[0009] In summary, based on Internet technology and the characteristics of the network environment in which it is located, the method of directly outputting the final image or photo after changing clothes from a single human body image is undoubtedly the most preferred. It is the most convenient and the user does not need to be present at the scene. With just one photo, the entire virtual dressing process can be completed. Then the problem that follows is that as long as the resulting photo effect can be guaranteed to be basically equivalent to the real 3D simulated dressing process, it will become the mainstream. Among them, (1) how to obtain a human body model that is closest to the real state of the human body through a photo, and (2) how to put the three-dimensional clothing model on the human body model in the state closest to the real state, have become the two most important and unavoidable problems in the virtual dressing method.

[0010] Regarding the first point. In the prior art, there are usually several methods for constructing human body models: (1) Regression-based methods, which use convolutional neural networks to reconstruct a voxel-represented human body model. The algorithm first estimates the positions of the main joints of the human body based on the input image, and then estimates whether each unit voxel in the voxel grid of a given size is occupied based on the key point positions, thereby using the entire shape of the occupied voxels to describe the reconstructed human body shape; (2) Human body reconstruction based on a single image, which simultaneously estimates the three-dimensional shape and posture of the human body. This method first roughly marks the simple key points of the human skeleton on the image, and then performs initial matching and fitting of the human body model based on these rough key points to obtain the approximate shape of the human body. (3) Use 23 bone nodes to represent the human skeleton, and then use the rotation of each bone node to represent the posture of the entire human body. At the same time, use 6890 vertex positions to express the human body shape. In the fitting process, the shape and posture parameters are fitted at the same time given the bone node positions, thereby performing three-dimensional human body reconstruction; or first use a CNN model to predict the key points on the image, and then use the SMPL model for fitting to obtain the initial human body model. Next, the fitted shape parameters are used to regress a bounding box for each joint. Each joint corresponds to a bounding box, represented by its axis length and radius. Finally, the initial model and the regressed bounding boxes are combined to create a 3D human reconstruction. However, these methods suffer from slow modeling speed, insufficient modeling accuracy, and a strong dependence on the established body and posture database.

[0011] Regarding the second point. There is such a virtual fitting solution in the prior art, including: obtaining a dressed reference human body model and an undressed target human body model; embedding skeletons of the same hierarchical structure for the reference human body model and the target human body model respectively; performing skin binding on the skeletons of the reference human body model and the target human body model; calculating the rotation amount of the bones in the skeleton of the target human body model, and recursively adjusting all the bones in the skeleton of the target human body model so that the posture of the target human body model skeleton is consistent with that of the reference human body model skeleton; using the LBS skinning algorithm to deform the skin of the target human body model according to the rotation amount of the bones in the skeleton of the target human body model; based on the skin deformation of the target human body model, migrating the clothing model from the reference human body model to the target human body model. This invention can reduce the difficulty of migrating the clothing model from the reference human body to the target human body after the postures of the target human body model and the reference human body model are adjusted to be consistent, and convert the inefficient non-rigid registration problem into an efficient rigid registration problem, thereby realizing the migration of the clothing model from the reference human body model to the target human body model. This approach solves the technical problem of automatically fitting clothing on different body types and in different postures while maintaining the same size before and after fitting. However, this skinning method places too much emphasis on a fixed distance between clothing and skin. While this method offers advantages in fitting speed, it suffers from significant disadvantages in terms of clothing matching and realism, making it suitable only for situations where clothing needs to be quickly and easily adapted to follow the movement of the skin mesh. Furthermore, the aforementioned method focuses solely on maximally restoring and reproducing the clothing model, without addressing the original appearance of the clothing model.

[0012] The first prior art discloses a method for reconstructing geometric details of human clothing from a single perspective with illumination separation, comprising: collecting human motion data through single RGB to obtain a single RGB image, extracting the human posture from the single RGB image, and solving the shape, posture and relative spatial position of the human in each frame; generating a two-dimensional clothing mesh model through a preset clothing template, and sewing different parts of the clothing and putting them on the human in the initial posture through a particle simulation method; transitioning the human posture to the posture of the first frame in the video, and performing a joint physical simulation on the three-dimensional clothing, and performing frame-by-frame clothing physical simulation based on the human posture for all subsequent frames; solving the problem of human body segmentation through a human body segmentation method. Solve the clothing parameters so that the simulated clothing shape and the segmentation map in the image meet the matching conditions; for each frame in the video, use the image illumination separation method to extract the image's intrinsic illumination image and intrinsic albedo image; use the physical simulation to obtain the initial clothing shape, and solve the vertex-by-vertex normal of the clothing mesh model, and obtain the illumination information through the assumption of spherical harmonic illumination; under the assumption of spherical harmonic illumination and the premise of preset spherical harmonic illumination coefficients, solve the vertex-by-vertex deformation of the clothing to obtain the geometric details of the clothing; project the solved vertex-by-vertex deformation of the clothing for each frame into the local coordinate system on each vertex, and perform time-domain smoothing of the projection coefficients of each frame to obtain the final dynamic clothing detail reconstruction result.

[0013] While this existing technology takes into account the impact of lighting conditions on clothing models, it only requires a single RGB camera to capture the human body and utilizes the intrinsic decomposition information of the image to obtain scene illumination. This allows for the combined solution of illumination and clothing surface detail information, enabling simultaneous modeling and simulation of both the person and clothing in a single RGB video. This framework of clothing modeling and surface detail solution enables relatively good reconstruction of clothing details in the input image. However, due to its excessive focus on achieving a high degree of restoration for each frame, the computational methods and principles employed are relatively complex, requiring continuous computation throughout the entire process. Consequently, the time control for clothing model reconstruction is poor, resulting in prolonged simulation and solution times, making it unsuitable for virtual dressing applications in internet scenarios.

[0014] Therefore, to keep pace with the development trends of the internet industry, in the virtual fitting market, the three fundamental goals are to minimize input information, minimize computational effort, and achieve optimal results. We need to find an optimal balance between these three factors, providing a clothing matching method that allows for simple input, minimizes computational effort within the device's capabilities, and achieves results similar to those achieved with professional equipment. Summary of the Invention

[0015] Based on the above problems, the present invention provides a three-dimensional clothing model matching method, device and storage medium that overcome the above problems.

[0016] The present invention provides a virtual fitting method for illumination matching, the method comprising:

[0017] 1) Obtain a two-dimensional image of the target human body;

[0018] 2) Obtaining a semantic segmentation map of the target human image;

[0019] 3) obtaining a two-dimensional image of the garment;

[0020] 4) Matching the illumination intensity and illumination angle of the prefabricated garment with the target human body image;

[0021] 5) Select appropriate clothing texture maps to create a three-dimensional model of the clothing.

[0022] Furthermore, it also includes a matching process between the clothing model and the human body model, combining the mathematical model to construct a three-dimensional standard human body model, the three-dimensional standard model is in an initial pose; the three-dimensional clothing model is fitted to the three-dimensional standard human body model of the initial pose, that is, the standard basic human body; according to the two-dimensional image of the target human body, the three-dimensional target human body model parameters are obtained through the neural network model calculation; the obtained several sets of POSE bases and body shape three-dimensional human body parameters are input into the three-dimensional standard human body model for fitting; and the target human body model with the same posture and body shape as the target human body and wearing the changed clothes is obtained.

[0023] Furthermore, a contour map of the target human body two-dimensional image is obtained, and the two-dimensional human body contour image is substituted into a deep learning neural network for regression to obtain a semantic segmentation map of the target human body image.

[0024] Furthermore, the matching process also includes obtaining the RGB color space values ​​of the target human body's two-dimensional image, converting the RGB values ​​of the target image into XYZ values ​​in the XYZ color space, converting the XYZ values ​​of the target image's XYZ color space into Lab color space values, combining the Lab image with the semantic segmentation map to obtain the L channel values ​​corresponding to the target human body part, and obtaining the illumination intensity level corresponding to the target human body part from the L channel values. According to different illumination intensity levels, texture maps of the clothing model under different illumination intensities are collected, and the texture map corresponding to the illumination intensity of the target human body is selected for rendering the clothing model.

[0025] The process also includes obtaining the RGB values ​​of the target human image, converting the RGB values ​​into grayscale values, combining the grayscale image with the semantic segmentation map to obtain a grayscale image corresponding to the target human part, calculating the horizontal and vertical brightness difference values ​​δ for each pixel in the image region of interest, and then obtaining the overall image brightness difference statistical feature γ in the grayscale image. By comparing a pre-established image difference statistical feature image library, the cosine similarity value between the current image and the feature library image is calculated, and the image with the highest feature similarity is selected to estimate the illumination angle of the current image. According to different illumination angles, texture maps of the clothing model at different illumination angles are collected, and the texture maps corresponding to the illumination angle of the target human body are selected for clothing model rendering.

[0026] In addition, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the aforementioned methods and steps are implemented.

[0027] An electronic device, characterized in that it includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement the methods and steps described above when executing the programs stored in the memory.

[0028] The beneficial effects of the present invention are:

[0029] 1. The lighting conditions of the virtual clothing are highly consistent with the target body image. When we initially created the 3D clothing models, we took into account the various lighting conditions of the target body photos, including varying intensities and angles. This can result in noticeable differences in the photos. For example, if the lighting intensity in the photo is poor, the face and skin color of the person appear darker. Since the clothing models are generated using good lighting conditions, the face and exposed limbs in the photos after the outfit change are likely to appear darker, while the clothing is very bright, resulting in a very unrealistic effect. To this end, we adopted a method of pre-creating models for multiple lighting conditions for a single set of clothing. Firstly, the time-consuming process of creating different clothing models is completed in advance, so the user does not notice any processing delay after clicking to start the outfit change. Secondly, the clothing model and the body photo match very naturally, with dark areas appearing darker and bright areas appearing brighter, and the transition between clothing and skin is very natural. This is particularly advantageous when the virtual outfit change involves only half of the body or a single item, and the photos after the outfit change look seamless. For example, a cheongsam will have more than ten types of lighting, combining different light intensity levels and lighting angle combinations. In other words, for a cheongsam, there will be more than ten such cheongsams in our clothing model library, which can be perfectly matched with the input photos under different lighting conditions.

[0030] 2. High realism of virtual clothing. Realism encompasses two aspects: first, the natural adaptation of the clothing to the human body; and second, the fidelity of the fabric texture. We first place the 3D clothing model on a standard 3D mannequin. We input the body parameters of the target mannequin to create the target mannequin. Through garment adaptation, the 3D clothing model adapts to the body shape changes from the standard mannequin to the target mannequin, and the clothing changes with remarkable fidelity. Furthermore, we employ cloth simulation, which simulates fabric effects close to reality (though this can also differ from real-world physical effects). We highlight the high fidelity of the fabric texture simulation, including the accuracy of the fabric print simulation. Cloth simulation combines cloth simulation with skinning. Skinning is used for parts of the garment that remain largely unchanged after movement, while cloth simulation is used for parts that deform during movement to ensure fidelity and computational speed. During cloth simulation, when the mannequin reaches the target pose, gravity calculations are performed on the garment fabric for several frames to ensure fidelity in the target pose.

[0031] 3. Simple user operation. The present invention provides a method for obtaining accurate three-dimensional human body model parameters by analyzing full-body photos of the human body through a deep neural network. Only an ordinary photo is needed to quickly model the human body. At the same time, the three-dimensional clothing model is obtained by processing the two-dimensional clothing picture in advance. The user does not need to participate in these behind-the-scenes work. He only needs to select the clothing style he wants to try on virtually, and the system will automatically match the corresponding clothing model. It adapts very well to the characteristics and trends of the Internet era. One is simple, and the other is fast. The user does not need to make any preparations. Uploading a photo is all the work the user needs to do. If the present invention is applied to scenarios such as entertainment applets or online shopping, it will greatly enhance the user experience and stickiness. The 3D model that can be obtained without the need for a depth of field camera or multiple sets of cameras corresponds to the real form of the human body, providing a wide range of application scenarios for various industries, such as clothing, health, etc.

[0032] 4. High-Frequency Use of Deep Neural Networks. This invention fully leverages the advantages of deep learning networks, enabling high-precision restoration of human posture and shape in a variety of complex scenarios. By using different neural networks for different purposes and utilizing neural network models with different input conditions and training methods, it achieves precise contour separation of the human body against complex backgrounds, semantic segmentation of the human body, and identification of key points and joints, eliminating the influence of loose clothing and hairstyles, achieving the greatest approximation to the true shape and form of the human body.

[0033] 5. The human body model is precise and controllable. The most commonly used parametric model is the Max Planck Institute's SMPL model, which contains two sets of 72 parameters, one for describing human posture and the other for describing human body shape. However, the SMPL model primarily uses a large number of human body model examples for deep learning and training. The relationship between body shape and shape basis is a holistic association, which is very difficult to decouple, making it impossible to arbitrarily control the desired body part. However, our human body model is not obtained through training, and the parameters have a correspondence based on mathematical principles. In other words, our sets of parameters are independent of each other and have no interdependence. Therefore, our model is more interpretable during transformation and can better represent the shape changes of a specific part of the body.

[0034] The present invention maintains the realism and restoration of three-dimensional clothing models through a series of methods, and fully considers the impact of light intensity and light angle on the quality of the final film when matching clothing models and human body models. Several groups of clothing models with different lighting conditions are set in advance, and the clothing model that best matches the input photo is used for virtual fitting, which not only maintains the fast processing speed but also maintains the authenticity of clothing simulation. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0036] Figure 1 is a flowchart of the entire process of one embodiment;

[0037] Figure 2 A processing flow chart of a light intensity parameter obtaining module according to one embodiment;

[0038] Figure 3 A processing flow chart of a module for obtaining illumination angle parameters according to one embodiment;

[0039] Figure 4 Schematic diagram of the system of the present invention. DETAILED DESCRIPTION

[0040] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In order to make the objects, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present invention by illustrating examples of the present invention.

[0041] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0042] The method for processing a human body image provided by an embodiment of the present invention is described in detail below with reference to the accompanying drawings.

[0043] like Figure 1As shown, an embodiment of the present invention discloses a virtual fitting method for illumination matching, the method comprising:

[0044] 1) Obtain a two-dimensional image of the target human body;

[0045] 2) Obtaining a semantic segmentation map of the target human image;

[0046] 3) obtaining a two-dimensional image of the garment;

[0047] 4) Matching the illumination intensity and illumination angle of the prefabricated garment with the target human body image;

[0048] 5) Select appropriate clothing texture maps to create a three-dimensional model of the clothing.

[0049] The above-mentioned method for matching clothing models with human models generally includes several steps. First, a 3D clothing model is generated; second, a standard human model is generated and fitted onto the 3D clothing model; third, the parameters of the target pose human model are obtained; and fourth, the standard human model's body shape and pose are adjusted to match those of the target human model, while simulating the realistic changes that occur in the 3D clothing model as the human model changes.

[0050] The first part primarily involves generating a 3D clothing model. Several different methods exist in the existing art for generating 3D clothing models. Currently, the more traditional method for creating 3D clothing models relies on the design and stitching of 2D clothing pieces. This method requires specialized clothing knowledge to design the pattern. Another relatively novel 3D modeling method is based on hand-drawing, which can generate a simple clothing model using the user's hand-drawn line information. Another method uses a combination of image processing and graphics simulation techniques based on garment image information to ultimately generate a virtual 3D clothing model. Contour detection and classification are used to obtain the garment's outline and dimensions from the image. Machine learning methods are then used to identify key points between edges within the outline. Stitching information is generated based on the key point correspondences. Finally, the garment is physically stitched together in 3D space to simulate the garment's appearance on the human body. Other methods include mapping and mathematical model simulation, and this section is not specifically limited in the present invention.

[0051] However, the 3D clothing model needs to be matched to a standard human body model. The general requirement is to adapt the clothing model, which has been adapted to the standard human body model, to the human body model in the target pose through fabric physics simulation, and to ensure the naturalness and rationality of the clothing. Therefore, some basic requirements must usually be met, including but not limited to the following: a. It must completely fit the initial pose of the standard mannequin without any penetration; b. The output is a uniform quadrilateral; c. The UVs of the model need to be unfolded, flattened and compactly aligned, and the textures must be manually aligned using Photoshop tools; d. Vertex merging must be performed; e. The output model should be uniformly reduced in facets, with the total facet count not exceeding 150,000 faces / set according to the reference standard; f. The material needs to be adjusted in mainstream clothing design software, and a 10-frame animation needs to be solved to observe the fabric effect. The expected material parameters must be saved; g. The rendering material needs to be adjusted in mainstream design software, and a preview of the rendering must be taken to ensure that the material Lambert properties are reasonable.

[0052] The virtual fitting method of the present invention with illumination matching comprises the following steps: 1) obtaining a two-dimensional image of a target human body; 2) obtaining a semantic segmentation map of the target human body image; 3) obtaining a two-dimensional image of clothing; 4) matching the illumination intensity and illumination angle of the prefabricated clothing with the target human body image; and 5) selecting an appropriate clothing texture map to produce a three-dimensional model of the clothing.

[0053] Among them, after obtaining the two-dimensional image of the target human body, the human body contour map can be output according to the algorithm or substituted into the neural network model, and the two-dimensional human body contour image is substituted into the deep learning neural network for regression to obtain the semantic segmentation map of the target human body image.

[0054] Step 3) is particularly critical and also includes the following steps: obtaining the RGB color space values ​​of the target human body's two-dimensional image, converting the RGB values ​​of the target image into XYZ values ​​in the XYZ color space, and converting the XYZ values ​​of the target image's XYZ color space into Lab color space values. The Lab image is combined with the semantic segmentation map to obtain the L channel values ​​corresponding to the target human body part, and the L channel values ​​are used to obtain the corresponding light intensity level of the target human body part.

[0055] Brightness refers to the degree of lightness or darkness in an image, contrast refers to the difference in lightness or darkness, and saturation refers to the richness of the image's colors. Image files are typically in RGB format, primarily for display purposes. RGB is an abbreviation for three colors: R stands for red, G stands for green, and B stands for blue. Modern color theory considers all colors to be combinations of red, green, and blue. In computers, each color is recorded using a byte. An RGB image file uses three bytes to record each of the three colors, red, green, and blue, respectively. Therefore, good image files are 24-bit. Some image files also support transparency, which can also be recorded using a single byte. Therefore, image files that support transparency are 32-bit. When recording color using a byte, the byte can be thought of as a number. A byte has 8 bits, each representing a binary value of 0 or 1. Converting an 8-bit binary number to decimal can represent a range from 0 to 255. Color values ​​can then be expressed as values ​​from 0 to 255, representing the brightness or darkness of the color. When the value is 0, the color is the darkest, and when the value is 255, the color is the brightest. When the red, green, and blue color values ​​are all 0, the image is black, and when the red, green, and blue color values ​​are all white, the image is white. Therefore, the changes in the red, green, and blue color values ​​can combine to create 16,777,216 colors, including black, white, and gray.

[0056] Compared to the RGB color space, Lab is a less commonly used color space. It was established based on the international color measurement standards established by the International Commission on Illumination (CIE) in 1931. In 1976, it was revised and officially renamed CIELab. It is a device-independent color system and a color system based on physiological characteristics. This means it uses a digital method to describe human visual perception. The L component in the Lab color space represents pixel brightness and has a value range of [0, 100], representing pure black to pure white; a represents the range from red to green and has a value range of [127, -128]; and b represents the range from yellow to blue and has a value range of [127, -128].

[0057] The main difference between the two is that RGB is composed of a red channel (R), a green channel (G), and a blue channel (B). The brightest red + the brightest green + the brightest blue = white; the darkest red + the darkest green + the darkest blue = black; and between the brightest and darkest, red of the same brightness + green of the same brightness + blue of the same brightness = gray. Within any RGB channel, white and black represent the brightness of that color. Therefore, where there is white or grayish white, the R, G, and B channels cannot be black, because these three channels are required to form these colors.

[0058] LAB is different. The lightness channel (L) in LAB is specifically responsible for the brightness and darkness of the entire image—simply put, it's the black and white version of the entire image. The a and b channels are solely responsible for the amount of color. The a channel represents the range from magenta (white in the channel) to dark green (black in the channel); the b channel represents the range from burnt sienna (white in the channel) to shimmering bluish blue (black in the channel). A 50% neutral gray in the a and b channels represents the absence of color, so the closer to gray, the less color there is. Furthermore, the colors in the a and b channels have no brightness. This explains why the outlines of red clothing are often very clear in the a and b channels, as red is composed of magenta and burnt sienna.

[0059] An image is made up of dots interwoven horizontally and vertically, each dot being called a pixel. The number of dots in the image, both horizontally and vertically, constitutes its resolution. The product of these two numbers is the number of pixels, which can be used to measure the resolution of an image. Cropping an image reduces the number of pixels, lowering its resolution. Each color value in each pixel can range from 0 to 255, with higher values ​​indicating a brighter color. Therefore, changing the brightness of an image effectively changes the value of each color in every pixel. Increasing the brightness of an image is like adjusting the values ​​of each color in every pixel in the image inversely, which decreases the brightness. For each color value in each pixel, a value less than 127 is considered dark, while a value greater than 127 is considered bright. If we reduce the color values ​​of all pixels on the image that are less than 127, and increase the color values ​​of all pixels on the image that are greater than 127, what we see is the adjustment of the contrast of the image, that is, making the dark parts of the image darker and the bright parts of the image brighter.

[0060] Light intensity estimation primarily relies on the values ​​of the Lab channel; this cannot be accomplished directly using RGB images. The L component in the Lab color space represents pixel brightness, with a value range of [0, 100], representing pure black to pure white. RGB images cannot be directly converted to Lab channels. Instead, the XYZ color space is used to convert the RGB color space to the XYZ color space, and then back to the Lab color space. The specific process is as follows:

[0061] (1) RGB to XYZ

[0062] Assuming that r, g, and b are the three channels of a pixel and their value range is [0, 255], the conversion formula is as follows:

[0063]

[0064]

[0065]

[0066] M=0.4124,0.3576,0.1805

[0067] 0.2126, 0.7152, 0.0722

[0068] 0.0193, 0.1192, 0.9505

[0069] Equivalent to the following formula:

[0070] X=var_R*0.4124+var_G*0.3576+var_B*0.1805

[0071] Y=var_R*0.2126+var_G*0.7152+var_B*0.0722

[0072] Z=var_R*0.0193+var_G*0.1192+var_B*0.9505

[0073] The gamma function above is used to perform nonlinear tone editing on an image in order to increase image contrast. This function is not the only one.

[0074] RGB is the color component after gamma correction: R = g(r), G = g(g), B = g(b), where rgb is the original color component.

[0075] g is the Gamma correction function: when x < 0.018, g(x) = 4.5318*x; when x > = 0.018, g(x) = 1.099* d^0.45-0.099

[0076] The value ranges of rgb and RGB are both [0, 1). After the calculation is completed, the value ranges of XYZ change slightly to [0, 0.9506), [0, 1), and [0, 1.0890).

[0077] (2) XYZ to LAB

[0078] L ★ =116f(Y / Y n )-16

[0079] a ★ =500[f(X / X n )-f(Y / Y n )]

[0080] b ★ =200[f(Y / Y n )-f(Z / Z n )] (5)

[0081]

[0082] In the above two formulas, L*, a*, and b* are the values ​​of the three channels of the final LAB color space.

[0083] Wherein f is a correction function similar to the Gamma function: when x>0.008856, f(x)=x^(1 / 3); when x<=0.008856, f(x)=(7.787*x)+(16 / 116).

[0084] x, Y, and Z are the values ​​calculated after linear normalization after RGB conversion to XYZ. The default values ​​for xn, Yn, and Zn are generally 95.047, 100.0, and 108.883, respectively. After calculation, L is in the range [0, 100), while a and b are approximately [-169, +169] and [-160, +160], respectively.

[0085] Through the above conversion calculations, we can obtain the L channel value corresponding to the 2D image of the human body. By comparing the L channel values ​​of the region of interest, we can quantify the light intensity into several different levels. For example, we can average the L values ​​of the region of interest and use this average L' value to quantify the light level. Because human visual aesthetics generally prefer photos with moderately bright lighting, when selecting the L' value, we generally choose a value that is brighter than the 80th percentile.

[0086] Since several sets of clothing texture maps have been collected under different lighting intensity conditions when the clothing model was made, after obtaining the lighting intensity level in the previous step, the texture maps of the clothing model under different lighting intensities can be collected according to different lighting intensity levels, and the texture maps with corresponding lighting intensity can be selected according to the target human body lighting intensity to render the clothing model.

[0087] Step 3) also includes the following parallel steps: converting the grayscale image through the RGB image and estimating the illumination angle by calculating the gradient.

[0088] Theoretically, there are many algorithms for calculating lighting direction, but basically no solution can extremely accurately solve the lighting problems in various scenarios in the real world. Therefore, this solution appropriately simplifies the lighting sources in the real environment. In many scenarios, the lighting source is from the left or right, and the influence of highlight, horizontal light or bottom light is relatively large. The specific difference in the degree of highlight is not that obvious. Experiments have shown that the lighting angle calculated by this solution can basically distinguish the approximate lighting direction.

[0089] Obtain the RGB values ​​of the target human image and convert them to grayscale. This first step eliminates any artifacts introduced by color differences in the original image. Since the focus is on the illumination direction of the human body, the semantic segmentation information from the 2D image is combined to extract only the human body region of interest (extending appropriately) for illumination direction calculation.

[0090] The grayscale image is combined with the semantic segmentation map to generate a grayscale image corresponding to the target human body. Appropriate filter operators are then used to calculate the horizontal and vertical brightness difference δ for each pixel in the image's region of interest, thereby determining the overall image brightness difference statistical feature γ within the grayscale image. By comparing the current image with a pre-established image library of image difference statistical features, the cosine similarity between the current image and the library images is calculated. The image with the highest feature similarity is selected as the current image's illumination angle. Texture maps of the clothing model are collected at different illumination angles, and the texture map corresponding to the target human body's illumination angle is selected for rendering the clothing model.

[0091] By integrating the two parameters of light intensity and light angle, we can obtain a three-dimensional clothing model that is very close to the lighting conditions of the target human body photo, laying a very good foundation for the overall realism of clothing simulation.

[0092] The second part involves pre-designing and modeling some basic mannequins based on our human modeling method, and then fitting the 3D clothing model onto the standard mannequin to achieve compatibility with our subsequent workflow. The main work involves combining mathematical models to construct a 3D standard mannequin, the basic mannequin. The Max Planck Institute's SMPL mannequin avoids surface distortion during human motion and accurately depicts the morphology of muscle extension and contraction. In this method, β and θ are the input parameters. β represents 10 parameters related to a person's height, weight, head-to-body ratio, and other proportions, while θ represents 75 parameters representing the overall motion pose and the relative angles of 24 joints. The β parameter is a ShapeBlendPose parameter that can be used to control the shape of the human body using 10 incremental templates. Specifically, the changes in the shape of each parameter can be captured through animated graphics. By studying the continuous animation of the parameter changes, we can clearly see that each continuous change in the parameter controlling the human shape leads to chain changes in the model, both locally and globally. To reflect the movement of the human musculature, linear changes in each parameter of the SMPL mannequin cause large-scale mesh changes. To illustrate this, for example, when adjusting the parameters of β1, the model will directly interpret these changes as changes to the entire body. You might only want to adjust the waist proportions, but the model will force adjustments to the legs, chest, and even the weight of the hands. While this working mode greatly simplifies the workflow and improves efficiency, it is indeed very inconvenient for projects that focus on modeling quality. This is because the SMPL human body model is ultimately trained using Western body photos and measurements to conform to Western body shapes. Its body shape changes generally conform to the typical curve of Western people. Applying this model to Asian human bodies can lead to many problems, such as arm-leg proportions, waist-to-body ratios, neck proportions, and leg and arm lengths. Our research has shown that there are significant discrepancies in these aspects. If the SMPL human body model is rigidly applied, the final generated results will not meet our requirements.

[0093] To achieve this, we employed a custom-built human model approach to achieve enhanced performance. The core of this approach is the creation of a custom blendshape base to enable precise, independent manipulation of the human body. Preferably, the three-dimensional standard human model (basic mannequin) consists of 20 body base parameters and 170 skeletal parameters. These bases comprise the entire human model, with each base independently controlled by its own parameters, without interfering with each other. Precise manipulation, on the one hand, involves increasing the number of control parameters, rather than relying on the Max Planck Institute's ten β control parameters. This allows for adjustable parameters beyond the standard body shape, including arm length, leg length, waist, hip, and chest shape. This more than doubles the number of skeletal parameters, significantly expanding the range of adjustable parameters and providing a solid foundation for the refined design of standard human models. Independent manipulation means that each base, such as the waist, legs, hands, and head, can be manipulated independently, and each bone can be independently adjusted in length, independently of each other and without any physical interaction. This allows for precise adjustments to the human model. The model no longer appears "cumbersome and clumsy," unable to be adjusted to the designer's satisfaction. Our current model embodies a mathematically based correspondence. This is essentially a combination of human aesthetics and statistical data analysis, designed to produce a model that we believe is accurate for the Asian body shape according to our design rules. This significantly differs from the SMPL body model trained on big data. Therefore, our parameter transformations are more interpretable and can better represent local shape changes in the body model. Furthermore, these changes are based on mathematical principles, ensuring that no parameters influence each other, and the arms and legs remain completely independent. The reason for designing so many different parameters is to avoid the drawbacks of body models trained on big data. This allows for precise control of the body model in multiple dimensions, beyond just a few parameters like height, significantly improving modeling performance. Setting so many independent control parameters is only meaningful when building a custom body base; both are essential to achieving designer-level performance.

[0094] As for putting the three-dimensional clothing model on the standard human body model, it is a conventional technology in this field. The present invention does not make too many restrictions on this, as long as the required effect can be achieved.

[0095] The third part is to process the acquired human body images to obtain the parameter information required to generate the human body model. In the past, the selection of these skeletal key points was usually done manually, but this method is very inefficient and does not adapt to the fast-paced requirements of the Internet era. Therefore, in today's era when neural networks are popular, it has become a trend to use deep learning neural networks instead of manual key point selection. However, how to use neural networks efficiently is a problem that requires further research. Generally speaking, we adopted the idea of ​​​​a secondary neural network plus data "fine-tuning" to construct our parameter acquisition system. Figure 2 As shown in the figure, we use a deep learning neural network to generate these parameters, which mainly includes the following sub-steps: 1) obtain a two-dimensional image of the target human body; 2) process and obtain a two-dimensional human body contour image of the target human body; 3) substitute the two-dimensional human body contour image into the first deep learning neural network for joint point regression; 4) obtain a joint point map of the target human body; obtain semantic segmentation maps of various parts of the human body; body key points; body bone points; 5) substitute the generated joint point map, semantic segmentation map, body bone points and key point information of the target human body into the second deep learning neural network for human posture and body shape parameter regression; 6) obtain the output three-dimensional human body parameters, including three-dimensional human body action POSE parameters and three-dimensional human body shape SHAPE parameters.

[0096] The two-dimensional image of the target human body can be a two-dimensional image including a human body in any posture and any clothing. The acquisition of the two-dimensional human body contour image utilizes a target detection algorithm, which is a target region rapid generation network based on a convolutional neural network.

[0097] Before inputting the two-dimensional human image into the first neural network model, a neural network training process is also included. The training sample includes a standard two-dimensional human image with the original joint point locations manually annotated with high accuracy on the two-dimensional human image. Here, a target image is first acquired and human body detection is performed on the target image using a target detection algorithm. Human body detection does not mean using a measuring instrument to detect a real human body. In this invention, it refers to a given image, typically a two-dimensional photograph, that contains sufficient information, such as a face, limbs, and body. A specific strategy is then used to search the given image to determine whether a human body is present. If a human body is present, parameters such as the human body's location and size are determined. In this embodiment, before obtaining the human body key points in the target image, human body detection is performed on the target image to obtain a human body frame that annotates the human body's location. Because the input image can be any image, some non-human background elements, such as tables, chairs, trees, cars, and buildings, are inevitably present. This unused background is removed using a sophisticated algorithm.

[0098] At the same time, we also need to perform semantic segmentation, joint detection, skeleton detection, and edge detection. By collecting these 1D point information and 2D surface information, we can lay a good foundation for generating a 3D human body model later. Use the first-level neural network to generate a human body joint map. Optionally, the target detection algorithm can quickly generate a network for the target area based on the convolutional neural network. This first neural network requires a large amount of data training. The joints of some photos collected from the Internet are manually annotated and then input into the neural network for training. After deep learning, the neural network can basically obtain a joint map with the same accuracy and effect as manually annotated joints immediately after inputting the photo, and the efficiency is dozens or even hundreds of times that of manual annotation.

[0099] In the present invention, obtaining the joint positions of the human body in the photo is only the first step. After obtaining 1D point information, 2D surface information must be generated based on this 1D point information. All of this can be accomplished using a neural network model and mature algorithms in the prior art. By redesigning the process and timing of the neural network model's involvement and rationally designing various conditions and parameters, the present invention makes parameter generation more efficient and reduces the degree of human involvement. This makes it very suitable for Internet application scenarios. For example, in a virtual dress-up program, users can obtain the dress-up results almost instantly without waiting, which plays a vital role in increasing the program's appeal to users.

[0100] After obtaining the relevant 1D point information and 2D surface information, these parameters or results, the target human body's joint point map, semantic segmentation map, body bone points and key point information can be substituted as input items into the second neural network that has undergone deep learning to regress human body posture and body shape parameters. After the regression calculation of the second neural network, several groups of three-dimensional human body parameters can be immediately output, including three-dimensional human body action POSE parameters and three-dimensional human body shape SHAPE parameters. Preferably, the loss function of the neural network is designed based on a three-dimensional standard human body model (basic mannequin), a predicted three-dimensional human body model, a standard two-dimensional human body image with the original joint point positions marked, and a standard two-dimensional human body image including predicted joint point positions.

[0101] The fourth part, which is also the most critical part, is to fit the parameters of the human body model to the human body model, and at the same time, ensure that the state of the clothes after moving is as realistic as possible. Figure 3As shown, the movement process includes the following substeps: First, the obtained 3D human pose and shape parameters are matched to several bases and skeletal parameters of a 3D standard human model; then, the obtained sets of bases and skeletal parameters are input into a standard 3D human parameter model for fitting; the 3D human model has a mathematical weight relationship between skeletal points and the model mesh; the determination of skeletal points can be correlated to determine the human model of the target human pose. In this step, the two parameters generated in the previous step can be substituted into the pre-designed human model to construct the 3D human model. These two types of parameters are similar in name to the parameters of the Max Planck Institute's SMPL human model, but their actual content differs significantly. This is because the two have different foundations. Specifically, the present invention uses a self-made 3D standard human model (basic mannequin), each base of which is designed based on the body shape and proportions of an Asian, including several areas not covered by the SMPL model. The Max Planck Institute's SMPL model uses a standard human model generated through big data training. The two models are generated using different calculation methods. Although both are ultimately reflected in the generated 3D human model, their connotations differ significantly. After this step, a preliminary 3D human body model will be obtained, including the mesh of the human body model with bone position and length information.

[0102] In this part, after fitting the 3D clothing model onto a standard human body model, we need to align the standard human body model's shape and posture with the target human body, while also simulating the realistic changes that the 3D clothing model undergoes as the human body model changes. We use several methods to ensure this is achieved.

[0103] First of all, equipment adaptation means that based on the clothing model that has been adapted to the standard human body model, a target posture human body model is given, that is, a human body model with the same posture and mesh structure as the standard human body model, the same number of mesh vertices and mesh units, and the same topological connection relationship between vertices. The human body model only has different heights, weights, and other body shapes. Through clothing adaptive matching, the three-dimensional clothing model is adapted to the human body model of the target body shape, and changes with the changes in the weight of the human body model, while ensuring the naturalness and rationality of the clothing.

[0104] During the equipment adaptation process, a field is generated for the mesh patches of the standard human body model, and a fixed field correspondence is established between the various patches of the three-dimensional clothing and the corresponding positions of the standard human body model. When the shape of the standard human body model changes toward the shape of the target human body model, the three-dimensional clothing model can also achieve corresponding uniform changes.

[0105] In this part, the human body model also needs to complete the transformation from the initial pose to the target pose. Because we only input a photo, the target human body posture in the photo is usually different from the basic human body. At this time, in order to fit the target human body posture, it is necessary to complete the transformation from the initial pose to the target pose. In order to simulate more realistically, in the fitting step, when fitting several sets of basis and bone parameters in the standard 3D human body parameter model, the following steps are also included.

[0106] 1) Obtain the position coordinates of the initial and target poses; the initial pose parameters are determined by the initialization parameters of the standard mannequin, and the target pose's skeletal information is predicted by regression using a neural network model. 2) Generate an animation sequence that moves from the initial pose to the target pose. 3) During the generation of the animation sequence, mesh interpolation is used. 4) The interpolation speed is set to be slow for the distance between the initial and target points, and fast during intermediate motion. 5) When driving to the final target pose, the animation sequence is paused for several frames to obtain the entire animation sequence. 6) Complete the movement of the skeleton from the initial pose to the target pose. This approach is closer to the real-world motion patterns than uniform interpolation, and produces better simulation results.

[0107] Combine Figures 1 to 3 The method for generating a clothing model and a three-dimensional human body model according to the embodiment of the present invention can be implemented by a human body image processing device. Figure 4 FIG. 3 is a schematic diagram showing a hardware structure 300 of a device for processing a human body image according to an embodiment of the present invention.

[0108] The present invention also discloses a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the aforementioned clothing model matching method and steps are implemented.

[0109] And an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement the aforementioned clothing model matching method and steps when executing the program stored in the memory.

[0110] like Figure 4 As shown, the device 300 for implementing virtual fitting in this embodiment includes: a processor 301, a memory 302, a communication interface 303 and a bus 310, wherein the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and communicate with each other.

[0111] Specifically, the processor 301 may include a central processing unit (CPU) or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.

[0112] Memory 302 can include a large capacity memory for data or instructions. For example, but not limitation, memory 302 can include an HDD, a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, memory 302 can include a removable or non-removable (or fixed) medium. In appropriate cases, memory 302 can be inside or outside the processing human body image device 300. In a specific embodiment, memory 302 is a non-volatile solid-state memory. In a specific embodiment, memory 302 includes a read-only memory (ROM). In appropriate cases, the ROM can be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM) or a flash memory or a combination of two or more of these.

[0113] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiment of the present invention.

[0114] Bus 310 comprises hardware, software or both, and the parts of the equipment 300 of processing human body image are coupled to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more above these combinations.In suitable case, bus 310 can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.

[0115] That is to say, Figure 4The device 300 for processing human body images shown can be implemented as including: a processor 301, a memory 302, a communication interface 303 and a bus 310. The processor 301, the memory 302 and the communication interface 303 are connected via the bus 310 and communicate with each other. The memory 302 is used to store program code; the processor 301 reads the executable program code stored in the memory 302 to run the program corresponding to the executable program code, so as to execute the virtual fitting method in any embodiment of the present invention, thereby realizing the combination of Figures 1 to 3 The invention describes a method and apparatus for virtual fitting.

[0116] An embodiment of the present invention further provides a computer storage medium having computer program instructions stored thereon; when the computer program instructions are executed by a processor, the method for processing a human body image provided by an embodiment of the present invention is implemented.

[0117] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0118] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0119] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.

[0120] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.

Claims

1. A virtual fitting method for illumination matching, characterized in that: The method comprises: 1) Obtain a two-dimensional image of the target human body; 2) Obtaining a semantic segmentation map of the target human image; 3) obtaining a two-dimensional image of the clothing, including: obtaining values ​​in an RGB color space of a two-dimensional image of a target human body, converting the RGB values ​​of the two-dimensional image of the target human body into XYZ values ​​of an XYZ color space, converting the XYZ values ​​of the XYZ color space of the target image into values ​​of a Lab color space, combining the Lab image with a semantic segmentation map to obtain an L channel value corresponding to the target human body part, and obtaining a light intensity level corresponding to the target human body part through the L channel value; simultaneously obtaining RGB values ​​of the two-dimensional image of the target human body, converting the RGB values ​​into grayscale values, combining the grayscale image with the semantic segmentation map to obtain a grayscale image corresponding to the target human body part, calculating horizontal and vertical brightness difference values ​​δ of each pixel in a region of interest of the image, and then obtaining an overall image brightness difference statistical feature γ in the grayscale image, calculating a cosine similarity value between a current image and the image brightness difference statistical feature image library by comparing the pre-established image brightness difference statistical feature image library, and selecting the one with the highest feature similarity to estimate the illumination angle of the current image; 4) matching the illumination intensity and illumination angle of the prefabricated garment with the target human body image according to the illumination intensity level and illumination angle; 5) Select appropriate clothing texture maps to create a three-dimensional model of the clothing.

2. The method according to claim 1, characterized in that It also includes a matching process between a clothing model and a human body model, constructing a three-dimensional standard human body model in an initial posture by combining a mathematical model; fitting the three-dimensional clothing model to a three-dimensional standard human body model in the initial posture, i.e., a standard basic human body; obtaining three-dimensional target human body model parameters through a neural network model calculation based on a two-dimensional image of a target human body; inputting the obtained three-dimensional target human body model parameters including several groups of postures and body shapes into the three-dimensional standard human body model for fitting; and obtaining a target human body model having the same posture and body shape as the target human body and wearing the changed clothing.

3. The method according to claim 1, characterized in that Obtain the contour image of the target human body two-dimensional image, substitute the two-dimensional human body contour image into the deep learning neural network model for regression, and obtain the semantic segmentation map of the target human body image.

4. The method according to claim 1, wherein The matching of illumination intensity and illumination angle between the prefabricated garment and the target human body image includes: collecting texture maps of the garment model under different illumination intensities according to different illumination intensity levels, and selecting the texture map with the corresponding illumination intensity according to the illumination intensity of the target human body to render the garment model.

5. The method according to claim 1, characterized in that The matching of illumination intensity and illumination angle between the prefabricated garment and the target human body image further includes: collecting texture maps of the garment model at different illumination angles according to different illumination angles, and selecting the texture map corresponding to the illumination angle according to the illumination angle of the target human body to render the garment model.

6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

7. An electronic device, characterized in that: The invention comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement any method described in claims 1-5 when executing the program stored in the memory.

Citation Information

Patent Citations

  • A virtual wear method and system with image deformation

    CN109035413A

  • Illumination-separated single-view-angle human body clothing geometric detail reconstruction method and device

    CN110310319A