Method for obtaining a digital representation of a person's hair
By employing deep convolutional neural networks to generate realistic hair representations from multiple viewpoints, the method enhances the accuracy and realism of virtual accessory fitting systems, addressing limitations in existing technologies.
Patent Information
- Application Number
- FR2023012918
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-11-23
AI Technical Summary
Existing virtual accessory try-on systems are limited to rendering images from a single viewpoint, primarily a frontal view, and fail to accurately represent accessories worn on the sides, such as earrings, due to the lack of realistic hair representation in 2D digital images.
A method using trained deep convolutional neural networks to obtain a digital representation of a person's hair by analyzing images from different viewpoints, adjusting parameters based on user input or image data, and integrating this representation into a digital twin for enhanced realism in virtual fitting systems.
Enables realistic virtual fitting of accessories by accurately depicting hair from multiple angles, improving the overall realism and functionality of digital twins in augmented reality applications.
Smart Images

Figure 00000026_0000 
Figure 00000026_0001 
Figure 00000027_0000
Abstract
Description
Title of the invention: Method for obtaining a digital representation of a person's hair technical field
[0001] The present invention relates to virtual accessory try-on systems for a digitally represented natural person. In particular, the present invention relates to obtaining a digital representation of a natural person's hair, said digital representation being intended to be used to add hair to a digital twin (avatar) of the natural person. Technological background
[0002] Applications or services that allow a person to virtually try on an accessory are known. These applications or services are based on augmented reality, that is, for example, on a combination of two-dimensional (2D) digital images captured of the person combined with a three-dimensional (3D) rendering of an accessory. It should be noted that augmented reality can also be achieved by combining a volumetric image with a rendering of a 3D scene.
[0003] The combination of 2D and 3D image data from augmented reality-based applications limits the display of the resulting image from the combination of this data to a single viewpoint or at best to different viewpoints whose view axes are very close to each other.
[0004] Current virtual accessory try-on applications or services thus allow a natural person to form an opinion concerning the purchase of these accessories by viewing on a screen an image formed by the combination of the 2D images of this natural person and these accessories.
[0005] These augmented reality-based applications or services are usually limited to rendering images resulting from a frontal view of a person. Indeed, image data representing a frontal view of a person is generally captured because that person is usually looking at the camera capturing the image data. The resulting image from combining this captured image data with data representing an accessory is then a 2D image of that person representing a frontal view. Augmented reality-based applications or services are therefore suitable when the accessory in question is intended to be worn on the person's face, for example, makeup. However, when the accessory is intended to be worn on the side of the person, For example, earrings, these applications or services are not well suited because they only offer 2D digital representations, whereas the physical person may want to have front and side views, for example, before deciding on their purchase of accessories.
[0006] In its French patent application No. FR2301334 filed on February 14, 2023, the applicant describes a virtual fitting system for an accessory by a digitally represented living being. This system makes it possible to create a three-dimensional digital twin JN of a part of a living being, a set of parts of a living being, or a living being in its entirety, and to calculate a 3D view of this digital twin JN wearing a digitally represented accessory.
[0007] When the virtual fitting system is used to try on head accessories such as earrings, makeup, hats, etc., the digital twin JN represents at least the head of the physical person (and possibly other parts of that person), and it is necessary to define a digital representation of the physical person's hair that will be superimposed onto the digital representation of the head of the digital twin JN. This digital representation of a physical person's hair must correspond to the physical person's hairstyle so that the digital twin JN is as realistic as possible and the virtual fitting system for accessories closely resembles a real-life accessory fitting. Summary of the present invention
[0008] One object of the present invention is to resolve at least one of the drawbacks of the technological background.
[0009] Another object of the present invention is to improve the realism of a digital twin representing at least the head of a physical person.
[0010] Another object of the present invention is to improve virtual accessory fitting systems.
[0011] According to a first aspect, the present invention relates to a method for obtaining a digital representation of a person's hair, said digital representation being intended to be used to add hair to a digital twin of the person, said method comprising the following steps: - obtaining at least one first parameter of a parametric hair model as output of a trained parameter estimation model when representative image data of at least the head of the natural person from different points of view are presented as input to the trained parameter estimation model; and - obtaining the digital representation of the hair of the natural person as a function of said at least one first parameter.
[0012] According to a particular and non-limiting embodiment of the present invention, the trained parameter estimation model can be implemented by a first deep convolutional neural network.
[0013] According to a particular and non-limiting embodiment of the present invention, said at least one first parameter can be modified as a function of at least one second parameter.
[0014] According to a particular and non-limiting embodiment of the present invention, said at least second parameter can be obtained from a user action on a human-machine interface or from image data.
[0015] According to a particular and non-limiting embodiment of the present invention, obtaining the digital representation of the hair of the natural person as a function of said at least one first parameter may comprise the following sub-steps: - obtaining a hair class identifier as output from a trained hair classification model among a set of hair classes when image data is presented as input to the trained hair classification model, each hair class identifier referencing a parametric hair model; and - adjustment of the hair model identified by the hair class identifier according to said at least one first parameter.
[0016] According to a particular and non-limiting embodiment of the present invention, the trained hair classification model can be implemented by a second deep convolutional neural network.
[0017] According to a second aspect, the present invention relates to a device for obtaining a digital representation of a person's hair comprising means for implementing the steps of the process according to the first aspect of the present invention.
[0018] According to a third aspect, the present invention relates to a virtual accessory fitting system by a natural person represented digitally by a digital twin comprising a device according to the second aspect of the present invention.
[0019] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0020] Such a computer program can use any programming language, and be in the form of source code, object code, or a intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0021] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0022] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, RAM, CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0023] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0024] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0025] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 6, in which:
[0026] [Fig-1] schematically illustrates a virtual accessory fitting system by a natural person represented digitally by a digital twin, according to a particular and non-limiting embodiment of the present invention;
[0027] [Fig.2] illustrates a flowchart of the different stages of a process for obtaining and rendering the digital twin from image data according to a particular and non-limiting embodiment of the present invention;
[0028] [Fig.3] illustrates a flowchart of the different stages of a process for obtaining a digital representation of the hair of a natural person from image data, according to a particular and non-limiting embodiment of the present invention;
[0029] [Fig. 4] illustrates a flowchart of the different steps of a process for obtaining at least one first parameter of a parametric hair model as output from a trained parameter estimation model when image data are presented as input to the model, according to a particular and non-limiting example of an embodiment of the present invention;
[0030] [Fig.5] illustrates a flowchart of the different steps of a process for obtaining a hair class identifier at the output of a trained hair classification model when image data are presented as input to the model, according to a particular and non-limiting embodiment of the present invention;
[0031] [Fig.6] schematically illustrates the device configured to obtain a digital representation of a person's hair represented digitally by a digital twin, according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0032] A method and device for obtaining a digital representation of a person's hair and a virtual accessory fitting system by a person represented digitally by a digital twin, will now be described in what follows with joint reference to Figures 1 to 6. The same elements are identified with the same reference signs throughout the following description.
[0033] Fig. 1 schematically illustrates a virtual accessory fitting system 1 by a physical person 10 represented digitally by a digital twin JN, according to a particular and non-limiting embodiment of the present invention.
[0034] The system 1 includes a device 101 for obtaining DIC image data representative of at least the head of the natural person 10 from different points of view, a device 102 for obtaining a digital RNT representation of at least the head of the natural person 10 in a three-dimensional geometric space from the DIC image data, a device 103 for obtaining a digital RNC representation of the hair of the natural person 10 in the three-dimensional geometric space from the DIC image data, a device 104 for obtaining DAC accessory data representative of a digital representation of at least one accessory in the three-dimensional geometric space and a device 105 for rendering the digital twin JN representative of at least the head of the natural person 10 from the RNC and RNT representations and the DAC accessory data.
[0035] The three-dimensional geometric space is a space in which the part or set of parts of the natural person 10 or the natural person 10 in its entirety is geometrically represented from the DIC image data and the camera or cameras used to capture the images from which this DIC image data is obtained.
[0036] The digital twin JN, also called a realistic avatar, is a set of three-dimensional digital data that represents in the metaverse at least the head of the physical person 10 that is represented by the DIC image data.
[0037] The geometry and texture of the digital twin JN are defined by the RNT representation. The DAC data can then be added to the digital twin JN to visualize the virtual accessory port by the digital twin JN.
[0038] According to one variant, system 1 may further include a human-machine interface (HMI).
[0039] For example, the HMI can be partly implemented by several of the devices 101 to 105.
[0040] For example, it may allow a user to enter parameters by graphical means such as a keyboard or a touch screen of device 101 or 105.
[0041] According to one variant, devices 101, 104 and 105 can be implemented in separate devices.
[0042] According to one variant, devices 101, 104 and / or 105 can be implemented in the same device.
[0043] According to an example of an embodiment of this variant, devices 101, 104 and / or 105 can be implemented in the same mobile communication device.
[0044] Depending on variations, device 101, 104 and / or 105 may be one of the following devices: - a telephone; - a computer; - a tablet; - a pair of glasses equipped with at least one camera and a screen or in communication with at least one camera and a screen; - a pair of ocular contact lenses equipped with at least one camera and a screen that adapts to a user's view.
[0045] The present invention is not limited to these examples of devices but can be extended to any device that would be configured to implement devices 101 to 105.
[0046] Devices 102 and 103 are computers (CPU for Computer Processing Units, or GPU for Graphical Processing Unit) which are associated with buffer memories to perform computational tasks.
[0047] Devices 101 to 105 can be connected, for example, via wired network communication (e.g., via Ethernet and / or a fiber optic link) and / or via a wireless network such as Wifi® (according to IEEE 802.11 or one of its variants) or via a cellular network such as LTE (Long-Term). Evolution (or in French, "Evolution à long ternie"), LTE-Advanced (or in French, LTE-avancé), and / or 5G and / or 6G. The term "cloud computing" refers to this type of architecture, which uses the memory and computing power of computers and servers distributed worldwide and linked by a communication network. Such an architecture is illustrated in [Fig. 1] by a cloud.
[0048] The devices lOlà 105 are thus configured to transmit data to the "cloud" 100 and / or to receive data from the "cloud" 100.
[0049] The mobile communication infrastructure enabling wireless data communication between devices 101 to 105 and the "cloud" 100 may include, for example, one or more communication devices (not shown in [Fig. 1]) of the relay antenna type (cellular network). In a communication mode using such a network architecture, data destined for devices 101 to 105 may be received, for example, from the "cloud" 100, and data may be transmitted by devices 101 to 105 via the "cloud" 100 through one or more relay antennas (each relay antenna being, for example, connected to the "cloud" 100 via a wired link).
[0050] DIC image data, RNT and RNC representational data and DAC accessory data are exchanged by devices 101 to 105 via the "cloud" 100.
[0051] The present invention is not limited to the communication architecture between devices 101 to 105 illustrated in [Fig.1] but extends to any type of communication architecture.
[0052] For example, when devices 101, 104 and 105 are the same mobile communication device, this mobile communication device can be configured to transmit DIC image data to devices 102 and 103 via the "cloud" 100 and to receive data representing the RNT and RNC representations via the "cloud" 100.
[0053] According to another example, devices 102 and 103 can be the same device which is then configured to receive DIC image data via the "cloud" 100 and to emit data representing the RNT and RNC representations via the "cloud" 100.
[0054] Other variants of communication architecture between devices 101 to 105 are obviously conceivable without going outside the scope of the present claimed invention.
[0055] Fig. 2 illustrates a flowchart of the different stages of a process for obtaining and rendering the digital twin JN from DIC image data according to a particular and non-limiting embodiment of the present invention.
[0056] In a first step 21, DIC image data representative of at least the head of the natural person 10 from different points of view are obtained from the device 101.
[0057] In a second step 22, the RNT digital representation is obtained from the DIC image data.
[0058] In a third step 23, the RNC digital representation is obtained from the DIC image data.
[0059] In a fourth step 24, the DAC accessory data are obtained.
[0060] In the fifth step 25, the digital twin JN defined from the digital representations RNT and RNC and the DAC accessory data is rendered, by calculating a view of this digital twin JN from a point of view.
[0061] Device 101 includes an image capture means.
[0062] According to an example of an embodiment of the present invention, the image capture means may include at least one camera corresponding for example to: - an infrared camera; - an RGB type image acquisition camera (from the English "Red, Green, Blue" or in French "Rouge, vert, bleu"); - a camera platform (called Lightfield) or RGB type Volumetric camera rig; - a LiDAR-type acquisition camera (known as "light detection and ranging" or "laser imaging detection and ranging" in English); or - an image acquisition camera associated with one or more devices (for example one or more LEDs (Light-Emitting Diode) emitting light in the infrared or near-infrared band.
[0063] According to one embodiment of the present invention, the image capture means may correspond to a camera of a mobile communication device comprising a camera.
[0064] According to variants, the captured images may correspond either to still images captured from different viewpoints of at least the head of the natural person 10 or to images from a video captured of the natural person 10 which represents at least the head of said natural person 10 from different viewpoints.
[0065] According to one embodiment of the present invention, the image capture means may be a scanning system comprising a plurality of cameras synchronized with each other to capture several images from different viewpoints of minus the head of the physical person. Captured image data then represents these captured images from different viewpoints.
[0066] In the case of a camera platform (sometimes referred to as "Lightfield" in English and "volumetric image data" in French), the cameras can be positioned in a circular or planar manner on a physical capture device and can capture in a single iteration or several (if motion capture) at least the head of the physical person.
[0067] According to one embodiment of the present invention, the image capture means may include at least one volumetric sensor enabling the acquisition of volumetric image data (in English, "lightfield").
[0068] According to one embodiment of the present invention, the device 101 can be configured so that the DIC image data corresponds to image data of the images captured by the image capture means.
[0069] Typically, three images are captured, one representing the head of the physical person 10 from the front, another representing it from a left profile and another representing it from a right profile.
[0070] According to one embodiment of the present invention, the device 101 can be configured to apply processing to the images captured by the image capture means and the DIC image data corresponds to the processed image data.
[0071] According to one embodiment of the present invention, if a video can be captured by means of image capture, a processing step can be the selection of images representing at least the head of the individual from among the images in this video. Typically, three images are selected, one representing the head from the front, another representing it from a left profile, and the third representing it from a right profile.
[0072] According to one embodiment of the present invention, the hair of the natural person depicted in the captured images can be segmented from the rest of the captured images; that is, at least one spatial area delimiting the hair on the natural person's head is identified in each of the captured images. The DIC image data can then correspond to the image data of the pixels contained within the spatial areas defined in the captured images.
[0073] For example, the segmentation of hair from the rest of a captured image can be implemented from a model based on a PSPNet-type architecture developed by Li et al. (Li, T., Bolkart, T., Black, M., Li, H., & Romero, J. (2017) or BiSeNet (https: / / arxiv.org / abs / 2004.02147) Leaming a Model of Facial Shape and Expression from 4D Scans. ACM Trans. Graph., 36(6)).
[0074] According to one variant, the DIC image data may further include information representing image content and / or descriptive geometry information of at least the head of the physical person in a three-dimensional geometric space.
[0075] The device 102 is configured to obtain the RNT digital representation of at least the head of the physical person 10 in three-dimensional geometric space from the DIC image data.
[0076] According to a particular and non-limiting embodiment of the present invention, the device 102 can be configured to implement the method described in French patent application no. FR2301334 filed on February 14, 2023.
[0077] According to this patent application, obtaining the RNT digital representation of at least the head of the natural person 10 includes obtaining DM mesh data (“mesh” in English) representative of a three-dimensional mesh structure represented in three-dimensional geometric space by a set of vertices of polygonal shapes forming a structure, each polygonal shape, for example triangles or quadrangles, sharing at least one side with another polygonal shape.
[0078] The device 102 is configured to calculate DM mesh data of the digital twin JN by modifying DMI mesh data of an initial digital twin JNI according to morphological information extracted from DIC image data.
[0079] The initial digital twin JNI is a digital representation of at least one head of an asexual human being. This digital representation comprises the DMI mesh data representative of a three-dimensional mesh structure represented in three-dimensional geometric space by a set of vertices of polygonal shapes forming a structure, each polygonal shape sharing at least one side with another polygonal shape.
[0080] According to a particular and non-limiting embodiment, the DMI mesh data can be modified by changing the spatial positions of the vertices defined in the three-dimensional geometric space.
[0081] The three-dimensional mesh structure of the initial JNI twin is then deformed (stretched, compressed, etc.) so that the modified three-dimensional mesh structure of the initial JNI twin resembles in three-dimensional geometric space the geometry of at least the head of the physical person 10.
[0082] The DM mesh data of the digital twin JN are then equal to the modified DMI data.
[0083] The DM data of the digital twin JN are obtained by an iterative process of deforming the DMI data of the digital twin JNI, which incorporates all the knowledge of the morphology of a living being such as, for example, all the morphology of a Caucasian, Afro, Asian human being and / or of an animal being, that is to say the position, a structure and a deformation of the eyes, mouth and skin folds, an underlying muscular system and a series of controllers enabling the deformations.
[0084] An iterative deformation process can be based on a neural network (Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications, arXiv preprint arXiv: 1704.04861, 2017) and / or on a deep learning method (Ayush Tewari, Michael Zollôfer, Hyeongwoo Kim, Pablo Garrido, Florian Bernard, Patrick Perez, and Theobalt Christian. MoFA: Model-based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction. In The IEEE International Conference on Computer Vision (ICCV), 2017.) and / or on a reinforcement learning method and / or using generative adversarial networks (GANs). (English "Generative Adversarial Network")•
[0085] In the case where a deep learning method is used for the reconstruction of the face of the physical person 10, a parametric face modeling method known as “3DMM” can be used. This is based on a statistical model of human faces that allows any face to be represented as a weighted sum of 3D mesh structures, a 3D mesh structure which presupposes the prior creation of a suitable 3D topology. Only the frontal part of the face – the only deformable part – is taken into account (this is referred to as the “monkey mask” in English). The geometric aspect is enriched by the representation of deformations corresponding to facial expressions. This involves a set of several parameters, typically around 150 parameters for facial identity and 100 parameters for expressions.Estimating these hundreds of parameters can be done by regression of these parameters using a deep learning method learned on a large set of face images (Ayush Tewari, Michael Zollôfer, Hyeongwoo Kim, Pablo Garrido, Florian Bernard, Patrick Perez, and Theobalt Christian. MoFA: Model-based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction. In The IEEE International Conference on Computer Vision (ICCV), 2017), and (Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Large-scale celebfaces attributes (celeba) dataset. Retrieved August 15, 2018).
[0086] Thus, during this iterative deformation process, a first three-dimensional (3D) mesh structure defined by the DNI data is positioned in the three-dimensional geometric space. This first mesh structure The three-dimensional image is then compared to three-dimensional mesh structures stored in a database, each stored 3D mesh structure representing at least the face of the physical person. The iterative process also compares each two-dimensional (2D) image calculated from the first 3D mesh structure to 2D images calculated from a 3D mesh structure stored in the database. The comparison process of these elements seeks to align the 2D images calculated from the first 3D mesh structure with 2D images calculated from 3D mesh structures stored in the database.The process iteratively improves the distances between the vertices of the first 3D mesh structure and the vertices of another 3D mesh structure. This is achieved through micro-displacements of the vertices of the first 3D mesh structure within three-dimensional geometric space, thereby optimizing the resulting first 3D mesh structure. This spatial deformation operation of the first 3D mesh structure proceeds through successive iterations until a resemblance is obtained between the first 3D mesh structure and a 3D mesh structure stored in the database that geometrically represents at least the head of the physical person. The resulting 3D mesh structure corresponds in resemblance to DIC image data.
[0087] According to this patent application, obtaining the RNT digital representation of at least the head of the natural person 10 may further include obtaining DT texture data from texture data extracted from DIC image data.
[0088] According to a particular and non-limiting embodiment of the present invention, the DT texture data can be texture data extracted from a defined area on the face of the physical person 10 from the DIC image data.
[0089] According to one embodiment, the area texture data can be obtained by assembling DIC image data, isolated, distributed and assembled to reconstruct a single texture.
[0090] For example, the assembly process can deform the texture of each of the images obtained from the DIC image data by modifying the color and contrast of the pixels of the texture images, corresponding to the texture of the images obtained from the DIC image data, and by correcting the boundaries between the texture images to generate a unique texture.
[0091] This unique texture can then be projected onto the surface of the 3D mesh structure corresponding to the DM data (a projection known as UV mapping, where U and V denote the axes of the 2D texture). This then allows for a reprojection of the assembled texture onto the 3D mesh structure.
[0092] For example, texture data can be obtained from a surface of a defined area between the forehead, ears and chin (“monkey mask” in English).
[0093] The device 105 is configured to obtain DAC accessory data representative of a three-dimensional representation of at least one accessory in three-dimensional geometric space. This three-dimensional representation is intended to be added to the digital twin JN.
[0094] DAC accessory data is known to those skilled in the art. It represents all types of objects, makeup. We often speak of a 2D / 3D layer which incorporates a material, a transparency, a relief and which is calculated by a shading and rendering process.
[0095] The device 103 is configured to obtain the RNC digital representation of the hair of the natural person 10 from the DIC image data. For this purpose, the device 103 is configured to implement the method of [Fig.3].
[0096] Figure 3 illustrates a flowchart of the different steps of a process for obtaining (step 23) the RNC digital representation of the hair of the natural person 10 from DIC image data, according to a particular and non-limiting embodiment of the present invention.
[0097] In a step 231, the device 103 obtains at least one first parameter PI of a parametric hair model as output from a trained MP parameter estimation model when DIC image data representing at least the head of the physical person from different viewpoints are presented as input to the trained MP parameter estimation model.
[0098] According to a variant of step 231, said at least one first parameter PI is modified as a function of at least one second parameter P2.
[0099] For example, said at least one parameter PI can be replaced by said at least second parameter P2, for example by concatenating said first and second parameters.
[0100] According to a particular and non-limiting example of an embodiment, said at least second parameter P2 can be obtained from a user action on the HMI of the virtual fitting system of at least one accessory.
[0101] This variant allows a user, for example a hairdresser, to modify a hairstyle (haircut) by changing at least one first parameter PI of its digital representation according to at least one second parameter P2 that they choose. This hairdresser can then present their client with different variations of the haircut using the virtual accessory try-on system 1.
[0102] According to a particular and non-limiting embodiment, said at least second parameter P2 can be obtained from DIC image data.
[0103] For example, said at least a second parameter P2 can be obtained by known image processing methods such as linear filtering, morphological operations for example to obtain color information.
[0104] According to another particular and non-limiting embodiment, said at least second parameter P2 may be a scene lighting parameter in which the digital twin JN evolves or is defined according to data that define a scene choice.
[0105] In a step 232, the numerical representation RNC is obtained as a function of said at least one first parameter PI.
[0106] According to a particular and non-limiting embodiment, the second step 232 may include a step 2321 of obtaining representations of hair strands by parametric curves from the DIC image data, and said at least one first parameter PI may be a parameter of said parametric curves. The RNC digital representation may then be defined by said parametric curves.
[0107] For example, the parametric curve can be an interpolation spline (sometimes called a circle in English). An interpolation spline is a piecewise function consisting of a polynomial on each interval defined between two nodes, which can be considered points in three-dimensional geometric space. The degree of the interpolation spline is defined by the highest-degree polynomial used: if the nodes are simply joined by straight lines, that is, if the interpolation spline consists of first-degree polynomials, the interpolation spline is of degree 1, and so on. If all the polynomials have the same degree, it is called a uniform interpolation spline. If all the polynomials are of degree three, it is called a cubic interpolation spline. The strands of hair represented by interpolation splines can then vary in length, thickness and curvature according to coefficients of the polynomials defining the interpolation splines.These coefficients can then be said to be at least one first parameter PI or be determined from said at least one first parameter PL.
[0108] According to another particular and non-limiting embodiment, the second step 232 may include a step 2322 of obtaining a hair class identifier (IDC) as output from a trained MC hair classification model among a set of hair classes when DIC image data are presented as input to the trained MC hair classification model, each hair class identifier referencing a parametric hair model. The second step 232 may further include a step 2323 of fitting the model to hair identified by the hair class identifier IDC as a function of said at least one first parameter PI.
[0109] According to a variant of the method in [Fig.3], the MP model for parameter estimation is implemented by a first deep convolutional neural network and the MC model for hair classification is implemented by a second deep convolutional neural network.
[0110] A deep convolutional neural network, also called a convolutional neural network and denoted CNN or ConvNet (from the English "Convolutional Neural Networks"), corresponds to an acyclic artificial neural network (from the English "feed-forward"). Such a convolutional neural network comprises a convolutional part implementing one or more convolutional layers and a densely connected part implementing one or more layers of densely connected (or fully connected) neurons ensuring the classification of information according to a model of the MLP type (from the English "Multi Layers Perceptron" or in French "Multilayer Perceptrons") for example.
[0111] According to the present invention, the first and second convolutional neural networks Each can incorporate a MIL (multi-instance learning) architecture that allows for the labeling of all input images for these convolutional neural networks, which represent the same head of hair from different viewpoints. The first and second convolutional neural networks then learn how a person's hair is represented from different viewpoints and which initial parameters best correspond to this representation of the person's hair.
[0112] Figure 4 illustrates a flowchart of the different steps of a process for obtaining (step 231) at least one first PI parameter of a parametric hair model at the output of the trained MP model for parameter estimation when DIC image data are presented as input to the MP model, according to a particular and non-limiting embodiment of the present invention.
[0113] In a substep 2311 of step 231, the device 103, implementing the first convolutional neural network, provides the DIC image data as input to a so-called convolutional part of the first convolutional neural network to determine DI data representative of the hair characteristics of the natural person 10. The convolutional part advantageously implements one or more convolutional layers whose objective is to detect the presence of these characteristics, which may be aesthetic and / or mechanical, of the natural person's hair, or to detect patterns associated with said characteristics sought in the DIC image data input to The convolutional part. For this purpose, a set of at least one convolution filter is applied to the DIC image data. This DIC image data can be provided, for example, as a two-dimensional matrix (corresponding to the pixel grid of an image, for instance). For example, 50, 100, 200, or more convolution filters can be applied, each filter having a specific size and stride. For example, each filter could have a size of 1x7 and a stride of 1. In other words, the input matrix containing the DIC image data passes through a first convolutional layer, and convolutional operations are applied to this input matrix based on the filters of determined size and stride. The stride corresponds to the number of pixels by which the window corresponding to the filter moves within the input tensor (the input matrix, for example).Each convolution filter can represent, for example, a specific feature sought in an image by dragging the window corresponding to the filter over the image and calculating the convolution product between that specific feature and each portion of the image scanned by the filter associated with that specific feature. The result of the convolution product allows us to determine the presence or absence of the specific feature in the input image.
[0114] According to one variant, one or more successive convolution operations can be applied to the DIC image data, with each convolution operation applying a set of convolution filters to the data obtained as output from the previous convolution operation.
[0115] According to one variant, one or more so-called "pooling" operations (for example, one or more "max pooling" operations and / or one or more "average pooling" operations) can be implemented and applied to the data obtained from the convolution operation(s). According to this variant, the "pooling" operation(s) can, for example, and optionally, be accompanied or associated with a random deactivation operation of a portion of the neurons in the network, with a determined probability, for example, a probability of 0.2 (20%), 0.3 (30%), or 0.4 (40%), this technique being known as "dropout".
[0116] In a substep 2312 of step 231, the device 103 provides the DI data as input to the so-called densely connected part of the first convolutional neural network to determine a set of at least one first PL parameter
[0117] The densely connected layer implements a DI data classification operation. To this end, the densely connected part of the first convolutional neural network comprises one or more layers of densely connected neurons or fully connected. Each layer of neurons provides output data obtained by applying a linear combination and optionally an activation function to the data present at its input.
[0118] The densely connected layer returns as output a vector of size N, where N corresponds to the number of sets of at least one first parameter PI, each set of at least one first parameter PI corresponding to a parametric hair model. Each element of the output vector provides the probability that the DI data corresponds to a set of at least one first parameter PI.
[0119] The input vector can pass for example through a layer of densely connected neurons, with for example 128, 256 or more neurons each connected to each of the neurons in a layer comprising as many neurons as there are sets of at least one first PI parameter, a set of at least one first PI parameter of a parametric hair model being associated with each neuron of the output layer of the densely connected part.
[0120] Figure 5 illustrates a flowchart of the different steps of a process for obtaining (step 2322) a hair class identifier (IDC) at the output of the trained MC hair classification model when DIC image data are presented as input to the MC model, according to a particular and non-limiting embodiment of the present invention.
[0121] In a substep 23221 of step 2322, the device 103 implementing the second convolutional neural network provides the DIC image data as input to a so-called convolutional part of the second convolutional neural network to determine DI data representative of the hair characteristics of the natural person 10. The convolutional part advantageously implements one or more convolutional layers whose objective is to detect the presence of these characteristics, which may be aesthetic and / or mechanical, of the hair of the natural person or to detect patterns associated with said characteristics sought in the DIC image data introduced as input to the convolutional part.To this end, a set of at least one convolution filter is applied to the DIC image data. This DIC image data can be provided, for example, as a two-dimensional matrix (corresponding to the pixel grid of an image, for instance). For example, 50, 100, 200, or more convolution filters can be applied, each filter having a specific size and stride. For example, each filter has a size of 1x7 and a stride of 1. In other words, the input matrix containing the DIC image data passes through a first convolution layer, and convolution operations are applied to this input matrix based on the filters of size and . The step size (or "stride") is the number of pixels by which the filter window moves within the input tensor (for example, the input matrix). Each convolution filter can represent, for example, a specific feature sought in an image by dragging the filter window across the image and calculating the convolution between that feature and each portion of the image scanned by the filter associated with that feature. The result of the convolution determines the presence or absence of the feature in the input image.
[0122] According to one variant, one or more successive convolution operations can be applied to the DIC image data, with each convolution operation applying a set of convolution filters to the data obtained as output from the previous convolution operation.
[0123] According to yet another variant, one or more so-called "pooling" operations (for example, one or more "max pooling" operations and / or one or more "average pooling" operations) can be implemented and applied to the data obtained from the convolution operation(s). According to this variant, the "pooling" operation(s) can, for example, and optionally, be accompanied or associated with a random deactivation operation of a portion of the neurons in the network, with a determined probability, for example, a probability of 0.2 (20%), 0.3 (30%), or 0.4 (40%), this technique being known as "dropout".
[0124] The convolutional part of the second convolutional neural network outputs a set of detected feature(s) from the DIC image data provided as input to the convolutional part. Information representing the presence of a given feature(s) is provided as input to a so-called densely connected part of the second convolutional neural network.
[0125] In a substep 23222 of step 2322, device 103 provides DI data as input to the so-called densely connected part of the second convolutional neural network to determine the hair class identifier IDC.
[0126] The densely connected layer implements a DI data classification operation. To this end, the densely connected part of the second convolutional neural network comprises one or more layers of densely connected or fully connected neurons. Each layer of neurons provides output data obtained by applying a linear combination and optionally an activation function to the data present at its input.
[0127] The densely connected layer returns as output a vector of size N, where N corresponds to the number of sets of at least one first parameter PI of hair classes, each hair class corresponding to a parametric hair model. Each element of the output vector provides the probability that the DI data correspond to a hair class.
[0128] The values of the convolutional filters of the first and second convolutional neural networks and the parameters of their densely connected layers are advantageously determined in a so-called learning phase, for example supervised learning, according to a method known to those skilled in the art. In a learning phase, a large number (for example, hundreds, thousands, tens of thousands, or more) of hair images whose associated characteristics are known are used to learn the different values or coefficients of the convolutional filters. In such a learning phase, a method known as backpropagation of the error gradient can, for example, be implemented.Similarly, in this learning phase, a large number (e.g., hundreds, thousands, tens of thousands or more) of data associations representative of hair characteristics are used to learn the parameters of the densely connected part enabling data classification in order to obtain a set of at least one first PI parameters of a parametric hair model (first convolutional neural network) or a hair class identifier (second convolutional neural network).
[0129] Learning can, for example, be carried out using synthetically obtained image data.
[0130] According to one embodiment, a set of at least one first PI parameter (first convolutional neural network) or a hair class (second convolutional neural network) can be determined for each image in a sequence of consecutive images, from the DIC image data of those images. A plurality of sets of at least one first PI parameter (or hair classes) can thus be obtained for the sequence, with as many sets of at least one first PI parameter (or hair classes) as there are images in the sequence. A temporal filtering (for example, an average) can advantageously be applied to this plurality of sets of at least one first PI parameter (or hair classes) to determine a single set of at least one first PI parameter (an identifier of a hair class).Such temporal filtering can, for example, be achieved or implemented by a digital filter known to a person skilled in the art, for example defined by a mathematical equation, such as a difference equation, or by a neural network, for example recurrent.
[0131] According to one variant, the image data used for training the first and second convolutional neural networks can be data from groups of three images taken from the front, right profile, and left profile of the same head of hair. Each group of images can be associated with annotations that indicate the expected hair class corresponding to the three images in the image group and that indicate a value of at least one first PI parameter of a parametric hair model.
[0132] According to one variant, the image data used for training the first and second convolutional neural networks can be synthetically obtained image data, which facilitates the generation of a large amount of image data required for training the first and second convolutional neural networks.
[0133] The synthetic image data used for training can be generated from parametric hair models and sets of parameters of these models chosen randomly to obtain hair variations. Variations in viewpoints can also be used to introduce variations in the viewpoints of the hair images.
[0134] A problem of generalizing the training of the first and second convolutional neural networks arises when image data for training is synthetically generated from images of hair variations, where several of these images contain the same face. To avoid this problem, the image data used for training can be segmented to identify at least one spatial region of these images that includes image data representative of the hair. The image data used for training can then correspond to created images that contain only these spatial regions of the images and therefore exclude the spatial regions corresponding to the faces.
[0135] Device 104 is configured to render a digital avatar representative of the physical person from the digital twin JN (of the RNT and RNC digital representations) and the DAC accessory data. To this end, device 104 calculates data representative of a view from a viewpoint of the RNT representation combined with the DAC accessory data and the RNC digital representation.
[0136] The device 104 for rendering a digital avatar representing the physical person wearing at least one accessory may include a rendering engine that calculates data representative of an image of the digital twin JN wearing said at least one accessory from a viewpoint. The image can then be displayed on a screen of the device 104.
[0137] According to one embodiment, the rendering engine can be a real-time rendering engine.
[0138] According to one embodiment, the rendering engine can be implemented on a GPU computing server.
[0139] For example, the real-time rendering engine can be a 3D simulation engine.
[0140] According to a particular and non-limiting example of an embodiment, said at least one The PI parameter can be applied to a parametric hair model using software.
[0141] The methods in Figures 2 to 5 allow a person to virtually try on accessories and decide whether or not to purchase them.
[0142] Through the user interface (UI) of the virtual accessory fitting system 1, a person can access a sales website and select accessories. They can also choose to virtually try on these accessories. The UI can then be configured so that this person adds the selected accessories to a shopping cart, which can be called a virtual fitting room. The process described above is then implemented so that a digital twin of this person is calculated and the person can view a picture of their digital twin wearing the accessories they have selected. Thus, the person can interact with their digital twin from this UI, which displays images of this digital twin wearing these accessories from viewpoints that can be selected by the person. The person can also choose the conditions of the scene (lighting, scene content) in which their digital twin is displayed.The person can also choose lighting conditions for the scene in which their digital twin is displayed, or even change the scene in which the digital twin is presented.
[0143] Figure 6 schematically illustrates the device 103 configured to obtain a digital representation of a person's hair represented digitally by a digital twin, according to a particular and non-limiting embodiment of the present invention.
[0144] According to a particular embodiment, the device 103 corresponds to a server or a computer of the "cloud" 100.
[0145] Device 103 is, for example, configured to carry out the operations described opposite Figures 1 and / or the steps of the processes described opposite Figures 2 to 5. The elements of device 103, individually or in combination, can be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 103 can be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0146] The device 103 comprises one (or more) processor(s) 1030 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 103. The processor 1030 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 103 further comprises at least one memory 1031, for example, volatile and / or non-volatile memory, and / or a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0147] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 1031.
[0148] According to various specific and non-limiting embodiments, the device 103 is coupled in communication with other similar devices or systems and / or with communication devices.
[0149] According to a particular and non-limiting embodiment, the device 103 includes a block 1032 of interface elements for communicating with external devices, for example a remote server or the "cloud" 100. The interface elements of the block 1032 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); - HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French).
[0150] According to a particular and non-limiting embodiment, the device 103 can provide output signals to one or more external devices, such as a display screen 1033, touch or not, one or more loudspeakers 1034 and / or other peripherals 3601034 (projection system for example) via output interfaces 1036, 1037, 1038 respectively. According to a variant, one or more of the external devices is integrated into the device 103.
[0151] Of course, the present invention is not limited to the embodiments described above but extends to a method for obtaining a digital representation of a person's hair which would include steps secondary applications without falling outside the scope of the present invention. The same would apply to a system configured for the implementation of such a method.
Claims
Demands
1. Method of obtaining a digital representation of a person's hair, said digital representation being intended to be used to add hair to a digital twin of the person, said method comprising the following steps: - obtaining (231) at least a first parameter of a parametric hair model as output of a trained parameter estimation model when image data representative of at least the head of the person from different viewpoints are presented as input to the trained parameter estimation model;and - obtaining (232) the digital representation of the natural person's hair as a function of said at least one first parameter, characterized in that obtaining (232) the digital representation of the natural person's hair as a function of said at least one first parameter comprises the following substeps: - obtaining (2322) a hair class identifier as output from a trained hair classification model from among a set of hair classes when image data are presented as input to the trained hair classification model, each hair class identifier referencing a parametric hair model; and - fitting (2323) the hair model identified by the hair class identifier as a function of said at least one first parameter.
2. A method according to claim 1, wherein the trained parameter estimation model is implemented by a first deep convolutional neural network.
3. A method according to claim 1 or 2, wherein said at least one first parameter is modified as a function of at least one second parameter.
4. Method according to claim 3, wherein said at least second parameter is obtained from a user action on a human-machine interface or from image data.
5. A method according to claim 4, wherein the trained hair classification model is implemented by a second deep convolutional neural network.
6. Device for obtaining a digital representation of a person's hair comprising means for carrying out the steps of the process according to any one of claims 1 Q
7. 1 d J. Virtual accessory fitting system by a natural person represented digitally by a digital twin comprising a device according to claim 6.
8. Computer program comprising instructions for carrying out the method according to any one of claims 1 to 5, when such instructions are executed by a processor.
9. Computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 5.