Method and device for three-dimensional reconstruction of a face with toothed part from a single image
By enhancing the photometric characteristics of the toothed portion in a 2D image using deep learning, the method achieves accurate 3D reconstruction for simulating dental treatments, addressing the limitations of existing techniques in reconstructing toothed portions from single images.
Patent Information
- Application Number
- FR2020005928
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-06-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-06-06
AI Technical Summary
Existing facial reconstruction techniques struggle to accurately reconstruct the toothed portion of a face from a single 2D image due to its high specularity and low texture, limiting the ability to simulate dental treatments effectively.
A method involving deep learning algorithms to enhance the photometric characteristics of the toothed portion in a 2D image, followed by 3D reconstruction, allowing for the simulation of dental treatments by merging the enhanced toothed portion with the facial portion to create a comprehensive 3D surface.
Enables realistic 3D reconstruction of faces with visible teeth from a single 2D image, facilitating accurate simulation of dental treatments such as tooth whitening, realignment, and prosthetics, enhancing patient consultation and treatment planning.
Smart Images

Figure 00000028_0000 
Figure 00000029_0000 
Figure 00000030_0000
Abstract
Description
Title of the invention: Method and device for three-dimensional reconstruction of a face with toothed part from a single image Technical field
[0001] The present invention relates generally to three-dimensional facial reconstruction, and more particularly to a method and a device for three-dimensional reconstruction of a face having a toothed part, from a single image, as well as to a computer program product implementing the method.
[0002] The invention finds applications, in particular, in digital processing techniques in the field of dentistry.
[0003] It also proposes a method for simulating the aesthetic result of a planned dental treatment for a human subject, from at least one two-dimensional (2D) color image of the subject's face with a visible toothed portion. The planned dental treatment may be of an aesthetic, orthodontic or prosthetic nature. Prior art
[0004] Three-dimensional (3D) facial reconstruction is a rapidly expanding field and has a wide variety of applications. Previously used primarily in the audiovisual industry, it is now finding other applications, particularly in the simulation of aesthetic treatments.
[0005] Among the aesthetic treatments, we can mention dental treatments (teeth whitening, veneer placement, dental realignment, prosthetic work, etc.). For these treatments, the patient must often undertake the treatment without being able to assess in advance for himself the aesthetic result that this treatment will produce. He must rely on the expertise of the practitioner for this. There is a need to be able to give the patient the benefit and advantages of a simulation capable of realistically presenting the aesthetic result of this type of treatment before his actual undertaking, to possibly choose to modify the planned treatment, in consultation with the practitioner.
[0006] Some facial reconstruction techniques already exist. But most of these techniques are based either on heavy technologies in terms of digital processing, i.e. in processing time and computing resources required to execute a reconstruction algorithm, or in terms of hardware. Other techniques are based on the acquisition of several passive 2D images in order to be able to work in photogrammetry. However, many commercial devices only have a single 2D sensor, which complicates the use of this type algorithm.
[0007] Some technical solutions can meet the need expressed above, notably thanks to artificial intelligence (AI). Increasingly advanced deep learning networks make it possible to reconstruct the face in 3D, from a single 2D image, and this with an increasingly realistic rendering. Unfortunately, the toothed portion very often remains an element inaccessible to deep learning algorithms due to their particular character, i.e. very specular and little textured.
[0008] In the article by Wu et al., "Model-based teeth reconstruction", ACM Transactions on Graphics 35, Article number 220, pp.1-13, November 2016, the authors disclose a parametric solution for the reconstruction of the toothed portion. However, the latter is only compatible with calibrated sensors. Due to its parametric nature, however, it is only applicable to teeth deviating only to a certain extent from a tooth taken as standard.
[0009] Document US2018174367A discloses an augmented reality visualization system for a model that allows the simulated result of a planned dental treatment to be seen directly, and also offers the possibility of interacting with this model to modify it in real time. The system operates by acquiring video data (therefore relating to a plurality of images), simulating dental treatments on this video data, and rendering the result in augmented reality. If a 3D scan of the toothed part is available, it can be registered on an image, with or without the simulation of the planned treatment. Otherwise, a simulation can be made on video data, with however the double disadvantage of the need to have two image sensors, on the one hand, and a result of the simulation that is coarse, on the other hand.
[0010] Document US2018110590A discloses a simulation method in which a dental arch is digitized in 3D on which it is planned to apply a dental treatment (fitting of rings, crowns, aligners, etc.), then, in an augmented reality system, the 3D dental arch is aligned, comprising the simulation of the projected dental treatment on the real image of the patient which is in 2D, with the aim of visualizing in this system no longer the real dental arch of the patient but this arch with the result of the projected dental treatment. However, since the method for aligning the 3D data of the modified arch in the 2D space of the real image of the patient is not explained, this method appears insufficiently described to be able to be reproduced by a person skilled in the art. Statement of the invention
[0011] The invention aims to make possible facial reconstruction, i.e., 3D reconstruction, of the face of a human subject with a visible toothed portion, from a series of any 2D images or possibly from any single 2D image of the face with the toothed portion, the 3D reconstruction thus obtained lending itself well to the affixing in the 3D domain of the result of the simulation of a projected dental treatment which modifies the toothed portion.
[0012] This aim is achieved by a method comprising the separation of the 2D image of the face into a part corresponding to the toothed part alone and another part corresponding to the rest of the face, the first part being subjected to a digital enhancement treatment before merging with the second, either at the 2D level or at the 3D level. The 3D reconstruction, or 3D surface, thus obtained is suitable for the simulation of a projected dental treatment to be applied to the toothed portion of the face, by substituting for the area of the 3D surface corresponding to said toothed portion another 3D surface corresponding to said toothed portion as it would be after said projected treatment. More specifically, a three-dimensional, 3D, reconstruction method is proposed for obtaining, from at least one two-dimensional, 2D, color image of a human face with a visible toothed portion, a single reconstructed 3D surface of the toothed portion and of the facial portion excluding the toothed portion of the face, said method comprising: - segmenting the 2D image into a first part corresponding to the toothed portion of the face only by masking in the 2D image the facial portion outside the toothed portion of the face, on the one hand, and a second part corresponding only to the facial portion, outside said toothed portion, of the face by masking in the 2D image the toothed portion of the face, on the other hand; - enhancing the first part of the 2D image in order to modify photometric characteristics of said first part; - the generation of a 3D surface of the face reconstructed from the first enhanced part of the 2D image, on the one hand, and from the second part of said 2D image, on the other hand, said 3D surface of the face being adapted for the simulation of a projected dental treatment to be applied to the toothed portion of the face, by substituting for the area of the 3D surface corresponding to said toothed portion another 3D surface corresponding to said toothed portion after said projected dental treatment.
[0013] The embodiments use the enhanced 2D image(s) with respect to the toothed portion, in order to produce a 3D facial reconstruction, with toothed portion, of the subject's face. It is the enhancement of the toothed portion of the patient's facial image that makes possible the 3D reconstruction not only of the facial portion (excluding the toothed portion) but also of the toothed portion itself, from a single 2D image of the face with this toothed portion visible.
[0014] A first mode of implementation provides that the facial reconstruction is decoupled from that of the toothed portion. In this first mode of implementation in fact, the generation of the 3D surface of the face can include: - implementing a first deep learning algorithm adapted to produce a 2D depth map representing a 3D reconstruction of the toothed portion of the face based on the first part of the enhanced 2D image; - the implementation of a second deep learning algorithm for facial reconstruction, adapted to produce a textured 3D reconstruction of the facial portion excluding the toothed portion of the face on the basis of the second part of the 2D image; and, - an algorithm for merging the 3D reconstruction of the toothed portion and the textured 3D reconstruction of the facial part of the 2D image of the 2D image, to obtain the 3D surface of the face with its toothed portion;
[0015] In this first mode of implementation, the second deep learning algorithm can be based on a 3DMM (3D Morphable Model) type method adapted to deform a generic 3D surface so as to approximate the 2D image on the photometric plane.
[0016] Optionally, the first deep learning algorithm may be adapted to predict a depth map for the dentate portion of the face from training data by masking a depth map associated with the 2D image with the same mask as a mask used on the 2D image to obtain the first portion of the 2D image corresponding to the dentate portion of the face, and the depth map for the dentate portion of the face may be converted into a 3D reconstruction that is merged with the 3D reconstruction of the non-dentate portion of the face to produce the 3D surface of the face.
[0017] The second algorithm may be further adapted to produce the relative 3D position of the camera having captured the face as presented in the 2D image as well as an estimate of the 2D area of said 2D image in which the toothed portion of the face is located, so that a consolidated 3D surface of the face may be obtained from a plurality of 2D images of the face taken by a camera according to different respective viewing angles and for each of which the steps of the method are repeated to obtain respective reconstructed 3D surfaces, said reconstructed 3D surfaces then being combined using the relative 3D position of the camera having captured the face as presented in each 2D image as well as the estimate of the 2D area of said 2D image in which the toothed portion of the face is located.
[0018] Conversely, a second mode of implementation provides that the 3D reconstruction of the facial part excluding the toothed part and that of the toothed portion are carried out by a single algorithm. In this implementation, the generation of the 3D surface of the face may comprise the implementation of a third deep learning algorithm, different from the first and second deep learning algorithms and adapted to globally produce a 3D reconstruction of the toothed portion and the facial portion excluding the toothed portion from the second part of the 2D image to which is added the first part of said enhanced 2D image with mutual registration of said second part of the 2D image and of said first part of said enhanced 2D image.
[0019] Modes of implementation, taken individually or in combination, further provide that: - the third deep learning algorithm may be based on a 3DMM type method adapted to deform a generic 3D surface so as to photometrically approximate the second part of the 2D image to which the first part of said enhanced 2D image is added; - modifying the photometric characteristics of the first 2D part of the image may include increasing the sharpness and / or increasing the contrast of said first part of the 2D image; - the enhancement of the toothed portion of the 2D image can be achieved using a series of purely photometric filters; - the 2D enhancement processing includes the extraction of the blue channel, a high-pass contrast enhancement filtering applied to the extracted blue channel, as well as a local histogram equalization filtering, for example of the CLAHE type, applied to the filtered blue channel; - the high-pass contrast enhancement filtering applied to the blue channel may include a sharpening algorithm, for example consisting of partially subtracting from said blue channel a blurred version of itself; - alternatively, the first part of the enhanced 2D image can be produced from the original 2D image as an intermediate output of a deep learning network for semantic segmentation, having a higher contrast than the original 2D image, and selected according to a determined quantitative criterion; - a contrast metric may be associated with the output of the convolution kernel of each of the convolution layers of the deep learning semantic segmentation network, and the selected intermediate output of the deep learning semantic segmentation network may be the output exhibiting maximum contrast with respect to the metrics associated with the respective intermediate outputs of said deep learning semantic segmentation network.
[0020] In a second aspect, the invention also relates to a device having means adapted to carry out all the steps of the method according to the first aspect above.
[0021] A third aspect of the invention relates to a computer program product comprising one or more sequences of instructions stored on a data carrier. machine-readable memory comprising a processor, said sequences of instructions being adapted to carry out all the steps of the method according to the first aspect of the invention when the program is read from the memory medium and executed by the processor.
[0022] In a fourth and final aspect, the invention also relates to a method for simulating the aesthetic result of a dental treatment planned for a human subject, for example an aesthetic, orthodontic or prosthetic treatment, from at least one two-dimensional, 2D, color image of the subject's face with a visible toothed portion, said method comprising: - the three-dimensional, 3D reconstruction, from the 2D image, of the face with the toothed portion, to obtain a single three-dimensional, 3D, reconstructed surface of the toothed portion and of the facial portion excluding the toothed portion of the face by the method according to the first aspect; - substituting the area of the 3D surface corresponding to the toothed portion of the face with another 3D surface corresponding to said toothed portion after said projected treatment; and, - display of the 3D surface of the face with the toothed portion after the planned dental treatment.
[0023] Modes of implementation, taken individually or in combination, further provide that: - the method comprises implementing an algorithm applied to a 3D reconstruction of the subject's total dental arch, said algorithm being adapted to realign the dental arch on the toothed portion of the 3D surface of the face as obtained by the method according to the first aspect, and to replace the toothed portion within said 3D surface of the face with a corresponding part of said 3D reconstruction of the subject's dental arch, i.e. with the part of the subject's dental arch which is visible in the 2D image; - the dental arch can undergo digital processing, either automatic or manual, before re-alignment on the toothed portion of the 3D surface of the face, in order to simulate within said 3D surface of the face the aesthetic result of the planned treatment; - 3D reconstruction of the subject's dental arch can be obtained with an intraoral camera; and / or - the planned treatment may include at least one of the following aesthetic, orthodontic or prosthetic treatments: a change in the shade of the teeth, a realignment of the teeth, the application of veneers to the teeth, the fitting of orthodontic material (for example, braces) or prosthetic material (for example, a crown, a bridge, an inlay-core, an inlay-onlay). Brief description of the drawings
[0024] Other characteristics and advantages of the invention will become apparent from reading the description which follows. This description is purely illustrative and should be read in conjunction with the appended drawings in which: [Fig.l] [Fig.l] is a functional diagram illustrating the segmentation, according to the method of the first aspect of the invention, of a 2D color image of a human face with a visible toothed portion, into a first part corresponding to the toothed portion of the face only and a second part corresponding to the facial portion, excluding said toothed portion, of the face; [Fig.2] [Fig.2] is a step diagram of a first embodiment of the method for obtaining a 3D reconstruction from the 2D image of [Fig.l], in which the 3D reconstruction is carried out separately for each of the first and second parts of the 2D image, after enhancement of the first part and before merging at the 3D level of the 3D reconstructions thus obtained; [Fig.3] [Fig.3] is a step diagram of a first embodiment of the method for obtaining a 3D reconstruction from the 2D image of [Fig.l], in which the 3D reconstruction is carried out together for the first and second parts of the 2D image, after enhancement of the first part and merging of the two parts at the 2D level; [Fig.4] [Fig.4] is a functional diagram illustrating a first method of enhancing the toothed portion of the face of the 2D image, using a processing which implements a series of photometric filters; [Fig.5] [Fig.5] is a functional diagram illustrating a second method of enhancing the dentate portion of the face of the 2D image, exploiting advances in artificial intelligence by using an intermediate output of a deep learning network; and, [Fig.6] [Fig.6] is a functional diagram illustrating an exemplary implementation of the simulation method according to the fourth aspect of the invention, in which the projected treatment is tooth whitening. Description of the embodiments
[0025] In the following description of embodiments and in the Figures of the accompanying drawings, the same or similar elements bear the same numerical references in the drawings.
[0026] The invention takes advantage of deep learning architectures such as deep neural networks and convolutional neural networks (or neural networks) or convolutional neural networks or CNN (from the English "Convolutional Neural Networks") to carry out three-dimensional reconstructions. from one (or more) 2D image(s) of a human face which includes a visible toothed portion, acquired by an acquisition device comprising a single 2D image sensor.
[0027] Before beginning the description of detailed embodiments, it appears useful to specify the definition of certain expressions or certain terms which will be used therein. Unless otherwise provided, these definitions apply notwithstanding other definitions which the person skilled in the art may find in certain works of specialized literature.
[0028] An “image”, or “view”, or even “scan”, consists of a set of points of the real three-dimensional scene. For a 2D image acquired by an image acquisition device, or imaging device (for example a CCD sensor or a CMOS sensor), the points concerned are the points of the real scene projected into the plane of the focal length of the 2D sensor used to acquire the 2D image, and are defined by the pixels of the 2D image. For a reconstructed 3D surface (also called “3D reconstruction”), this term designates the product or result of the 3D reconstruction processing, the points concerned being a 3D point cloud obtained by a transformation of a “depth map” (see definition given below), or by triangulation in the case of stereoscopy, or by 3D deformation of a generic 3D model in the case of a 3DMM type method (see definition given below).Such a point cloud defines a skeleton of the three-dimensional scene. And a 3D mesh of this point cloud, for example a mesh of triangulated 3D points, can define an envelope of it.
[0029] A "monocular" image acquisition device is a device having only a single image sensor and capable of acquiring images of a three-dimensional scene from only a single viewing angle at a given device position.
[0030] "Registration" consists of determining the spatial relationship between two representations (2D image or 3D surface) of the same object so as to make the representations of the same physical point superimpose.
[0031] “Pose calculation” is the estimation of the position and orientation of the imaged scene relative to the imager (image sensor). This is one of the fundamental problems in computer vision, often called “Perspective-n-Points” (PnP). This problem consists of estimating the pose (2-tuple [Rj, ; tj] formed from the rotation matrix Rj and the translation vector tj) of the camera relative to an object in the scene, which amounts to finding the pose that reduces the reprojection error between a point in space and its 2D correspondent in the image. A more recent approach, called ePNP (from the English “Efficient Perspective-n-Point”), assumes that the camera is calibrated, and takes the decision to overcome calibration problems in nor malizing the 2D points by multiplying them by the inverse of the intrinsic matrix. This approach adds to this the fact of parameterizing the pose of the camera by passing through 4 control points, ensuring that the estimated transformation is rigid. Proceeding in this way makes the calculation times shorter.
[0032] By "enhancement" of the toothed portion is meant a 2D level processing specific to the toothed portion aimed at improving the photometric characteristics of said toothed portion. In embodiments, this processing specific to the toothed portion may comprise the application of a sequence of image processing filters. In other embodiments, it comprises the use of an intermediate output of a learning network.
[0033] A "sharpening" algorithm is an image processing algorithm aimed at increasing the sharpness of the image.
[0034] The acronym "3DMM" (from the English "3D Morphable Model") designates a method of generating 3D poses via a morphable (i.e. modifiable) 3D model. This method is particularly suited to processing information from the face of a human being (skin, wrinkles, illumination, relief, etc.). The 3DMM method consists of placing a 3D face (mask) on the 2D image, and modifying it to make it correspond with a face on the 2D image. The information corresponding to the modified mask is then extracted, which will make it possible to create the 3D representation of the face from the 2D image.
[0035] A "depth map" associated with a 2D image is a form of 2D representation of the reconstructed 3D information, corresponding to the portion of the 3D scene reprojected into the 2D image. In practice, it is a set of values, coded in the form of levels (or shades) of gray, respectively associated with each pixel pt of the 2D image: the greater the distance between the point of the three-dimensional scene and the plane of the 2D image, the darker the pixel.
[0036] A "convolutional neural network" or "convolutional neural network" or CNN (from the English "Convolutional Neural Networks"), is a type of acyclic artificial neural network ("feed-forward" in English), consisting of a multi-layer stack of perceptrons, the purpose of which is to pre-process small amounts of information. A CNN consists of two types of artificial neurons, arranged in "strata" or "layers" which successively process the information: - processing neurons, which process a limited portion of the image (called the “receptive field”) through a convolution function; and, - the neurons for pooling (total or partial) outputs, called “pooling” neurons pooling” (which means “grouping” or “pooling” in English), which allow information to be compressed by reducing the size of the intermediate image (often by subsampling). The outputs of a processing layer are used to reconstruct an intermediate image, which serves as the basis for the next layer. Non-linear and ad hoc corrective processing can be applied between each layer to improve the accuracy of the result. CNNs are currently widely used in image recognition.
[0037] With reference to [Fig.l], the embodiments of the method of the invention comprise the segmentation of the two-dimensional (2D) image 21 of the face of a human subject, here a young woman, into a first part 22, on the one hand, and into a second part 22, on the other hand. The first part 22 corresponds only to the toothed portion 1 of the face, which is visible in the image 21. It is obtained by masking and blacking out, in the image 21, the facial portion 4 excluding the toothed portion 1 of the face. The second part 23 corresponds only to the facial portion 4, excluding the toothed portion 1, of the face. It is obtained by masking and blacking out in the 2D image of said toothed portion 1 of the face.
[0038] The toothed portion 1 is shown in [Fig.l] in detail 10 of image 21, which corresponds to the area of the subject's mouth, which area is also identified by the same reference 10 in part 22 and in part 23 of image 21. Those skilled in the art will appreciate that the toothed portion excludes the lips and the gums, to truly include only the portion visible in image 21, where appropriate, of the upper arch and / or the lower arch of the subject's dentition. This toothed portion has, compared to the rest of the face, a high specularity and a particular texture which make 3D reconstruction difficult with conventional 3D facial reconstruction techniques.
[0039] This segmentation of the 2D image into two parts makes it possible to implement image processing specific to the toothed portion 1 which is the sole object of the first part 22, in order to compensate for the poor photometric properties of said toothed portion 1 compared to the other portions of the face. The image processing is adapted to enhance these properties, in particular the contrast. Such processing is referred to as "enhancement". It is only applied to the toothed portion 1, i.e., only to the part 22 of the image 22 of the face. The toothed portion after enhancement and the facial part excluding the toothed portion are then merged, i.e. recombined to ultimately give the 3D reconstruction of the two-dimensional image 21 of the face with the toothed portion.
[0040] Essentially two embodiments are proposed, depending on whether the above fusion is performed at the 2D level, i.e. before a 3D reconstruction applied to the recomposed image, or that the fusion is carried out at the 3D level, that is, after 3D reconstructions applied to each of the two parts of the image respectively. These two modes of implementation will now be described with regard to the step diagrams of [Fig.2] and [Fig.3], respectively.
[0041] Referring first to the step diagram of [Fig.2], the method begins, at step 201, by acquiring at least one image (i.e., a 2D view) of a subject's face that includes a visible toothed portion. This is the case, in particular, when the subject smiles. A smile is the result of a natural expression of emotion, which can also be commanded by the subject. In general, smiling uncovers all or part of the subject's upper dental arch, and generally also the subject's lower arch, due to the opening of the mouth and the stretching of the lips that smiling causes. The image 21 of the subject smiling can be taken by the subject himself, or by another person using, for example, the on-board camera of a portable device of the subject such as his mobile phone, or by any other similar imaging device, for example a camera, a webcam, etc.In embodiments, step 201 comprises taking a plurality of images of the patient's face such as image 21, taken from different respective viewing angles. These embodiments, which will be discussed later, make it possible to improve the accuracy of the 3D reconstruction of the subject's face.
[0042] In step 202, the segmentation of the image 21 is carried out into a first part 22 and a second part 24. As previously explained above with reference to [Fig.l], the first part 22 corresponds to the toothed portion 1 of the face only. And the second part 24 corresponds only to the facial portion 4, excluding said toothed portion 1, of the face. This segmentation step 202 can be carried out by digital processing applied to the data of the image 21, via an algorithm 51 which implements the detection of external limits of the toothed portion 1 of the face using a deep learning network for detecting characteristic points on a face. This makes it possible to generate a mask for each of said first and second parts 22 and 24, respectively, of the image 21. The effect of these masks is as follows: - the first part 22 of the image 21 is obtained from said image 21 by masking, that is to say by putting in black the facial portion 4 outside the toothed portion of the face; and, - the second part 24 of the image 21 is obtained from said image 21 by putting the toothed portion 4 of the face in black. In fact, what are called parts 22 and 24 of image 21 are 2D images each corresponding to said image 21 but in which part of the pixels are replaced by black pixels.
[0043] This technique is known per se and its implementation is within the reach of the skilled person. profession, which is why it will not be described in more detail in this description. It will simply be noted that the deep learning network of algorithm 51 is, in particular, adapted to exclude the lips and gums from the first part 22, so that it only includes the toothed part 1 itself, whose specularity and texture are very different from those of organic tissues, whether soft or hard, such as the skin, lips or mucous membranes of the mouth. An example of such a deep learning network is described in the article Bulat et al. "How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks)”, ICCV, 2017. The algorithm described finds characteristic points distributed along the lips. By isolating the part of the image within these points, the toothed part is isolated.
[0044] In step 203, the implementation of a facial reconstruction algorithm is carried out, which can also be implemented in the form of a deep learning network 42. This CNN is adapted to predict a textured 3D reconstruction 34 of the facial portion 4 excluding the toothed portion 1 of the face. This reconstruction is obtained on the basis of the second part 24 of the image 21.
[0045] In embodiments, the algorithm 42 is based for example on the concept of 3DMM (from the English “3D Morphable Model”), according to which the 3D surface corresponding to the three-dimensional reconstruction of any face can be obtained by deformation of an average face, the deformation being parameterized by a vector comprising a number KJace of real values.
[0046] More particularly, the deep learning network 42 has been trained for this purpose to be able to predict, given a 2D image provided as input, the set of KJace parameters which deforms the average 3D face model so that it resembles as much as possible, photometrically, the face of the 2D image provided as input. In other words, the algorithm implemented by the deep learning network 42 implements a 3DMM type method adapted to deform a generic 3D surface so as to approximate the 2D image photometrically. To approximate the 2D image as closely as possible, the algorithm can be based on a photometric proximity metric between the deformed 3D model and the initial 2D image, in connection with an optimization process based on this metric.The learning of this network 42 is done from 2D images of faces whose 3D surface is also known by a spatially precise means (for example a structured light facial scanner).
[0047] Various examples of such algorithms are known to those skilled in the art. One may cite, for example, the algorithm described in the article by Deng et al. “Accurate 3Dface reconstruction with weakly supervised learning: from single image to image set”, IEEE Computer Vision and Pattern Recognition Workshop (CVPRW) on Analysis and Modeling of Faces and Gestures (AMFG), 2019.
[0048] It will be noted that, in addition to the three-dimensional surface 34 of the face (excluding the toothed portion), the learning network 42 is also adapted to predict an illumination model (represented by 9 parameters) and a pose (represented by 6 parameters), which make it possible to estimate the relative 3D position of the camera having taken the face as presented on the 2D image provided as input. This pose estimation can be advantageously used in the case of using the method with several 2D images as input, which will be explained later.
[0049] The limitation of using this type of deep learning algorithm is that it cannot predict a plausible reconstruction of the toothed portion 1, due to the photometric character (very specular, little textured) of the latter. This is why the invention proposes to circumvent this problem, by enhancing the toothed portion 1 of the 2D images in order to make it usable on the photometric level to carry out a satisfactory three-dimensional reconstruction.
[0050] Indeed, step 204 comprises the application of a digital processing 54 to the data of the first part 22 of the image 21, which corresponds to the toothed portion 1 of the subject's face. This processing 54 comprises an enhancement of the first part 22 of the image 21 in order to modify photometric characteristics of this first part. Essentially, this enhancement aims to improve the contrast of the image 22. The processing 54 therefore makes it possible to generate an enhanced version 23 of the image 22 corresponding to the toothed portion of the face. Two modes of implementing the enhancement will be described below, with reference to [Fig.4] and [Fig.5], respectively.
[0051] In step 205, the implementation of another deep learning algorithm 41 is carried out, adapted to produce a depth map (in the 2D domain) of the toothed portion 1 of the face on the basis of the enhanced image 23 corresponding to the first part 22 of the two-dimensional image 21. In one embodiment, the deep learning algorithm 41 is adapted to predict a depth map for the toothed portion of the face from training data, by masking a depth map associated with the image 21 with the same mask as a mask used on the image 21 to obtain, in step 202, the first part 22 of the image 21 corresponding to the toothed portion 1 of the face. This depth map for the toothed portion 1 of the face is then converted into a 3D reconstruction.
[0052] In implementations, the deep learning algorithm 41 may implement a particular example of a CNN, which is in fact a FCN (Fully Convolutional Network) inspired by the article by J. Long, E. Shelhamer and T. Darrell, "Fully convolutional networks for semantic segmentation", IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, 2015, pp. 3431-3440. Such a deep learning network is specifically trained to produce a map of depth of the toothed portion 1. It takes as input 2D images of which the toothed portion 1 is isolated as described above in connection with step 202 (the rest of the image being masked and put in black) then enhanced by the processing 54 as explained above in connection with step 204. At the output, the deep learning network predicts the expected depth map on the toothed portion 1, generated from the network's training data by masking the global depth map with the same mask used on the 2D image in the enhancement step 204.
[0053] Step 206 then comprises the implementation of an algorithm 56 for merging the three-dimensional reconstruction 23 of the toothed portion 1 and the textured three-dimensional reconstruction 34 of the facial portion 4 of the face represented by the two-dimensional image 21, to obtain the three-dimensional reconstruction 35 of the complete face, with its toothed portion 1. In other words, in step 206, the three-dimensional reconstruction 33 corresponding to the depth map produced by the algorithm 41 for the toothed portion 1 of the face, is merged with the three-dimensional reconstruction 34 obtained by the algorithm 42 for the facial portion excluding the toothed portion of the face, in order to produce the three-dimensional surface 35 of the complete face.
[0054] The fusion algorithm 56 can again implement a deep learning network.
[0055] To train this network, it is necessary to build a database, with n-tuples of data acquired for different people, and which associate, for each 2D image of a person, the 3D surface of their face as well as the toothed portion.
[0056] More particularly, in order to constitute a triplet of training data, the 2D image of each person can be acquired by any commercially available device (camera, mobile phone, digital tablet, etc.). For the 3D reconstruction of the facial portion 4 of the face excluding the toothed part, it is possible to use a 3D facial reconstruction scan by structured light. Furthermore, for the toothed portion 1 which is not or which would be poorly imaged by this type of device, a 3D reconstruction can be obtained by a real color intraoral scanner (for example a WoW™ scanner available from the company BIOTECH DENTAL), thus producing a complete, textured and precise 3D dental arch. It will be noted that such a scanner can restore the texture of the teeth by amalgamating the colors of the 2D images (coded by RGB coding, for example) used for the 3D reconstruction.It is then easy to retexture the 3D model using not the raw 2D images, but images enhanced by algorithm 54 of step 204 of the method. The 3D model then presents a much more contrasted surface and is better suited to subsequent image processing algorithms based on photometry, which can ultimately be implemented within the framework of the use made of the reconstructions. facial expressions which are obtained using the method of the invention, for example for simulating the aesthetic effect of a planned dental treatment.
[0057] In the training data triplets, the 3D reconstruction of the part of the image corresponding to the toothed portion 1 of the face, enhanced or textured in RGB depending on the use to be made of it, is manually realigned on the 3D reconstruction of the facial part 4 of the face, in order to produce a single 3D reconstruction comprising the facial portion 4 and the toothed portion 1 of the face. Finally, the relative pose of the 2D image with respect to the 3D reconstruction can be calculated semi-automatically, by choosing 3D points of interest on the 3D surface as well as their corresponding point on the 2D image. Thanks to these pairs, a relative pose algorithm, for example ePNP, makes it possible to find the pose. By this method, training data are obtained per triplet {2D image; 3D reconstruction; pose}.These training data can easily be converted into other triplets {2D image; depth map; pose}, with the depth map being preferred in some implementations.
[0058] Thanks to the deep learning network 56 trained as just explained, the 3D surface of the face generated in step 206 of the method from the enhanced version 23 of the first part 22 of the 2D image, on the one hand, and from the second part 24 of said 2D image, on the other hand, is a 3D reconstruction of good quality including for the toothed portion 1 of the face. This 3D reconstruction is therefore well suited for the simulation of a projected treatment to be applied to the toothed portion of the face, by substituting for the area of the 3D surface corresponding to said toothed portion another 3D surface corresponding to said toothed portion as it would be after said projected treatment.
[0059] In summary, the enhanced image 23 corresponding to the part 22 of the input image 21 which corresponds to the toothed portion 1 of the face, is used by the deep learning algorithm 41 in order to produce a three-dimensional reconstruction 33 of the toothed portion 1 of the face in the image 21. In parallel, the deep learning algorithm 42, which is for example based on a 3DMM method, generates a three-dimensional reconstruction 34 of the facial portion 4 alone. Such an algorithm, for example, is advantageously adapted to, in addition, produce the relative 3D position of the camera having taken the face as presented in the 2D image, as well as an estimate of the 2D zone of said 2D image in which the toothed portion of the face is located.This can be used, in certain implementations of the method, to obtain at step 206 a consolidated 3D surface of the face from a plurality of 2D images of the face such as image 21, taken by a camera from different respective viewing angles. Each of these images is subjected to the 3D reconstruction method according to steps 202 to 205 of [Fig.2]. In other words, the implementation of the . method of [Fig.2] can be repeated to obtain respective reconstructed 3D surfaces. These reconstructed 3D surfaces can then be combined, in step 206, using the relative 3D position of the camera having taken the face as presented on each 2D image as well as the estimation of the 2D area of said 2D image in which the toothed portion of the face is located. The consolidated 3D surface of the face which is obtained by this type of implementation from a plurality of 2D images of the subject's face is a more accurate 3D reconstruction of the face and teeth than that obtained from a single 2D image of said face.
[0060] In a second embodiment, which will now be described with reference to [Fig. 3], the generation of the 3D surface of the face comprises the implementation of another deep learning algorithm 43 capable of predicting a 3D reconstruction from a 2D image, which differs from the deep learning algorithms 41 and 42 of the implementation mode of [Fig. 2]. This other algorithm is adapted to globally produce a 3D reconstruction of the toothed portion 1 and of the facial portion 4 excluding the toothed portion, from the second part 24 of the 2D image to which is added the first part 22 of said enhanced 2D image, with mutual registration of said second part of the 2D image and of said first part of said enhanced 2D image. This third algorithm 43 can be derived from algorithm 42 used in step 203 of the implementation mode illustrated by [Fig.2].
[0061] The first step 301 and the second step 302 of the embodiment according to [Fig. 3] are identical, the first step 201 and the second step 202, respectively, of the embodiment according to [Fig. 2]. Furthermore, the third step 303 of the embodiment of [Fig. 3] corresponds to step 304 of the embodiment of [Fig. 2]. Thus, the first step 301 corresponds to taking a 2D image of a patient's face with a visible toothed portion 1. The second step 302 is the step of segmenting the acquired 2D image, into a first part 22 corresponding to the toothed portion alone, and a second part 24 corresponding to the facial part 4 except the toothed portion 1. And the third step 303 comprises the enhancement processing of the part 22 of the image corresponding to the toothed portion 1, which makes it possible to produce an enhanced version 23 of said image 22. These steps 301, 302 and 304 are therefore not described again in detail here.
[0062] The sequence of steps in implementing the method according to [Fig.3] however differs from the implementation according to [Fig.2].
[0063] In step 304, in fact, the enhanced image 23 which corresponds to the image 22 of the toothed portion alone on which a specific treatment has been applied to enhance its photometric characteristics, is reinjected into the original 2D image 21. More particularly, this result can be obtained by merging the enhanced image 23 and the part 24 of the original 2D image 21 corresponding to the facial part 4 except the toothed part 1 of the face, by a fusion algorithm 52. The result of this fusion is a refused two-dimensional image 25, in which the toothed part 1 is enhanced. In other words, the image 25 produced by the fusion algorithm 52 is still a 2D image, like the original image 21, but it differs in that the toothed part 1 of the face is enhanced therein.
[0064] Then, in step 305, the facial reconstruction and that of the toothed portion are carried out by the common implementation of a three-dimensional reconstruction algorithm 43, applied to the refused two-dimensional image 25 in which the toothed portion 1 is enhanced. This algorithm can be derived from the algorithm 42 used in step 303 of the implementation of the method according to [Fig.2], but on the condition of adding the image of a toothed portion with enhanced texture in the training data. Following the training process described above with regard to the algorithm 42 of [Fig.2], it is possible under this condition to use the same training database to train the network 43 to predict, from 2D images of faces with enhanced toothed portion, the total 3D surface including the textured toothed portion.In other words, the algorithm implemented by the deep learning network 43 implements a 3DMM type method applied to the refused 2D image 25, and which is adapted to deform a generic 3D surface so as to approximate the 2D image on the photometric plane. To approximate the 2D image as closely as possible, the algorithm 43 can be based on a photometric proximity metric between the deformed 3D model and the initial 2D image, in connection with an optimization process based on this metric.
[0065] Referring for example to the principle of the scientific article by Deng et al. already cited above, the total reconstructions (of an image with facial part and with toothed part) exhibiting an enhanced texture on the teeth, are recalibrated with each other. A restricted parameterization is then put in place on this recalibrated data in order to best account for inter-individual deformations. At the output of this process, the deformations are parameterized by a number K_total of deformation parameters which is greater than the number KJace of parameters of the algorithm 42 of [Fig.2], accounting for both the face and the teeth. Once the deep learning network 43 is thus trained, it is capable of reconstructing, for any 2D face image including or not a toothed portion, the corresponding 3D surface.
[0066] As the person skilled in the art will have understood, the modification of the photometric characteristics of the first part 22 of the original 2D image 21 which is generated in the enhancement step 204 of [Fig.2] as in step 303 of [Fig.3], comprises the increase in the sharpness and / or the increase in the contrast of said first part 22 of the 2D image.
[0067] According to a first example of implementation, illustrated by the functional diagram of [Fig.4], the enhancement of the toothed portion of the 2D image can be achieved using a series of purely photometric filters.
[0068] More particularly, the 2D level enhancement treatment 54 which is applied to the toothed portion 1 comprises: - in step 401, the extraction of the blue channel from the color image coded in RGB format; - a high-pass contrast enhancement filter 402 applied to the blue channel extracted in step 401; as well as - in step 403, a CLAHE type local histogram equalization filtering applied to the filtered blue channel which is obtained by step 402.
[0069] Regarding step 401, those skilled in the art will appreciate that the blue channel is, spectrally, the one that contains the most contrast on dental tissue.
[0070] Further, in one example, the contrast enhancement high-pass filtering applied in step 402 to the blue channel may include a sharpening algorithm such as a sharpening algorithm applied to the blue channel. Such an algorithm involves partially subtracting a blurred version of itself from the blue channel, thereby enhancing high spatial frequency detail.
[0071] Finally, the local histogram equalization filtering of step 403 may for example be of the CLAHE type, as described in the chapter by Karel Zuiderveld “Contrast Limited Adaptive Histogram Equalization”, in the book Graphics Gems IV, P. Heckbert editions, Cambridge, MA. (Academy Press, New York), August 1994, pp. 474-485.
[0072] According to a second exemplary implementation, illustrated by the functional diagram of [Fig. 5], the enhanced version 23 of the first part 22 of the original two-dimensional image 21 corresponding to the toothed portion of the image 21 can be obtained from said original 2D image 21, as an intermediate output of a deep learning network 50 for depth map prediction, having a higher contrast than the original 2D image according to a determined quantitative criterion. Those skilled in the art will appreciate that it is perfectly accepted that the deep learning network 50, applied to an unenhanced image, can only produce depth maps as output which are not usable as such, but that this does not prevent its intermediate outputs from being used as those of a contrast enhancer in accordance with implementations of the invention, notwithstanding the fact that its outputs are not usable and are actually not used.
[0073] The deep learning architecture 50 is for example a convolutional neural network (CNN) which can have a completely conventional structure. This type of CNN is available in libraries known to those skilled in the art which are freely accessible. As input, the two-dimensional image 21 is provided in the form of a matrix of pixels. Color is encoded by a third dimension, of depth equal to 3, to represent the fundamental colors [Red, Green, Blue].
[0074] The CNN in [Fig.5] is in fact an FCN (from the English "Fully Convolutional Network") inspired by the scientific article already mentioned above, by J. Long, et al., "Fully convolutional networks for semantic segmentation", IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, 2015, pp. 3431-3440. This FCN has two very distinct parts, according to an encoding / decoding architecture.
[0075] The first part of the encoding FCN is the convolutional part itself. It comprises the “convolutional processing layer” 51, which has a succession of filters, or “convolution kernels”, applied in layers. The convolutional processing layer 51 functions as a feature extractor of the 2D images admitted as input to the CNN. In the example, the input image 21 is passed through the succession of convolution kernels, creating each time a new image called a convolution map. Each convolution kernel has two convolution layers 511 and 512, and a layer 513 for reducing the resolution of the image by a pooling operation also called a local maximum operation (“max pooling” in English).
[0076] The output of the convolutional part 51 is then provided as input to a final convolution layer 520 capturing the entire visual field of action of the previous layer, and thus mimicking a fully connected layer.
[0077] Finally, a final deconvolution layer 530 produces as output a depth map 22'. As already stated, this type of CNN is unfortunately not suitable for the 3D reconstruction of the toothed part 1 in the image 22, due to the high specularity and the low texture of the teeth. This is why the depth map 22' generated by this network 50 cannot be used for the intended application.
[0078] On the other hand, each convolution kernel of the convolutional processing layer 51 of the network 50 is adapted to extract determined photometric characteristics from the 2D image admitted as input to the CNN. In other words, each kernel generates a convolution map in the form of a new image constituting a version of the input image which is enhanced from the point of view of said characteristics.
[0079] Thus, the enhanced image 23 corresponding to the enhanced version of the image 21 at the input of the deep learning network 50 can be extracted as a determined intermediate output of said network 50, having a higher contrast than the original 2D image according to a determined quantitative criterion. This intermediate output can be selected from the outputs of the convolution kernels by a selection engine 52, on the basis of the values of a contrast metric which are respectively associated with the output of each of the convolution kernels of each convolution layer of the network 50. For example, the selected intermediate output of the network 50 may be the output of the kernel of the convolutional processing layer 51 of said network which exhibits maximum contrast with respect to the metrics associated with the respective intermediate outputs of the network, i.e. the outputs of the respective kernels of the layer 51. The image delivered by this intermediate output has a higher contrast than the original 2D image 21 provided as input to the CNN.
[0080] The invention described above makes it possible to perform facial reconstruction with a toothed portion based on any single 2D image, or on any series of 2D images. In the latter case, several images are taken from different viewing angles, and a final multi-view stereoscopic reconstruction procedure is conducted to produce a more accurate 3D reconstruction of the face and teeth.
[0081] The method finds very varied applications, notably in the simulation of dental treatments having aesthetic implications.
[0082] Thus, for example, the functional diagram of [Fig.6] illustrates an example of a method for simulating the aesthetic result of an aesthetic, orthodontic or prosthetic dental treatment, which is projected for a human subject, i.e., a patient from at least one two-dimensional, 2D, color image of the subject's face with a visible toothed portion. In this example, the treatment envisaged is an aesthetic treatment consisting of tooth whitening.
[0083] The method comprises: - the three-dimensional reconstruction 60, from an original 2D image 71 of the patient's face with a visible toothed portion (or from a plurality of such images), to obtain a single reconstructed three-dimensional surface 73 of the toothed portion 1 and of the facial portion 4 excluding the toothed portion of the face by the method as described above; - obtaining 61 a three-dimensional reconstruction 75 of the patient's dental arch, at least of the portion 1 of said dental arch concerned by the planned treatment and which is visible in the original 2D image 71; - the substitution 63 for the area of the three-dimensional surface 73 corresponding to the toothed portion 1 of the face in the original 2D image 71, of another three-dimensional surface 77 corresponding to the toothed portion 2 as it would be after said projected processing; and, - the display of the three-dimensional surface 73 of the face with the toothed portion 2 as it would be after the projected treatment.
[0084] In one example, the three-dimensional reconstruction 75 of the patient's dental arch 1 that is obtained in step 61 may be a 3D reconstruction of the arch complete image of the patient. This 3D reconstruction can, for example, be reconstructed by a 3D intraoral scanner (IOS) 72. Alternatively, the three-dimensional reconstruction 75 of the patient's dental arch 1 can be obtained by cone beam volumetric imaging (or CBCT, for "Cone Beam Computed Tomography"). CBCT is a computed tomography technique for producing a digital radiograph, located between the dental panoramic and the scanner.
[0085] In a step 62, a dental practitioner (such as a dental surgeon or an orthodontist, for example) develops a dental treatment project 74. Consequently, the dental arch 1 undergoes automatic or manual digital processing which generates a simulation 2 of said dental arch after treatment. Then, the treated dental arch 2 (here we can speak of the whitened dental arch), is realigned on the toothed portion of the three-dimensional surface 73 of the patient's face, in order to simulate within said 3D surface the aesthetic result of the projected treatment 74.
[0086] Indeed, in step 63, the three-dimensional surface 77 of the toothed portion 2 as it would appear after the planned dental treatment, here a tooth whitening, is realigned on the toothed portion of the three-dimensional reconstruction 73 of the patient's face, thanks to a realignment algorithm 76. Thus, it replaces within the three-dimensional reconstruction 73 of the patient's face, the toothed portion 1 which is visible in the original 2D image 71. In other words, the realignment algorithm 76 which is applied to a three-dimensional reconstruction 77 of the dental arch is adapted to realign the whitened dental arch 2 on the toothed portion of the three-dimensional surface 73 of the patient's face as obtained by the method according to the first aspect of the invention.This allows the patient to appreciate the relevance of the planned dental treatment 74, from the aesthetic point of view, on the basis of an overall view 73 of his face with the toothed part as it would be after said planned dental treatment.
[0087] The display of the three-dimensional surface 73 of the face with the toothed portion 2 as it would be after the projected treatment may be a 3D display, for example in 3D software of the Meshlab™ type (which is free software for processing 3D meshes), in CAD software (standing for (“Computer Aided Design”). It may also be the display of a 2D image, or a display on virtual reality glasses, or on augmented reality glasses. These examples are not limiting.
[0088] The example of teeth whitening is not limiting of the dental treatments which can be simulated using the method as described above with regard to [Fig.6]. In addition, several treatments can be simulated simultaneously. Thus, the planned treatment can include at least one of the following aesthetic, orthodontic or prosthetic treatments: a change in the shade of the teeth, a re alignment of teeth, application of veneers to teeth, fitting of orthodontic material (braces) or prosthetic material (crown, bridge, inlay core, inlay onlay), etc.
[0089] More generally, the present invention has been described and illustrated in the present detailed description and in the figures of the attached drawings, in possible embodiments. The present invention is not limited, however, to the embodiments presented. Other variants and embodiments can be deduced and implemented by the person skilled in the art upon reading the present description and the attached drawings.
[0090] In the claims, the term "comprise" or "comprise" does not exclude other elements or other steps. A single processor or several other units may be used to implement the invention. The different features presented and / or claimed may be advantageously combined. Their presence in the description or in different dependent claims does not exclude this possibility. The reference signs should not be understood as limiting the scope of the invention.
Claims
Claims
1. Method of three-dimensional, 3D, reconstruction for obtaining, from at least one two-dimensional, 2D, color image of a human face with a visible toothed portion, a single reconstructed 3D surface of the toothed portion and of the facial portion excluding the toothed portion of the face, said method comprising: - segmenting the 2D image into a first part corresponding to the toothed portion of the face only by masking in the 2D image the facial portion outside the toothed portion of the face, on the one hand, and a second part corresponding only to the facial portion, outside said toothed portion, of the face by masking in the 2D image the toothed portion of the face, on the other hand; - enhancing the first part of the 2D image in order to modify photometric characteristics of said first part; - generating a 3D surface (35) of the face reconstructed from the first enhanced part (23) of the 2D image, on the one hand, and from the second part (24) of said 2D image, on the other hand, said 3D surface of the face being adapted for the simulation of a projected treatment (74) to be applied to the toothed portion (1) of the face, by substituting for the area of the 3D surface (73) corresponding to said toothed portion (1) another 3D surface (77) corresponding to said toothed portion (2) after said projected dental treatment, said generation of the 3D surface of the face comprising: - implementing a first deep learning algorithm (41) adapted to produce a 2D depth map representing a 3D reconstruction (33) of the toothed portion of the face on the basis of the first part (22) of the enhanced 2D image; - implementing a second deep learning algorithm (42) for facial reconstruction, adapted to produce a textured 3D reconstruction (34) of the facial portion excluding the toothed portion of the face on the basis of the second part (24) of the 2D image; and, - an algorithm (56) for merging the 3D reconstruction of the toothed portion and the textured 3D reconstruction of the facial part of the 2D image of the 2D image, to obtain the 3D surface (35) of the face with its toothed portion.
2. The method of claim 1, wherein the first deep learning algorithm (2) is adapted to predict a map of depth for the toothed portion of the face from training data by masking a depth map associated with the 2D image with the same mask as a mask used on the 2D image to obtain the first part of the 2D image corresponding to the toothed portion of the face, and wherein the depth map for the toothed portion of the face is converted into a 3D reconstruction which is fused with the 3D reconstruction of the non-toothed facial portion of the face to produce the 3D surface of the face.
3. Method according to claim 1 or claim 2, wherein the second deep learning algorithm (42) is based on a 3D pose generation method type method via a morphable 3D model adapted to deform a generic 3D surface so as to approximate the 2D image on the photometric plane.
4. The method of claim 3, wherein the second algorithm (42) is further adapted to produce the relative 3D position of the camera having captured the face as presented in the 2D image as well as an estimate of the 2D area of said 2D image in which the toothed portion of the face is located, and wherein a consolidated 3D surface of the face is obtained from a plurality of 2D images of the face taken by a camera from respective different viewing angles and for each of which the steps of the method according to any one of claims 1 to 4 are repeated to obtain respective reconstructed 3D surfaces, said reconstructed 3D surfaces being combined using the relative 3D position of the camera having captured the face as presented in each 2D image as well as the estimate of the 2D area of said 2D image in which the toothed portion of the face is located.
5. Method according to claim 1, in which the generation of the 3D surface of the face comprises the implementation of a third deep learning algorithm (43), different from the first and second deep learning algorithms and adapted to globally produce a 3D reconstruction of the toothed portion and of the facial portion excluding the toothed portion from the second part of the 2D image to which is added (304) the first part of said enhanced 2D image with mutual registration of said second part of the 2D image and of said first part of said enhanced 2D image.
6. The method of claim 5, wherein the third deep learning algorithm (43) is based on a method type method 3D pose generation via a morphable 3D model adapted to deform a generic 3D surface so as to photometrically approximate the second part of the 2D image to which the first part of said enhanced 2D image is added.
7. A method according to any one of claims 1 to 6, wherein modifying the photometric characteristics (54) of the first 2D portion of the image comprises increasing the sharpness and / or increasing the contrast of said first portion of the 2D image.
8. A method according to claim 7, wherein the enhancement of the toothed portion of the 2D image is carried out using a series of purely photometric filters (401,402,403).
9. Method according to claim 8, in which the 2D enhancement processing comprises the extraction of the blue channel (401), a high-pass contrast enhancement filtering applied to the extracted blue channel (402), as well as a local histogram equalization filtering, for example of the CLAHE type (403), applied to the filtered blue channel.
10. The method of claim 9, wherein the contrast enhancement high-pass filtering (402) applied to the blue channel comprises a sharpening algorithm, for example, partially subtracting from said blue channel a blurred version of itself.
11. The method of claim 7, wherein the first portion of the enhanced 2D image (23) is produced from the original 2D image as an intermediate output of a semantic segmentation deep learning network (50), having a higher contrast than the original 2D image, selected (52) according to a determined quantitative criterion.
12. The method of claim 11, wherein a contrast metric is associated with the output of the convolution kernel of each of the convolution layers (51) of the semantic segmentation deep learning network (50), and wherein the selected intermediate output of the semantic segmentation deep learning network is the output exhibiting maximum contrast with respect to the metrics associated with the respective intermediate outputs of said semantic segmentation deep learning network.
13. A device having means configured to implement all the steps of a method according to any one of claims 1 to 12.
14. Computer program product comprising one or more sequences of instructions stored on a memory medium readable by a machine comprising a processor, said sequences of instructions being adapted to carry out all the steps of the method according to any one of claims 1 to 12 when the program is read from the memory medium and executed by the processor.
15. Method for simulating the aesthetic result of a planned dental treatment for a human subject, from at least one two-dimensional, 2D, color image of the subject's face with a visible toothed portion (1), said method comprising: - the three-dimensional, 3D, reconstruction from the 2D image of the face with the toothed portion, to obtain a single 3D surface (73) of the toothed portion and of the facial portion excluding the toothed portion of the face reconstructed by the method according to any one of claims 1 to 12; - the substitution (76) for the area of the 3D surface corresponding to the toothed portion (1) of the face of another 3D surface (77) corresponding to said toothed portion (2) after said projected treatment; and, - the display of the 3D surface of the face with the toothed portion after the projected dental treatment.
16. Method according to claim 15 comprising the implementation of an algorithm (76) applied to a 3D reconstruction (77) of the dental arch of the subject, said algorithm being adapted to re-align the dental arch on the toothed portion of the 3D surface (73) of the face as obtained by the method according to any one of claims 1 to 12, and to replace the toothed portion (1) within said 3D surface of the face by a corresponding part (2) of said 3D reconstruction of the dental arch of the subject.
17. Method according to claim 16, in which the dental arch (75) undergoes digital processing before registration on the toothed portion of the 3D surface (73) of the face, in order to simulate within said 3D surface of the face the aesthetic result of the projected treatment.
18. A method according to claim 16 or claim 17, wherein the 3D reconstruction (75) of the subject's dental arch is obtained with an intraoral camera.
19. A method according to any one of claims 15 to 18, wherein the planned treatment comprises at least one of the following aesthetic, orthodontic or prosthetic treatments: a change in the shade of the teeth, a realignment of the teeth, an apposition of veneers on the teeth, an installation of orthodontic material such as braces, or an installation of prosthetic material such as a crown, a “bridge”, an “inlay core” or an “inlay onlay”.