A learning method for machine learning systems to detect and model objects in images, corresponding computer program products, and equipment.
The learning method for a machine learning system uses augmented reality images to enhance object and feature region detection and modeling by combining segmentation models and contour point learning, addressing inaccuracies in existing methods and improving detection and modeling performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FITTINGBOX
- Filing Date
- 2022-02-10
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for detecting and modeling objects and feature regions in images suffer from inaccuracies due to the use of feature points, leading to incomplete or inconsistent annotations, particularly when objects are obscured or hidden.
A learning method for a machine learning system that utilizes augmented reality images containing virtual elements to train a neural network, combining segmentation models and contour point learning to enhance accuracy and consistency in object and feature region detection and modeling.
The method achieves accurate segmentation and point distribution across object contours, resolving ambiguities in manual annotation and improving detection and modeling performance by leveraging augmented reality images with virtual elements.
Smart Images

Figure 0007853309000001 
Figure 0007853309000002 
Figure 0007853309000003
Abstract
Description
Technical Field
[0001] The field of the present invention is the field of image processing.
[0002] More specifically, the present invention relates to a method for detecting and modeling objects and / or feature regions (e.g., eyes, nose, etc.) detected in an image.
[0003] The present invention has, in particular, but not exclusively, numerous application fields for the virtual testing of glasses.
Background Art
[0004] In the remainder of this specification, in particular, the existing problems in the field of virtual testing of glasses that the inventors of the present application have faced are described. Of course, the present invention is not limited to this specific field of application, but is interested in the detection and modeling of any type of object represented in an image and / or any type of feature region (i.e., a part of the image of interest) of such an image.
[0005] To detect the objects and / or feature regions under consideration, it is known from the prior art to use feature points of some objects and / or some feature regions. For example, the corners of the eyes have conventionally been used as feature points that enable the detection of an individual's eyes in an image. Other feature points such as the nose or the corners of the mouth may also be considered for face detection. In general, the quality of face detection depends on the number and position of the feature points used. Such techniques are described, in particular, in the French patent published in French Patent Invention No. 2955409 and in the international patent application published in International Publication No. 2016 / 135078 of the company that has filed the present patent application.
[0006] For manufactured objects, for example, edges or corners may be considered as feature points.
[0007] However, using such feature points can lead to a lack of accuracy when detecting and, therefore, when modeling the object and / or feature region being considered, where applicable.
[0008] Alternatively, manual image annotation may be considered in some cases to artificially generate feature points for the objects and / or feature regions under consideration. However, even in this case, a lack of accuracy in detecting the objects and / or feature regions under consideration is noted. Where applicable, such inaccuracies can cause problems when modeling the objects and / or feature regions thus detected. [Prior art documents] [Patent Documents]
[0009] [Patent Document 1] French Patent No. 2955409 Specification [Patent Document 2] International Publication No. 2016 / 135078 [Non-patent literature]
[0010] [Non-Patent Document 1] Ronneberger, Fischer & Brox, "U-Net: Convolutional Networks for Biomedical Image Segmentation," 2015. [Non-Patent Document 2] Chen, Zhu, Papandreou, Schroff, & Adam, "Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation," 2018. [Overview of the Initiative] [Problems that the invention aims to solve]
[0011] As a result, there is a need for techniques that can accurately detect and model one (or more) objects shown in an image and / or one (or more) feature regions present in the image being considered. [Means for solving the problem]
[0012] Embodiments of the present invention provide a learning method for a machine learning system for detecting and modeling at least one object represented in at least one given image and / or at least one feature region of the at least one given image. According to such a method, the machine learning system - A step of generating a plurality of augmented reality images comprising a real image representing the at least one object and / or the at least one feature region and at least one virtual element, -For each augmented reality image, with respect to at least one given virtual element of the augmented reality image, - A segmentation model of a given virtual element obtained from a given virtual element, and - A set of contour points obtained from a given virtual element, corresponding to the parameter representation of the given virtual element, or the parameter representation Steps to obtain learning information that includes, -The machine learning system learns from multiple augmented reality images and training information, and detects the at least one object and / or the at least one feature region in the at least one given image, - A segmentation model of the at least one object and / or the at least one feature region, and - A set of contour points corresponding to the parameter representation of the at least one object and / or the at least one feature region, or the parameter representation A step of providing a set of parameters that make it possible to determine the corresponding modeling information comprising To carry out.
[0013] As a result, the present invention provides a novel and ingenious solution for training a machine learning system (e.g., a conventional neural network) to detect one or more objects (e.g., one pair of glasses) and / or one or more feature regions (e.g., the contour of an eye or the contour of an iris) in a given image (e.g., a figure showing the face of a person wearing glasses) and determining the corresponding modeling information.
[0014] More specifically, the system learns from augmented reality images containing virtual elements representing one (or more) objects and / or one (or more) feature regions. As a result, segmentation (e.g., by binary masks of objects and / or feature regions simulated by one or more virtual elements) and point distribution across the entire contour are performed accurately. For example, this resolves the ambiguity inherent in manual annotation of such points. For example, the point distribution corresponds to a 2D or 3D parameterized representation of an object simulated by one or more virtual elements (e.g., a 3DMM model ("3D Morphable Model")) and / or a 2D or 3D parameterized representation of a feature region. With this method, annotations are accurate, and it is easy to return from contour points to parameterized representations.
[0015] Furthermore, using such virtual elements solves the problem of obscuration, thus avoiding incomplete annotations, such as those that may occur with the ends of eyeglass temples hidden by ears or irises hidden by eyelids. The mixture of real data (images) and virtual objects allows machine learning systems to be specialized for real-world applications. As a result, augmented reality images offer a trade-off between the realism of the image and the ease of creating images and annotations with sufficient variability.
[0016] In an embodiment, for each augmented reality image, the learning of the machine learning system comprises co-learning from a segmentation model of a given virtual element and a set of contour points corresponding to the parameter representation of the given virtual element.
[0017] As a result, a synergistic effect is obtained with respect to the learning of the machine learning system, the learning of the segmentation model, and the learning of a set of contour points, which reinforce each other. By maximizing the number of pixels of the virtual object appropriately detected by the segmentation model (i.e., by minimizing the number of pixels erroneously detected as belonging to the object), the accuracy can be improved. Further, the set of points detected in this way corresponds to the shape of a consistent object. In this example, this consistency is enhanced by the fact that the points result from a parametric model. As a result, a consistent object shape is obtained regardless of the position of the camera capturing the real image and regardless of the setup of the objects within the augmented reality image.
[0018] In an embodiment, the co-learning executes a cost function that depends on a linear combination between the cross entropy associated with the segmentation model of a given virtual element and the Euclidean distance associated with a set of contour points corresponding to the parameter representation of the given virtual element.
[0019] For example, the machine learning system comprises a branch for learning the segmentation model and a branch for learning a set of contour points. As a result, the cross entropy is associated with the branch for learning the segmentation model, and the Euclidean distance is associated with the branch for learning a set of contour points.
[0020] In some embodiments, the real image includes an example of a face. The learning information comprises visibility information indicating whether a contour point is visible or hidden by the face with respect to at least one of the set of contour points corresponding to the parameter representation of the given virtual element.
[0021] As a result, the visibility of the contour points is taken into account.
[0022] In some embodiments, the cost function further depends on the binary cross-entropy related to the visibility of the contour points.
[0023] In some embodiments, the learning information includes a parameter representation of the given virtual element.
[0024] As a result, the machine learning system can directly provide a representation of the parameters being considered.
[0025] The present invention also relates to a method for detecting and modeling at least one object represented in at least one image and / or at least one feature region of said at least one image. Such a detection and modeling method is performed by a machine learning system trained by performing the learning method described above (by any one of the embodiments described above). According to such a detection and modeling method, the machine learning system performs the detection of said at least one object and / or said at least one feature region in said at least one image and performs the determination of modeling information for said at least one object and / or said at least one feature region.
[0026] As a result, training on augmented reality images containing virtual elements representing one (or more) objects and / or one (or more) feature regions ensures consistency of modeling information when modeling one (or more) objects and / or one (or more) feature regions. Furthermore, when both one (or more) objects and one (or more) feature regions (e.g., eyes, irises, nose) are detected and modeled simultaneously, a synergistic effect is achieved, resulting in improved performance for both objects and feature regions compared to detecting and modeling only one of them.
[0027] In some of the embodiments described above, the learning of the machine learning system comprises collaborative learning from a set of contour points corresponding on the one hand to a segmentation model of a given virtual element and on the other hand to a parameter representation of a given virtual element. In some of these embodiments, the decisions made by the detection and modeling methods are - A segmentation model of the at least one object and / or the at least one feature region, and - A set of contour points corresponding to the parameter representation of the at least one object and / or the at least one feature region. This will be a joint decision.
[0028] In the embodiments described above, the collaborative learning of the machine learning system performs a cost function that depends on a linear combination between, on the one hand, the cross-entropy associated with a segmentation model of a given virtual element, and on the other hand, the Euclidean distance associated with a set of contour points corresponding to a parameterized representation of a given virtual element. In some of these embodiments, the collaborative decision performed by the detection and modeling method performs a given cost function that depends on a linear combination between, on the one hand, the cross-entropy associated with a segmentation model of the at least one object and / or the at least one feature region, and on the other hand, the Euclidean distance associated with a set of contour points corresponding to a parameterized representation of the at least one object and / or the at least one feature region.
[0029] In the embodiments described above, the machine learning system trains an augmented reality image having a real image with a face example. In some of these embodiments, the at least one image has a given face representation, and the machine learning system further determines visibility information indicating whether a given contour point is visible or hidden by a given face, with respect to at least one of a set of contour points corresponding to the parameter representation of the at least one object and / or the at least one feature region.
[0030] As a result, the machine learning system further determines the visibility of the contour points.
[0031] In the embodiments described above, the cost function performed during the collaborative learning of the machine learning system further depends on the binary cross-entropy related to the visibility of the contour points. In some of these embodiments, the decisions performed by the detection and modeling methods perform a given cost function, which further depends on the binary cross-entropy related to the visibility of the given contour points.
[0032] In some of the embodiments described above, the learning information comprises a parameter representation of a given virtual element. In some of these embodiments, the modeling information comprises a parameter representation of the at least one object and / or the at least one feature region.
[0033] In some embodiments, the at least one image comprises a plurality of images, each representing a different view of the at least one object and / or the at least one feature region. Not only detection but also determination is performed jointly for each of the plurality of images.
[0034] As a result, the performance of detecting and determining modeling information for one (or more) objects and / or one (or more) feature regions is improved.
[0035] The present invention also relates to a computer program comprising program code instructions for performing a method described above herein, according to any different embodiment of the present invention, when executed on a computer.
[0036] The present invention also relates to an apparatus for detecting and modeling at least one object shown in at least one image and / or at least one feature region of said at least one image. Such an apparatus comprises at least one processor and / or at least one dedicated computer configured to perform the steps of a learning method according to the present invention (according to any one of the different embodiments described above). As a result, the features and advantages of this apparatus are the same as those of the corresponding steps of the learning method described herein. Consequently, the features and advantages of this apparatus will not be described in further detail.
[0037] In some embodiments, the at least one processor and / or the at least one dedicated computer is further configured to perform the steps of the detection and modeling method according to the present invention (according to any one of the different embodiments described above). As a result, the features and advantages of this device are the same as those of the corresponding steps of the detection and modeling method described herein. Consequently, the features and advantages of this device will not be described in further detail.
[0038] In some embodiments, the above-described device includes the aforementioned machine learning system.
[0039] In some embodiments, the above-mentioned device is the aforementioned machine learning system.
[0040] Other objects, features, and advantages of the present invention will become clearer from reading the following description, which is provided with reference to the drawings and is merely illustrative and not limiting. [Brief explanation of the drawing]
[0041] [Figure 1]The present invention provides steps for learning a machine learning system to detect and model one (or more) objects shown in at least one image and / or one (or more) feature regions of the at least one image being considered, according to embodiments of the present invention. [Figure 2a] This provides an example of a real image with a facial feature. [Figure 2b] Figure 2a illustrates a real image and an augmented reality image with one pair of glasses. [Figure 2c] Figure 2b illustrates a segmentation model for one pair of glasses in an augmented reality image. [Figure 2d] Figure 2b illustrates a set of contour points corresponding to the parameter representation of one pair of glasses in the augmented reality image. [Figure 3] The present invention provides a step-by-step method for detecting and modeling one (or more) objects shown in at least one image and / or one (or more) feature regions of the at least one image being considered. [Figure 4a] An example image is provided showing a face and a pair of glasses. [Figure 4b] In addition to the segmented model of one pair of glasses in the image in Figure 4a, we also illustrate a set of contour points corresponding to the parameterized representation of one pair of glasses in the image in Figure 4a. [Figure 5] In addition to a set of contour points corresponding to the parameter display of the eye in the image, the iris of the eye, which is also being considered, is also illustrated. [Figure 6] An example of the structure of a device that enables the execution of several steps of the learning method shown in Figure 1 and / or the detection and modeling method shown in Figure 3, according to an embodiment of the present invention, is shown. [Modes for carrying out the invention]
[0042] The general principle of the present invention is based on the steps of using augmented reality images to train a machine learning system (e.g., a conventional neural network) to detect one or more objects (e.g., one pair of glasses) and / or one or more feature regions (e.g., the contour of the eye or the contour of the iris, the nose) in a given image (e.g., a figure showing the face of a person wearing glasses), and determining the corresponding modeling information.
[0043] More specifically, such an augmented reality image comprises a real image and at least one virtual element representing one (or more) objects and / or one (or more) feature regions under consideration.
[0044] Training convolutional neural networks requires a large amount of annotated data. The cost of acquiring and annotating this data is very high. Furthermore, annotation accuracy is not guaranteed, thereby limiting the robustness and accuracy of the inference models thus created. By using images of synthetic objects obtained from parameterized 2D or 3D models, it becomes possible to have a large amount of training data, and it also becomes possible to guarantee the positioning and visibility of 2D or 3D annotation points. These virtual objects are illuminated by realistic environment maps ("environment mappings") which may be set up or estimated from real images. Moreover, by using such virtual elements, the problem of hidden elements can be solved, and thus annotation operators can avoid obtaining incomplete or inconsistent annotations in order to arbitrarily select annotations.
[0045] Furthermore, we propose using training information that complements the augmented reality images, including not only training information with segmentation models related to the corresponding virtual elements, but also a set of contour points corresponding to the parameter representation of the virtual elements being considered.
[0046] As a result, segmentation and point distribution on contours (for example, by a binary mask of the object and / or feature region simulated by one or more virtual elements) are performed accurately, eliminating the need for annotation of the real image.
[0047] For the remainder of this application, the “machine learning system” should be understood as a system configured not only to perform training on a learning model, but also to use the model under consideration.
[0048] Referring to Figure 1, the steps of the learning method PA100 of a machine learning system (e.g., a conventional neural network) for detecting and modeling one (or more) objects represented in at least one image and / or one (or more) feature regions of the at least one image considered, according to the present invention. Examples of performing the steps of the considered method PA100 are also discussed with reference to Figures 2a, 2b, 2c, and 2d. More specifically, according to the examples in Figures 2a, 2b, 2c, and 2d, the real image 200 comprises an example of a face 220, and the virtual element 210 is a pair of glasses. Correspondingly, Figure 2c illustrates a segmented model 210ms of the single virtual glasses in Figure 2b, and Figure 2d illustrates a set of contour points 210pt corresponding to the parameterized representation of the single virtual glasses in Figure 2b. For clarity, the characteristics of Method PA100 will be illustrated below using references to elements in Figures 2a, 2b, 2c, and 2d, in a manner not limited thereto.
[0049] Returning to Figure 1, during step E110, the machine learning system obtains a plurality of augmented reality images 200ra, each comprising a real image 200 representing one or more objects and / or one or more feature regions, and at least one virtual element 210.
[0050] For example, each augmented reality image 200ra is generated thanks to a tool that specifically inserts virtual elements 210 into the real image 200. In some variant forms, the generation of augmented reality images 200ra comprises adding at least one virtual element 210 (e.g., by adding Gaussian noise, blur) and then inserting it into the real image 200. Such an addition includes, for example, illumination of the virtual object using a realistic environment map that may be set up or estimated from the real image. As a result, the realism of the virtual element 210 and / or the integration of the virtual element 210 into the real image 200 is improved. For example, such improved realism makes it easier to improve detection performance with respect to the real image, as well as learning.
[0051] For example, the augmented reality image 200ra generated in this way is stored in a database that the machine learning system accesses to obtain the augmented reality image 200ra being considered.
[0052] Returning to Figure 1, during step E120, the machine learning system obtains training information for each augmented reality image, with respect to at least one virtual element 210 of the augmented reality image 200ra being considered, comprising the following: - A segmentation model 210ms of a given virtual element 210. For example, such a segmentation model is a binary mask of one (or more) objects and / or one (or more) feature regions simulated by the given virtual element 210. - A set of contour points 210pt corresponding to the parameterized representation of a given virtual element 210. For example, the point distribution corresponds to a 2D or 3D parameterized representation of one (or more) objects and / or one (or more) feature regions simulated by the given virtual element 210. For example, such a parameterized representation (also called a parametric model) is a 3DMM model (meaning "3D Morphable Model"). For example, in relation to a 3D parameterized representation, contour points of a 3D object may be referenced in 2D by projection and parameterization of geodesic curves onto the surface of the 3D object representing the contour from the perspective of a camera capturing a real image.
[0053] As a result, segmentation and point distribution across the 210pt contour are performed accurately, and the segmentation model 210ms and a set of contour points 210pt are obtained directly from the corresponding virtual element 210, but not through post-processing of an augmented reality image with the virtual element 210 being considered. For example, this not only resolves the ambiguity inherent in manual annotation of such points, but also allows for easy return to parameter representation from the 210pt contour points.
[0054] In some embodiments, the learning information comprises a segmentation model 210ms and a parameterized representation of the contour points 210pt (instead of the coordinates of these contour points 210pt). This parameterized representation may be derived from a modeling of a virtual element 210, or it may be a specific inductive modeling of the virtual element 210. For example, the machine learning system learns the control points of one or more splines whose contour points 210pt are later found. In this case, the output of the machine learning system consists of these modeling parameters (e.g., control points) with an invariant cost function (e.g., the Euclidean distance between the control points and the ground truth). For this to be possible, the transformation that allows switching from the modeling parameters to contour points 210pt must be differentiable, so that the gradient can be backpropagated by the learning algorithm of the machine learning system.
[0055] In some embodiments, the learned information further includes an additional item: consistency between the segmented model 210ms and the contour points 210pt. This consistency is measured by the intersection of the segmented model 210ms and the surface defined by the contour points 210pt. For this purpose, a mesh is defined over this surface (for example, by the well-known Delaunay algorithm of the prior art), and the mesh is then used by a differential rendering engine that colors ("fills") this surface with uniform values. Subsequently, consistency items (e.g., cross-entropy) measuring the proximity of the segments and the pixels of the rendered surface can be defined.
[0056] For example, the training information is stored in the aforementioned database by referring to the corresponding augmented reality image 200ra.
[0057] Returning to Figure 1, during step E130, the machine learning system performs a learning phase based on multiple augmented reality images 200ra and the learning information. Such learning enables the machine learning system to generate a set of parameters (or a learned model) that allows it to detect one or more objects and / or one or more feature regions under consideration in at least one given image and determine the corresponding modeling information.
[0058] For example, during a given iteration of such learning, the input to the learning system is an augmented reality image 200ra comprising a virtual element 210. Learning also performs learning information related to the augmented reality image 200ra. For example, the learning information is a segmentation model 210ms of the virtual object 210 and contour points 210pt of the virtual object 210. Knowledge of the virtual element 210, for example, by a parametric 3D model of the virtual element 210, allows these learning information to be generated in a preferred manner, for example, by projection of the 3D model in the image with respect to the segmentation model 210ms and by sampling points of the model with respect to the contour points 210pt. The output of the learning system is the segmentation model of the virtual element and the contour points 210pt of the virtual element, or a parameterized representation of the virtual element. Learning is performed by comparing the output of the learning system with the learning information until convergence occurs. If the output of the learning system is a parameterized representation (e.g., 3DMM, spline, Bézier curve, etc.), the contour points are determined from these parameters and compared with the true contour points.
[0059] In some embodiments, for every 200 augmented reality images, the machine learning system's training involves co-learning from a segmentation model 210ms and a set of contour points 210pt. As a result, a synergistic effect is achieved with respect to the training of the machine learning system, the training of the segmentation model 210ms, and the training of the set of contour points 210pt, which reinforce each other.
[0060] For example, a machine learning system might have a branch for learning a segmentation model 210ms and another branch for learning a set of contour points 210pt. Cross-entropy is associated with the branch for learning the segmentation model, and Euclidean distance is associated with the branch for learning the set of contour points. Co-learning runs a cost function that depends on a linear combination between the cross-entropy associated with the segmentation model 210ms and the Euclidean distance associated with the set of contour points 210pt.
[0061] In some embodiments, the machine learning system is a convolutional semantic segmentation network. For example, the machine learning system is a "U-Net" type network, as described in the 2015 paper "U-Net: Convolutional Networks for Biomedical Image Segmentation" by Ronneberger, Fischer & Brox, or a "Deeplabv3+" type network, as described in the 2018 paper "Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation" by Chen, Zhu, Papandreou, Schroff, & Adam.
[0062] In the case of "U-Net," the network structure can be modified to allow for the joint learning of a segmentation model and a set of contour points. For example, the last convolutional layer of the decoder section is split into two branches (a branch for learning the segmentation model and a branch for learning the set of contour points). As a result, consistency between learning the segmentation model and learning the set of contour points is ensured. Furthermore, a pooling layer that follows a fully connected layer can reduce the dimensionality of the branch dedicated to learning the set of contour points.
[0063] In Deeplabv3+, the step of concatenating low-level and encoder characteristics is performed at four times the resolution. For example, it is at this level that the split into two branches occurs (a branch for training the segmentation model and a branch for training a set of contour points). In some implementations, convolutional layers, pooling layers with maximum pooling (or "max pooling"), and finally fully connected layers for training a set of contour points can be added.
[0064] As shown in the examples in Figures 2a, 2b, 2c, and 2d, the set of contour points 210pt in Figure 2d comprises, in detail, contour points 210pt hidden by a face 220 when a virtual pair of glasses is inserted into the real image 200 to generate an augmented reality image 200ra. As a result, by using such virtual elements, the problem of hijacking can be solved, and thus it becomes possible to avoid obtaining incomplete annotations, as can be the case in this example where the ends of the glasses temples are hidden by the ears.
[0065] As a result, in an embodiment of method PA100 in which the real image 200 comprises an example of a face, the learning information includes visibility information indicating whether at least one contour point 210pt of a set of points is visible or hidden by the face 220. Consequently, the visibility of the contour point is taken into consideration.
[0066] For example, in the above embodiment where the machine learning system is a conventional semantic segmentation network of, for example, the "Unet" type or "Deeplabv3+" type, the cost function further depends on the binary cross-entropy related to the visibility of the contour points 210pt.
[0067] In some embodiments, the learning information includes a parameter representation of a given virtual element 210, and therefore, indirectly, a parameter representation of one or more objects and / or one or more feature regions simulated by the given virtual element 210. As a result, the machine learning system can directly provide the parameter representation being considered.
[0068] In some embodiments, method PA100 comprises a learning refinement step in which the machine learning system refines a set of parameters given during the execution of step E310 from real data having annotated real images. Such annotations are performed manually or automatically (for example, by performing a face parsing algorithm).
[0069] Next, referring to Figure 3, the steps of a method for detecting and modeling one (or more) objects represented in at least one image and / or one (or more) feature regions of the at least one image, according to an embodiment of the present invention, are described.
[0070] More specifically, the detection and modeling methods of this technique are performed by the aforementioned machine learning system, which is trained by performing the learning method PA100 described above (by any one of the embodiments described above).
[0071] As a result, during step E310, the machine learning system performs the detection of one or more objects and / or one or more feature regions within at least one image (real or augmented image), and determines modeling information for one or more objects and / or one or more feature regions.
[0072] As a result, training from an augmented reality image 200ra having virtual elements 210 representing one (or more) objects and / or one (or more) feature regions under consideration guarantees consistency of the modeling information associated with modeling one (or more) objects and / or one (or more) feature regions.
[0073] Furthermore, in embodiments in which both at least one object (e.g., a pair of glasses) and at least one feature region (e.g., an eye, iris, or nose) of the at least one image are detected and modeled simultaneously, a synergistic effect is obtained, resulting in improved performance in detecting and modeling the at least one object and the at least one feature region compared to detecting and modeling only one of the at least one object and the at least one feature region.
[0074] In some embodiments, the modeling information is - A segmentation model for one (or more) objects and / or one (or more) feature regions, and - A set of contour points corresponding to the parameter representation of one (or more) objects and / or one (or more) feature regions. It is equipped with.
[0075] As a result, it is easy to return to the parameterized representation of the model obtained from the contour points.
[0076] In some embodiments described above with reference to Figure 1, the training of the machine learning system comprises co-learning from a segmentation model 210ms and a set of contour points 210pt. In some of these embodiments, step E310 comprises co-determination of a segmentation model for one or more objects and / or one or more feature regions, as well as a set of contour points corresponding to a parameter representation of one or more objects and / or one or more feature regions.
[0077] In some embodiments described above with reference to Figure 1, collaborative learning performs a cost function based on a linear combination between the cross-entropy associated with the segmented model 210ms and the Euclidean distance associated with a pair of contour points 210pt. In some of these embodiments, detection and modeling in step E310 performs the aforementioned cost function which depends on a linear combination between the cross-entropy associated with the segmented model and the Euclidean distance associated with a pair of points.
[0078] In some embodiments described above with reference to Figure 1, the training information includes visibility information indicating whether a contour point 210pt is visible or hidden by a face 220. In some of these embodiments, detection and modeling in step E310 further determines, for at least one contour point, visibility information indicating whether the contour point is visible or hidden by a face in the image analyzed during step 310. As a result, the machine learning system further determines the visibility of the contour point. In some of these embodiments, for training, a cost function is performed that depends on the loss of binary cross-entropy related to the visibility of the contour point, as performed in the corresponding embodiments described above with reference to Figure 1.
[0079] In some embodiments described above with reference to Figure 1, the learned information comprises a parameter representation of a virtual element used in an augmented reality image, and therefore indirectly, a parameter representation of one or more objects and / or one or more feature regions simulated by the virtual element under consideration. As a result, in some of these embodiments, the determination and modeling in step E310 determines the parameter representation of one or more objects and / or one or more feature regions detected in the image being analyzed during step E310.
[0080] In some embodiments, one (or more) objects and / or one (or more) feature regions are represented in multiple images, each representing a different view of the one (or more) objects and / or feature regions being considered. In this method, during step E310, the joint execution of detection and modeling in different images among the multiple images allows the modeling information of the one (or more) objects and / or feature regions being considered to be determined by an improved method.
[0081] In some embodiments, the position of the image to be analyzed by the machine learning system is normalized during step E310. For example, when the image represents a face with respect to a portion of a given subdivision of the face (e.g., the eyes), the size of the region of the subdivision under consideration is changed, for example, using facial markers ("markers"). These markers may be obtained by any known marker detection or face recognition method. As a result, step E310, in which modeling information is detected and determined, is performed for each resized region. This method facilitates not only the detection of one or more objects and / or one or more feature regions, but also the determination of the corresponding modeling information.
[0082] In some embodiments, markers (e.g., facial markers) are added across the image to be analyzed by the machine learning system during step E310 to indicate feature points (e.g., the location of the nose, the location of the temple point). For example, these markers may be obtained by a face analysis algorithm. This method facilitates not only the detection of one (or more) objects and / or one (or more) feature regions, but also the determination of corresponding modeling information.
[0083] Next, we will first discuss an example of performing the steps of the detection and modeling method with reference to Figures 4a and 4b. In this example, image 400 contains an example of face 420. Furthermore, we assume that the machine learning system is trained to detect and model a pair of glasses from an augmented reality image as described above herein with reference to, for example, Figures 2a, 2b, 2c, and 2d. As a result, the object 410 to be detected and modeled in image 400 is a pair of glasses. By performing step E310, the machine learning system determines modeling information comprising a segmented model 410ms of the glasses under consideration and a pair of contour points 410pt. In detail, the pair of contour points 410pt in Figure 4b comprises contour points 410pt that are obscured by face 420 in image 400. As a result, the method described herein can solve the problem of obscuration and thus avoid obtaining incomplete annotations, as may be the case in this example where the ends of the glasses temples are obscured by the ears.
[0084] Next, with reference to Figure 5, we discuss another example of performing the steps of the detection and modeling method. More specifically, according to this example, image 500 contains a partial example of a face. The feature region 510zc to be detected and modeled is the eye represented in image 500. Furthermore, we assume that the machine learning system is trained to detect and model an eye from an augmented reality image that has one or more virtual elements placed at eye level to model the eye. In this method, by performing step E310, the machine learning system determines modeling information about the feature region 510zc, which in this case also includes the iris, as well as a set of contour points 510pt of the eye, which are considered in detail.
[0085] Next, referring to Figure 6, we describe an apparatus 600 that enables the execution of several steps of the learning method PA100 of Figure 1 and / or the detection and modeling method of Figure 3, according to an embodiment of the present invention.
[0086] The device 600 comprises random access memory 603 (i.e., RAM memory) and a processing unit 602 equipped with, for example, one (or more) processors and controlled by a computer program stored in read-only memory 601 (e.g., ROM memory or hard disk). During initialization, the code instructions of the computer program are loaded into the active memory 603 before being executed, for example, by the processor of the processing unit 602.
[0087] Figure 6 illustrates just one specific method, among several possible methods, for making the device 600 perform some steps of the learning method PA100 in Figure 1 and / or the detection and modeling method in Figure 3 (by any one of the embodiments and / or variations described above herein with reference to Figures 1 and 3). In practice, these steps may be performed without distinction on a programmable computer (PC computer, one or more DSP processors, or one or more microcontrollers) that runs a program comprising a set of instructions, or on a dedicated computer (e.g., a set of logic gates such as one or more FPGAs or one or more ASICs, or any other hardware module).
[0088] If the device 600 is made at least partially using a reprogrammable computer, the corresponding program (i.e., a permutation of instructions) may or may not be stored on a removable storage medium (such as a CD-ROM, DVD-ROM, or flash disk), which is partially or fully readable by a computer or processor.
[0089] In some embodiments, the device 600 includes a machine learning system.
[0090] In some embodiments, the device 600 is a machine learning system.
Claims
1. A learning method for a machine learning system for detecting and modeling at least one object represented in at least one given image and / or at least one feature region of the at least one given image, The aforementioned machine learning system, The steps of generating a plurality of augmented reality images comprising a real image representing at least one object and / or at least one feature region and at least one virtual element, For each of the augmented reality images, with respect to at least one given virtual element of the augmented reality image, A segmentation model of the given virtual element obtained from the given virtual element, and A set of contour points, or the parameter representation, obtained from the given virtual element, corresponding to the parameter representation of the given virtual element. Steps to obtain learning information that includes, The machine learning system learns from the plurality of augmented reality images and the learning information, and detects the at least one object and / or the at least one feature region in the at least one given image, The segmentation model of the at least one object and / or the at least one feature region, and A set of contour points corresponding to the parameter representation of the at least one object and / or the at least one feature region, or the parameter representation A learning method characterized by performing the steps of providing a set of parameters that enable the determination of corresponding modeling information comprising a certain element.
2. The learning method according to claim 1, wherein for each augmented reality image, the learning of the machine learning system comprises co-learning from, on the one hand, the segmentation model of the given virtual element, and on the other hand, the set of contour points corresponding to the parameter representation of the given virtual element.
3. The learning method according to claim 2, wherein the collaborative learning implements a cost function that depends on a linear combination between, on the one hand, the cross-entropy associated with the segmentation model of the given virtual element and, on the other hand, the Euclidean distance associated with the set of contour points corresponding to the parameter representation of the given virtual element.
4. The learning method according to any one of claims 1 to 3, wherein the real image comprises an example of a face, and the learning information comprises visibility information indicating whether at least one of the set of contour points corresponding to the parameter representation of the given virtual element is visible or hidden by the face.
5. The learning method according to claim 4, dependent on claim 3, wherein the cost function further depends on the binary cross-entropy associated with the visibility information of the contour points.
6. A method for detecting and modeling at least one object represented in at least one image and / or at least one feature region of said at least one image, A machine learning system trained by performing the learning method according to any one of claims 1 to 5, a detection and modeling method that performs the detection of the at least one object and / or the at least one feature region in the at least one image, and performs the determination of modeling information for the at least one object and / or the at least one feature region.
7. The machine learning system is trained by performing the learning method described in claim 2 or any one of claims 3 to 5 when dependent on claim 2. The aforementioned decision is, The segmentation model of the at least one object and / or the at least one feature region, and The set of contour points corresponding to the parameter representation of the at least one object and / or the at least one feature region With joint decision-making, The detection and modeling method according to claim 6.
8. The detection and modeling method according to claim 6 or 7, wherein the machine learning system is trained by performing the learning method according to claim 4 or claim 5 when dependent on claim 4, wherein the at least one image comprises a representation of a given face, and the machine learning system further determines visibility information indicating whether a given contour point is visible or hidden by a given face, with respect to at least one given contour point of a set of contour points corresponding to the parameter representation of the at least one object and / or the at least one feature region.
9. The detection and modeling method according to any one of claims 6 to 8, wherein the at least one image comprises a plurality of images each representing a different view of the at least one object and / or the at least one feature region, and the detection and modeling are performed jointly for each of the plurality of images.
10. A computer program product comprising, when the program is executed on a computer, program code instructions for executing the method according to any one of claims 1 to 9.
11. A device comprising a machine learning system for detecting and modeling at least one object represented in at least one image and / or at least one feature region of said at least one image, The steps of generating a plurality of augmented reality images comprising a real image representing at least one object and / or at least one feature region and at least one virtual element, For each of the augmented reality images, with respect to at least one given virtual element of the augmented reality image, A segmentation model of the given virtual element obtained from the given virtual element, and A set of contour points obtained from the given virtual element, corresponding to the parameter representation of the given virtual element, or the parameter representation and Steps to obtain learning information that includes, The machine learning system learns from the plurality of augmented reality images and the learning information, and detects at least one object and / or at least one feature region in at least one given image, The segmentation model of the at least one object and / or the at least one feature region, and A set of contour points corresponding to the parameter representation of the at least one object and / or the at least one feature region, or the parameter representation A step of providing a set of parameters that make it possible to determine the corresponding model information comprising A device comprising at least one processor and / or at least one dedicated computer configured to perform the following:
Citation Information
Patent Citations
METHOD FOR INTEGRATING A VIRTUAL OBJECT INTO PHOTOGRAPHS OR VIDEO IN REAL TIME
FR2955409A1
Method and system for creating custom products
JP2016537716A
Method, device and computer program for virtual fitting of eyeglass frames
JP2020525844A
Virtual try-on system and method for eyeglasses
JP2021534516A
Process and method for real-time physically accurate and realistic-looking glasses try-on
WO2016135078A1