Device and method for determining size information relating to a face
The method addresses the challenge of locating individuals and detecting gaze orientation in a three-dimensional space by adjusting a reference face model based on morphological typing and using key points to calculate the face's location and orientation relative to the camera.
Patent Information
- Application Number
- FR2023013304
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for detecting people in images or videos cannot accurately locate individuals in a three-dimensional space relative to the camera and simultaneously detect the orientation of their gaze in a three-dimensional frame of reference.
A method that involves detecting faces in images, adjusting a reference face model based on morphological typing, locating key points on the detected face, and obtaining the location and orientation of the face relative to the camera in a three-dimensional measurement system, using a focal length and pixel coordinates of key points.
Enables accurate localization of individuals in a three-dimensional space and detection of gaze orientation, simplifying calculations and reducing computing time by minimizing the same error function.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Device and method for determining information on quantities relating to a face Technical field
[0001] The present invention relates to the field of computer vision and more specifically to the processing of images captured by conventional type cameras. Prior art
[0002] Many techniques for detecting people in images or videos exist. Indeed, more and more applications, such as surveillance and security applications, are using these methods to identify people. Several of these approaches are based on artificial intelligence and in particular on "deep learning". One of the limitations of existing methods is that they allow people to be located in the image frame of reference (for example, a two-dimensional space whose distances are expressed in pixels), and not to be located relative to the camera, in a three-dimensional space, whose distances are expressed in a spatial coordinate system, such as, for example, a metric system.
[0003] There are also methods for refining face detection, by proposing to detect the orientation of the gaze in a two-dimensional frame of reference of the image. However, the known methods do not allow both locating people relative to the camera in a three-dimensional frame of reference equipped with a length measurement system, and detecting the orientation of the person's gaze in this same measurement system.
[0004] Thus, there is a need to locate people in a three-dimensional frame of reference, and optionally to detect the orientation of their gaze in this frame of reference. Statement of the invention
[0005] The invention proposes to overcome at least one drawback of the prior art by proposing a method for determining information on quantities relating to a face captured by at least one camera comprising - a detection of at least one face in at least one image captured by the at least one camera - an adjustment of a modeling of at least one reference face, on which a plurality of key points are located, from an estimation of a morphological typing of said at least one detected face; - a location, on said at least one detected face, of corresponding key points to at least a first subset of said key points of said reference face modeling and, - obtaining a location relative to the camera, in a length measurement system, of said at least one detected face relative to said at least one camera, said location being obtained from a focal length of said camera, coordinates in pixels of at least a second subset of said key points located on said at least one detected face, coordinates in said length measurement system of the key points of said first subset located on said adjusted reference face model.
[0006] According to at least one embodiment, the method comprises: - obtaining an orientation of said at least one detected face from the focal length of said camera, the pixel coordinates of at least said second subset of said key points on said at least one detected face, the coordinates in said length measurement system of the key points of said first subset located on said adjusted reference face model.
[0007] According to at least one embodiment, obtaining a location and obtaining an orientation are obtained by minimizing the same error function.
[0008] According to at least one embodiment, the adjustment of a modeling of at least one reference face comprises a selection of a modeling of a reference face corresponding to said morphological typing.
[0009] According to at least one embodiment, the adjustment of a modeling of at least one reference face comprises: - a selection of a reference face from a plurality of reference faces, the selection taking into account at least one characteristic of said morphological typing, - an adjustment of the selected reference face based on one or more characteristics of said morphological typing not taken into account when selecting the reference face.
[0010] According to at least one embodiment, the adjustment of a modeling of at least one reference face comprises an adjustment of the coordinates of the key points of said modeling of the reference face as a function of said morphological typing.
[0011] According to at least one embodiment, said distance and said orientation are obtained by minimizing an error function from the RANSAC method.
[0012] According to at least one embodiment, the adjustment of a model of the reference face comprises: - obtaining the coordinates of said key points on said adjusted reference face model by multiplying the coordinates of said key points on the reference model by a coefficient relating to an estimated age of said at least one detected face and by a coefficient relating to an estimated gender of said at least one detected face.
[0013] According to at least one embodiment, said first subset comprises at least 5 key points.
[0014] According to at least one embodiment, said key points of said at least one detected face are located on salient points of said at least one detected face, at least some of which are located on one or the other or several of the eyes, the mouth, the nose, or the contours of said at least one detected face.
[0015] According to at least one embodiment, the detection of a face comprises a determination of a bounding box around the detected face.
[0016] According to at least one embodiment, the length measuring system is a three-dimensional system whose origin of the reference frame is centered on a position of the camera.
[0017] The characteristics presented in isolation in the present application in connection with certain embodiments of the method of the present application can be combined with each other according to other embodiments of the present method.
[0018] The present invention also relates to a computer program comprising instructions for executing the steps of the method according to the invention, according to any one of its embodiments, when said program is executed by a computer.
[0019] The present invention also relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for executing the steps of the method according to the invention, according to any one of its embodiments.
[0020] The present invention also relates to a device for determining information on quantities relating to a face captured by at least one camera, the device comprising one or more processors configured together or separately to - detect at least one face in at least one image captured by said at least one camera, - adjusting a model of at least one reference face, on which a plurality of key points are located, from an estimation of a morphological typing of said at least one detected face - locating, on said at least one detected face, key points corresponding to at least a first subset of said key points of said modeling of the reference face, - obtaining a location, in a length measurement system, of said at least one detected face relative to said at least one camera, said location being obtained from a focal length of said camera, pixel coordinates of at least at least a second subset of said keypoints located on said at least one detected face, coordinates in said length measurement system of the keypoints of said first subset located on said adjusted reference face model.
[0021] Other characteristics and advantages of the present invention will emerge from the description given below, with reference to the accompanying drawings which illustrate an exemplary embodiment thereof without any limiting character. Brief description of the drawings
[0022] [Fig-1] [Fig.l] represents a method according to certain embodiments of the present invention,
[0023] [Fig.2] [Fig.2] represents an example of adaptation of a reference model of a face to the gender and age of the person,
[0024] [Fig.3] [Fig.3] represents an example of the localization of key points on the face of the captured person,
[0025] [Fig.4] [Fig.4] shows a device according to certain embodiments of the present invention. Description of the embodiments
[0026] The present invention mainly concerns the determination of magnitude information relating to a face. It can in particular allow the localization of people in a three-dimensional reference frame, for example in the form of a vector of x, y and z coordinates expressed in the metric system of units (for example in meters); Thus, it is a localization in the real world. Such a three-dimensional reference frame differs from an image reference frame in which the dimensions can for example be expressed in pixels. Such a three-dimensional reference frame can for example have as its origin the position of the camera.
[0027] The present application relates to may also allow in certain embodiments to also detect the direction of the gaze of the person, the direction also being determined in the three-dimensional reference frame. By reference frame, we can understand, in other words, a reference frame in a length measurement system making it possible to determine a distance between the camera and the detected person. In the same way, the orientation of the face of the detected person is determined in a length measurement system such as a three-dimensional reference frame in which the camera is located and more particularly according to certain embodiments, a reference frame in which the camera is located at the origin of the three axes of the reference frame, therefore with coordinates the vector (0,0,0).
[0028] During a step E1, one or more faces are detected in an image captured by a camera. The detected face is also called the target face in the remainder of the description. Description. Several people, therefore faces, can be detected in the same image or in the same video, representing a scene acquired by one or more cameras. Depending on the context of the present disclosure, this may be, for example, an indoor or outdoor scene. This scene may include one or more moving or static people.
[0029] For example, in the case of a panography, consisting of reconstituting an image representation of a real scene by assembling a plurality of shots, taken for example by a plurality of cameras, it is possible to reconstitute a virtual shooting camera with which a focal length and a viewing angle are associated. The steps of the method can therefore be applied from this reconstituted image and the virtual camera, the center of the three-dimensional reference then being the position of the virtual camera.
[0030] According to certain embodiments, the detection of a face may comprise a step of tracing (or determining or delimiting) an area around the face. This area may for example be a rectangle, known in English as a “bounding box” or encompassing box in French.
[0031] Thus, at the input of step E1, one or more images or videos are received and at the output one or more images or videos are obtained (or alternatively one or more portions of at least one of these still images or videos) on which detected faces or certain detected faces are identified (or marked, or located or identified, or determined).
[0032] According to certain embodiments, images can be selected based on criteria such as: -the image includes at least N faces, N being greater than or equal to 1, -the image includes at least one face whose size, in pixels, is greater than a first value (for example a first value determined previously, for example during a configuration). The closer the faces are to the camera, the larger their size - the image includes a type of face that we wish to detect, man, woman, child, - criteria relating to facial recognition: we can enter a certain number of visual or morphological criteria or a robot portrait of an individual and only select images based on these criteria or meeting one or more of these criteria.
[0033] According to certain embodiments, it is possible to select from the images selected as indicated above, and comprising one or more faces, one or more faces according to criteria such as those defined above and to position a bounding box, for example, only on the face(s) selected in the selected images.
[0034] We can therefore, for example, determine a zone per face or a zone only for certain faces, by configuring step E1.
[0035] According to certain embodiments, if several faces are detected, the following steps E2 to E5 of the method can be implemented simultaneously for each of the detected faces or iteratively, face by face. At least some of the detected faces may have been selected, following their detection, according to criteria as mentioned above and the following steps can be applied only to these selected faces. It is thus possible to detect the distance of at least some of the faces (for example of each of the faces) relative to the camera and their orientation relative to the camera, or a frame of reference of the camera.
[0036] According to certain embodiments, the face detection of step E1 can be carried out using methods based on a deep neural network. Such methods are for example methods based on networks of the convolutional neural network (CNN), Transformer, or support vector machines (SVM) type trained on labeled databases.
[0037] The following steps E2 and E3 can be carried out simultaneously or sequentially depending on the embodiments. These steps can be carried out from the contents of the determined or selected bounding box(es).
[0038] The still or video image(s), or portions of still or video images, in which the faces are identified or possibly selected, following their identification, are analyzed, step E2, so as to estimate a morphological typing (or morphological type) of this face or of the person to whom this face belongs. By morphological typing, we can understand for example the shape or size of the face, the endomorphic, mesomorphic and ectomorphic traits, the age, the gender of the person.
[0039] Thus, depending on the morphological typing criteria that one wishes to estimate, step E2 may comprise one or more estimates. This or these estimates may for example determine one or more coefficients representative of the morphological typing. This or these coefficients obtained may be combined or used during the subsequent step E4.
[0040] Thus, for example, one may wish to determine the age and / or gender of the person or face, during this step.
[0041] For example, according to a first embodiment, the still or video image(s), or portions of still or video images, in which the faces are identified or possibly selected following their identification, can be analyzed, step E21, so as to estimate or determine an age of this face or of the person to whom this face belongs.
[0042] At the end of step E21, an integer representative of the estimated age of the person. Alternatively, a range of values can be obtained in which the age of the person whose face is being identified lies.
[0043] According to certain embodiments, the detection of the age of step E21 can be carried out using methods based on a deep neural network. Such methods are for example methods based on networks of the convolutional neural network (CNN) type, or transform or making it possible to solve a regression problem. Such neural networks can receive as input for example the bounding box determined during step E1, or the selected face, and give as output an estimate of the age of the individual or an age range (for example an integer relating to an age range rather than an age). According to certain variants, the output of the neural network can also be a probability density and the age associated with the greatest probability can be retained as the estimated age.
[0044] For example, according to a second embodiment, the still or video image(s), or portions of still or video images, in which the faces are identified or possibly selected following their identification, can be analyzed, step E22, so as to estimate or determine a gender of this face or of the person to whom this face belongs. At the end of step E22, a variable representative of the gender of the person to whom the face belongs, whether male or female, can be obtained. For example, a “0” if it is a man, a “1” if it is a woman (or vice versa) and a “2” if the face is androgynous (non-gendered).
[0045] For example, according to another embodiment, during step E22, the gender of the person whose face has been detected is estimated or determined. At the end of step E22, a variable representative of the gender of the person to whom the face belongs can be obtained, whether male, female, or androgynous. For example, a “0” if it is a man, a “1” if it is a woman (or vice versa) and a “2” if the face is androgynous (non-gendered).
[0046] According to certain embodiments, the gender detection of step E22 can be carried out using methods based on a deep neural network. Such methods are for example methods based on networks of the convolutional neural network (CNN) type or making it possible to solve a regression problem. Such neural networks can receive as input for example the bounding box determined during step E1, or the selected face, and give as output an estimate of the gender of the individual. According to certain variants, the output of the neural network can also be a probability density and the gender associated with the greatest probability can be retained as the estimated age.
[0047] According to certain embodiments, not shown in [Fig.l], step E2 could include other determinations, such as the shape of the face, long, thin, narrow, wide, etc., and thus provide varied morphological typings.
[0048] At the output of step E2, we obtain a morphological typing Typ.
[0049] During step E3, key points of a reference model (or face) (as defined later with respect to step E4) are determined or located on the detected face(s). The key points are characterized by their coordinates (u, v) in pixels in the image. Here, key points located on a reference model are compared with key points on the detected or selected face(s). In other words, this involves positioning key points of a reference model on the same areas on the detected face. For example, the key points, or part of the key points, located on the eyes of the reference model are placed on the eyes of the detected face. It is thus possible to match a key point i of the reference model with a point j of the detected face.
[0050] According to certain embodiments, the number of key points is between 5 and 100.
[0051] According to certain embodiments, the minimum number of keypoints is 5. The minimum number of keypoints may for example be located in a range between 5 and 15 according to the embodiments.
[0052] According to certain embodiments, the optimal number of keypoints is 68. The optimal number of keypoints may for example be located in a range between 55 and 75 according to the embodiments.
[0053] According to certain embodiments, one or more key points, preferably 5 key points, are distributed or located: - on the nose and / or, - on the mouth and / or - on the eyes, or the eyebrows and / or, - on the contours of the face.
[0054] According to some embodiments, the keypoints may represent all keypoints of a reference face model.
[0055] According to certain embodiments, the key points may represent a subset of the key points of a reference face model. This subset preferably comprises key points located on salient areas of the face. By salient area of the face, one may in particular understand one or more areas comprising a salient element when establishing a saliency map of the face. It may thus be estimated that areas having contrasts or different colors may constitute salient areas. Such areas may for example include the nose, the mouth, the eyes, the eyebrows, the cheekbones, the contours of the face. Such areas may indeed offer the advantage of being easily detectable automatically in a face (for example by artificial intelligence techniques).
[0056] According to certain embodiments, the localization of the points of the reference model on the target face, carried out during step E3, can be carried out using methods based on a deep neural network.
[0057] According to certain embodiments, the coordinates are two-dimensional coordinates, expressed in pixels. The two-dimensional reference frame or coordinate system having its origin in the top left corner of the image.
[0058] At the end of step E3, a set of key points of the detected face is obtained, defined by their coordinates in pixels in a two-dimensional space, the axes of which are defined relative to a reference point defined in the image.
[0059] During a step E4, one or more facial reference models, MR, are adapted or modified from the morphological typing estimated during step E2. More generally, step E4 consists of obtaining an adjusted reference model, as a function of the morphological typing. In other words, step E4 consists of adjusting a modeling of at least one reference face, on which a plurality of key points, with coordinates (x;, y;, Zi), are located, from an estimation of a morphological typing of the detected face;
[0060] According to a first embodiment, a single reference model is obtained and adapted to the morphological typing as determined during step E2 to obtain an adjusted reference model, also called adjusted reference face or modeling of the adjusted reference face. This adaptation may comprise a determination of one or more coefficients representative of the typing and a modification of the coordinates of the key points on the face by applying the coefficient(s), or a combination of these coefficients, to the coordinates of the key points of the reference face.
[0061] For example, the reference model can be adapted from the age estimated during step E21 and the gender estimated during step E22. As mentioned previously, the reference face is a face on which a set of key points is positioned. The location of the key points is determined on salient areas of the image, such as, for example, one or more of: - the nose, - the mouth - the eyes, and / or the eyebrows, - the contours of the face.
[0062] The coordinates of the key points of the reference face model are expressed in a length measurement system, in a 3-dimensional reference frame. It is possible to estimate, from cohorts of faces, a reference model of a human face. On the reference model, one can for example arbitrarily place the origin of the 3-dimensional reference frame by choosing the point with coordinates (0,0,0) in the center of the face, for example on the nose. We can thus deduce the position of other points from this point, in a system of length measurement, such as the metric system, expressed for example in meters or centimeters, or an imperial system, expressed in feet.
[0063] Coefficients relating to the age determined during step E21 and to the gender determined during step E22 are applied to this reference model.
[0064] Kage and Kgender coefficients can thus be applied to the coordinates of the key points of the reference model to obtain a new reference model adapted to the age and gender of the person whose face was identified during step E1. The coefficients correspond to a weighting applied to the coordinates of the reference model to adjust the reference model. According to certain embodiments, step E4 can comprise the application of weighting coefficients to the coordinates of the key points or of a part of the key points of the reference model (or of a modeling of the reference face), to correspond to the morphological typing or to adapt a selected reference model to a morphological typing.
[0065] According to certain embodiments, the coefficient Kgenre can represent a ratio between the size of a man's face and that of a woman's.
[0066] The reference model can be a reference model corresponding to the face of a man, the face of a woman or an androgynous face (i.e. “non-gendered”), we can for example use the following values for the Kgenre coefficient: - if the reference model corresponds to a female face and the detected gender of the target face is female, then the coefficient Kgenre is equal to 1, - if the reference model corresponds to a female face and the detected gender of the target face is male, then the coefficient Kgenre is equal to 1.025, - if the reference model corresponds to a female face and the detected gender of the target face is androgynous, then the coefficient Kgenre is equal to 1.0125, - if the reference model corresponds to a male face and the detected gender of the target face is female, then the coefficient Kgenre is equal to 0.975, - if the reference model corresponds to a male face and the detected gender of the target face is male, then the coefficient Kgenre is equal to 1, - if the reference model corresponds to a male face and the detected gender of the target face is androgynous, then the coefficient Kgenre is equal to 0.9875, - if the reference model corresponds to an androgynous face, and the detected gender of the target face is female, then the Kgenre coefficient is equal to 0.9875. - if the reference model corresponds to an androgynous face, and the detected gender of the target face is male, then the Kgenre coefficient is equal to 1.0125. - if the reference model corresponds to an androgynous face, and the detected gender of the target face is androgynous, then the Kgender coefficient is equal to 1.
[0067] These values can be summarized in the table below: Kgender Reference Model Female Androgynous Male Detected Gender Female 1 0.9875 0.975 Androgynous 1.0125 1 0.9875 Male 1.025 1.0125 1
[0068] These values are values derived from physiological statistics studies and constitute a simple example of values that can be used. Of course, other values can be used depending on the embodiments. For example, the Kage coefficient can be determined according to the following formula:
[0069] , _ 55-3CC? 55
[0070] With “age” representing the age of the person whose face was detected and obtained during step E21. This formula can be obtained for example by regression of a curve expressing the evolution of the size of the head as a function of age. It is derived from charts of physiological observations of the size of the skull of an individual as a function of his age. Of course, other values can be used depending on the embodiments.
[0071] The application of the coefficients Kage and Kgender to the reference model can consist of multiplying the coordinates of the points of the reference face by the product (Kage * Kgender)•
[0072] As mentioned previously with regard to step E2, when this comprises the determination of the shape or a type of shape of the face, for example, face, thin, long, wide, narrow, a coefficient relating to this typing can be determined. When the face is narrow, a coefficient can be applied to the coordinates of the reference face, the effect of which makes it possible to bring the key points closer together, in other words to reduce the distance between the key points. In the same way, if the typology of the face corresponds to a wide face, a coefficient can be applied to the coordinates of the reference face, the effect of which makes it possible to move the key points apart, in other words to increase the distance between the key points.
[0073] In other words, step E4 may comprise the application of weighting coefficients to the coordinates of the key points positioned on a reference model in order to obtain a modeling of a reference face adjusted to the morphological typing of the detected face. The reference model may have been previously selected from a set of reference models.
[0074] According to certain embodiments, several models of reference. For example, but not limited to, one can consider a model by gender, or a model by age or age group, or a model by facial morphology, such as for example a long, wide, thin face, or a model combining at least some of these characteristics. A reference model can, for example, be predefined by morphological typing. The morphological typing obtained during step E2 can then make it possible to select one from several reference models or to match a reference model to a morphological typing. In such an embodiment, step E4 can comprise a selection of the reference model corresponding to the determined morphological typing.
[0075] According to certain embodiments, it may be envisaged to combine the two approaches previously described for the adjustment of a modeling of the reference face of step E4. In other words, step E4 may comprise: - the selection of a reference face from a plurality of reference faces, the selection taking into account at least one characteristic of a morphological typing, - the adjustment of the selected reference face based on one or more characteristics of the morphological typing not taken into account when selecting the reference face.
[0076] According to a first example, when the determination of the morphological typing, step E2, estimates the age and gender of the detected face, a reference model relating to the gender can be selected and a coefficient as mentioned above can be applied to adjust this reference model relating to the gender, to the estimated age. Conversely, in this example, a reference model relating to the age could be selected and a coefficient as mentioned above can be applied to adjust this reference model relating to the age, to the estimated gender.
[0077] According to a second example, when the determination of the morphological typing, step E2, estimates the age and the shape of the detected face, a reference model relating to the age can be selected and a coefficient as mentioned above can be applied to adjust this reference model relating to the age, to the estimated shape. Conversely, in this example, a reference model relating to the shape (wide, narrow, thin...) could be selected and a coefficient as mentioned above can be applied to adjust this reference model relating to the shape, to the estimated age.
[0078] [Fig.2] illustrates an example of adaptation of a reference model corresponding to an adult man to that of an 8-year-old boy. We thus obtain a reference model adapted to morphological typing and in this example, adapted to age and gender.
[0079] The coordinates illustrated in [Fig.2] are coordinates expressed in a unit of the metric system (for example in meters or millimeters), from an axis centered on the face.
[0080] We thus obtain at the end of step E4 an adjusted reference model, MRA, at
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087] morphological typing and in this example, adapted to the gender and age of the target face. In step E5, the coordinates of the key points of the face are determined, for example in a three-dimensional length measurement system, from - of the adjusted reference model, - the focal length of the camera, and - two-dimensional pixel coordinates, keypoints positioned on the target face. More precisely, a location and therefore a distance are obtained, in a length measurement system, of the detected face(s) relative to the at least one camera. The location is obtained from a focal length of the camera, the pixel coordinates of at least a second subset of the key points located on the detected face(s), the coordinates in the length measurement system of the key points located on the adjusted modeled reference face. It can also be noted that the direction of gaze can be obtained simultaneously, for example in the form of a vector expressing the direction of gaze, which can be expressed by the rotation of the face. For a key point of index i, we can note: - (^, Vi,) its pixel coordinates in the captured image of the target face, - (Xi, yi5 Z;) its coordinates in a length measurement system in the fitted model. The relationship between these quantities can be expressed as follows: why [ / oi vi = 0 f % 1 LO 0 1 With >n >'i2 >3 *1 ^21 ^2^ ^23 3 ,r31 r32 Z. !
[0088] - f the focal length in pixels of the camera, - w and h representing respectively the horizontal and vertical resolution in pixels of the captured image, - x, y and z are the coordinates in the length measurement system, in a frame of the camera expressing the position of the face in said frame, and the coefficients pj correspond to a rotation matrix expressing the rotation of the face relative to the axis of the camera. The present disclosure aims in particular to determine the location of the detected face in the three-dimensional reference frame, by determining the coordinates (x,y,z) making it possible to locate the face, as well as to determine the coefficients pj making it possible to obtain the direction of gaze. One of the advantages of the present invention is to help locate the face and determine the direction of gaze from the minimization of the same error function and can therefore help simplify calculations and reduce computing time necessary to obtain these quantities.
[0089] The previous equation can be solved, that is, the coordinates (x,y,z) as well that the coefficients r^ can be obtained by minimizing the error function next:
[0090]
[0091] According to certain embodiments, this function can be minimized using a RANSAC method (an acronym for “RANdom SAmple Consensus”). This method advantageously makes it possible to consider only a subset of key points of the adjusted reference model. Certain associations between (u^V;) and points of the adjusted model (Xj,yiz;) can be very imprecise and can therefore be discarded by the RANSAC method.
[0092] It can also make it possible not to take into account key points that are very specific to the face, for example expression lines, non-standard faces, presenting for example certain disproportions on certain organs, linked for example to one or more malformations of the target face. As such, it takes into account a second subset of key points, the second subset being obtained from the set of key points in which the key points i whose coordinates (u^V;) on the detected face and coordinates (xi, yi, zi) on the adjusted modeled reference face are too far apart, are discarded.
[0093] According to some embodiments, this function can be minimized using a least median method.
[0094] The coordinates (x, y, z) representative of the location of the detected face relative to the camera can therefore be obtained when minimizing this function. The coordinates are expressed in the distance measurement system of the coordinates of the reference model, and not in pixels. Thus, the coordinates (x, y, z) are also expressed in the length measurement system. The distance between the face and the camera is then obtained by calculating the Euclidean distance between the camera and the point with coordinates (x,y,z).
[0095] We can also, from the minimization of the same error function, obtain the direction of a vector expressing the direction of gaze of the target face in the form of the following vector:
[0096] look= 'Hi H2 r2i r22 .r31 r32
[0097] According to some embodiments, the length measurement system may be the metric system in which distances may be expressed in meters.
[0098] According to some embodiments, the length measurement system may be the imperial system in which distances may be expressed in feet.
[0099] Other examples of length measurement systems can be used. It can be understood that one of the goals is to measure (or more precisely estimate) real distances between the camera and the face and therefore not to determine a distance in the image frame of reference which could be expressed in pixels.
[0100] [Fig.4] represents a device configured to implement at least one of the embodiments of the invention. The device 10 comprises in particular a processor 1, a random access memory 2, a read-only memory 3 and a non-volatile memory 4. They may further comprise communication means 5, and / or a user interface 6. More generally, the device 10 has the hardware architecture of a computer.
[0101] The read-only memory 3 constitutes a recording medium in accordance with at least one embodiment of the present disclosure, readable by the processor 1 and on which is recorded a computer program PROG in accordance with at least one embodiment of the present disclosure, comprising instructions for executing steps of the method according to at least one embodiment of the present disclosure. The program PROG defines functional modules of the device.
[0102] The communication means 5 allow in particular the device 10 to exchange data with capture means, such as a camera configured to capture a scene and in particular one or more people present in this scene in accordance with the present disclosure or one of its embodiments. For this purpose, the communication means 5 may for example comprise a computer data bus capable of transmitting data. Alternatively, said data may be transmitted via a communication interface, wired or wireless, capable of implementing any suitable protocol known to those skilled in the art.
Claims
Claims
1. Method for determining quantity information relating to a face captured by at least one camera comprising - a detection (El) of at least one face in at least one image captured by said at least one camera - an adjustment (E4) of a modeling of at least one reference face (MR), on which a plurality of key points are located, from an estimation (E2) of a morphological typing (typ) of said at least one detected face; - a localization (E3), on said at least one detected face, of key points corresponding to at least a first subset of said key points of said modeling of the reference face (MR) and, - an obtaining (E5) of a location (x,y,z) relative to the camera, in a length measurement system, of said at least one detected face relative to said at least one camera, said location being obtained from a focal length (f) of said camera, coordinates (u;,Vi) in pixels of at least a second subset of said key points located on said at least one detected face, coordinates (Xj, yi5 Z;) in said length measurement system of the key points of said first subset located on said adjusted reference face modeling.;
2. Method according to claim 1 comprising - obtaining (E5) an orientation of said at least one detected face from the focal distance (f) of said camera, coordinates (u;,Vi) in pixels of at least said second subset of said key points on said at least one detected face, coordinates (x;, y;, z;) in said length measurement system of the key points of said first subset located on said adjusted reference face model.
3. Method according to one of the preceding claims in which obtaining (E5) a location and obtaining (E5) an orientation are obtained by minimizing the same error function.
4. Method according to one of the preceding claims in which the adjustment (E4) of a modeling of at least one reference face (MR) comprises a selection of a modeling of a reference face (MR) corresponding to said morphological typing.
5. Method according to one of claims 1 to 3 in which the adjustment (E4) of a modeling of at least one reference face (MR) comprises: - a selection of a reference face (MR) from a plurality of reference faces, the selection taking into account at least one characteristic of said morphological typing, - an adjustment of the reference face (MR) selected according to one or more characteristics of said morphological typing not having been taken into account during the selection of the reference face (MR).
6. Method according to one of the preceding claims in which the adjustment (E4) of a modeling of at least one reference face (MR) comprises an adjustment of the coordinates of the key points of said reference face (MR) as a function of said morphological typing (typ).
7. Method according to one of the preceding claims in which the adjustment of a model of a reference face (MR) comprises: - obtaining the coordinates of said key points on said model of the adjusted reference face by multiplying the coordinates of said key points on said reference face (MR) by a coefficient (Kage) relating to an estimated age of said at least one detected face and by a coefficient (Kgender) relating to an estimated gender of said at least one detected face.
8. Method according to one of the preceding claims in which said key points of said at least one detected face are located on salient points of said at least one detected face, at least some of which are located on one or the other or more of the eyes, the mouth, the nose, or the contours of said at least one detected face.
9. Method according to one of the preceding claims in which the length measuring system is a three-dimensional system whose origin of the reference frame is centered on a position of the camera.
10. Device for determining information on quantities relating to a face captured by at least one camera, the device comprising one or more processors configured together or separately to - detect at least one face in at least one image captured by said at least one camera, - adjust a modeling of at least one reference face (MR), on which a plurality of key points are located, from an estimation of a morphological typing of said at least one detected face - locating, on said at least one detected face, key points corresponding to at least a first subset of said key points of said modeling of the reference face (MR), - obtaining a location (x,y,z), in a length measurement system, of said at least one detected face relative to said at least one camera, said location being obtained from a focal length (f) of said camera, coordinates (u^V;) in pixels of at least a second subset of said key points located on said at least one detected face, coordinates (xi5 y;, Zj) in said length measurement system of the key points of said first subset located on said adjusted reference face model.
Citation Information
Patent Citations
Imaging apparatus, imaging apparatus control method, and computer program
US20090073304A1