Computer-based methods and apparatus for determining the size and distance of head features.

By identifying the true size probability distribution of image features and using machine learning, combined with Gaussian processes, the complexity of existing technologies requiring specific equipment and multiple images is solved, achieving high-precision head feature measurement from a single image.

CN115428042BActive Publication Date: 2025-12-02CARL ZEISS VISION INTERNATIONAL GMBH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202180021425.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-16
Filing Date
2021-03-15
Publication Date
2025-12-02
Estimated Expiration
2041-03-15

AI Technical Summary

Technical Problem

Existing technologies require specific equipment or multiple images to determine the true size or distance of features in an image, and the methods are complex and lack flexibility.

Method used

By identifying multiple features in an image, utilizing the true size probability distribution and pixel size of the features, and combining machine learning and Gaussian processes, the true size of the features or their distance from the camera device can be estimated, enabling accurate measurement using a single image.

Benefits of technology

It can determine the true size of head features or the distance between the camera and the feature with high precision without the need for additional hardware and multiple images, simplifying the measurement process and improving the flexibility and accuracy of measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115428042B_ABST
    Figure CN115428042B_ABST
Patent Text Reader

Abstract

A computer-implemented method and apparatus for determining the size or distance of head features are provided. The method includes identifying multiple features in an image of a human head. Based on a probability distribution of the true size of at least one of the multiple features and the pixel size of the at least one of the multiple features, the true size of at least one target feature or the true distance between the at least one target feature and a camera device used to capture the image is estimated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application relates to a computer-implemented method and apparatus for determining the size and distance of head features based on one or more head images.

[0002] Various applications in the field of eyeglass lens fitting or eyeglass frame fitting require knowledge of the dimensions of various head features. For example, the interpupillary distance (PD) (i.e., the distance between the centers of the pupils when the eyes are looking straight ahead at an object at infinity) as defined in Section 5.29 of DIN EN ISO 13666:2013-10 may be required both for custom eyeglass frames and, in some cases, for fitting eyeglass lenses themselves to a specific person.

[0003] Many recent methods that use technologies such as virtual eyeglass frame fitting rely on one or more images taken from a person's head.

[0004] When an image of a human head is captured, the dimensions of various head features (such as interpupillary distance) can be determined from that image in pixels (image elements). However, some scaling is required to convert that pixel-level size to a real-world size (e.g., in millimeters).

[0005] In the following text, the term “real size” will be used to refer to the size of a feature in the real world (in millimeters, centimeters, etc.), rather than the size given in pixels, which can be directly obtained from the image and will be referred to as “pixel size”.

[0006] In this regard, the term "feature" refers to one or more parts of the head that can be identified in an image, and / or the relationships between these parts. For example, interpupillary distance can be considered such a feature, as well as individual parts like the nose, ears, and eyes.

[0007] There are several common methods known for achieving such scaling. For example, US 9,265,414 B2 uses an object of known size captured along with other objects in the image. For instance, the object of known size could be a plastic credit card placed against the head to be captured in the image, and then the number of pixels the credit card spans in the image (its true size) is determined. Since the true size of the credit card (e.g., in millimeters) is also known, this corresponds to scaling that converts the pixel-to-millimeter size.

[0008] Another publication that uses a reference object of known size (in this case, for virtual try-on technology) is US6,535,233 A. These methods require an object of known size, which can be inconvenient in some cases.

[0009] US 6,995,762 B1 enables the reconstruction of solid objects from objects found in two-dimensional images. This method requires knowledge of the exact physical size of the imaged object.

[0010] US 6,262,738 B1 describes a method for estimating a volume distance map from a two-dimensional image. This method uses physical rays parallel and perpendicular to the object's projection into three-dimensional space, and then reconstructs the size estimate from the resulting two-dimensional image. For this method, specific rays are required to obtain the depth image.

[0011] JP 6 392 756 B2 discloses a method for estimating the size of an object by taking multiple images while the object to be measured is rotated on a platform having a known axis of rotation and a known rate of rotation. This method requires a rotating platform.

[0012] In another method disclosed in US 2010 / 0 220 285 A1, the distance between a camera used to capture an image of a person's head and the head is measured, for example, using an ultrasonic sensor. Knowing this distance and the pixel size of a feature in the image, the actual size, for example, in millimeters, can be calculated for a given camera optics. In other words, for a specific camera with specific optics, the pixel size of a feature depends on the distance between the feature and the camera, as well as the actual size of the feature. If two of these three quantities (i.e., pixel size, actual size of the feature (e.g., in millimeters), and distance between the camera and the feature) are known, the third quantity can be calculated.

[0013] For some applications, it is also necessary to determine the distance between the feature (especially the eye) and the camera, for example, for the eccentric photorefraction measurement described in WO 2018 / 002 332 A1.

[0014] Kumar MS Shashi et al., “Face distance estimation from a monocular camera,” IEEE International Conference on Image Processing, September 15, 2013, pp. 3532-3536, XP032966333, DOI: 10.1109 / ICIP.2013.6738729; Sai Krishna Pathi et al., “A Novel Method for Estimating Distances from a Robot to Humans Using Egocentric RGB Camera,” Snesors, Vol. 19, No. 14, July 17, 2019, pp. 1-13, XP055725398, DOI: 10.3390 / s19143142; and Bianco Simone et al., “A unifying representation for pixel-precise distance.” "Estimation [A Unified Representation for Precise Pixel Distance Estimation]", Multimedia Tools and Applications, Kluwer Academic Press, Boston, USA, Vol. 78, No. 10, August 24, 2018, pp. 13767-13786, XP036807892, ISSN: 1380-7501, DOI: 10.1007 / S11042-018-6568-2, each discloses a method for estimating the distance between the camera and the face based on facial features.

[0015] As can be seen from the above description, traditional methods for determining the true size of features in an image or the distance between a feature and the camera when the image is captured have various drawbacks, such as requiring specific equipment (such as a rotating platform) or specific lighting rays, multiple images, or additional objects of known size.

[0016] Therefore, the purpose of this application is to provide methods and corresponding devices for determining the true size or distance of head features, which do not require special additional hardware or the capture of multiple images.

[0017] This objective is achieved by the method as claimed in claim 1 or 3 and the device as claimed in claim 12 or 13. The dependent claims define other embodiments, as well as computer programs and storage media or data signals carrying such computer programs.

[0018] According to the present invention, a computer-implemented method for determining the size or distance of head features is provided. The method includes:

[0019] Provide images of human heads.

[0020] Identify multiple features in the image, and

[0021] Based on the probability distribution of the true size of at least one of the plurality of features and the pixel size of the at least one of the plurality of features, the true size of at least one of the plurality of features or the true distance between the at least one of the plurality of features and the camera device used to capture the image is estimated.

[0022] As already mentioned, "true" size and "true" distance refer to sizes and distances in the real world, such as the interpupillary distance of a person, measured in millimeters. Pixel size refers to sizes that can be directly obtained from an image. For example, the pixel size corresponding to interpupillary distance is the distance between pupils in an image, measured in pixels.

[0023] "At least one target feature among the plurality of features" refers to one or more features for which the true size or true distance is to be determined. "At least one feature among the plurality of features" refers to one or more features used as the basis for the determination. These two are not mutually exclusive; that is, a feature can be either "at least one target feature among the plurality of features" or "at least one feature among the plurality of features," but different features can also be used as the basis for the determination and as target features (multiple features).

[0024] In this application, a probability distribution is a type of information, specifically a mathematical function that gives information about the probability of different true sizes of a feature occurring. It can be a function whose integral is normalized to 1 to give a mathematical probability. The probability distribution constitutes prior knowledge about the true size of the feature. Such a probability distribution can be obtained using data measured from many heads. Generally, human body dimensions, including the head (especially the face), have been studied extensively in medicine and biology, resulting in a wealth of data and published probability distributions. Examples of such publications include Patrick Caroline's "The Effect of Corneal Diameter on Soft Lens Fitting," November 17, 2016; blog www.contamac-globalinsight.com; CC Gordon et al.'s "2012 Anthropometric Survey of USArmy Personnel: Methods and Summary Statistics," 2012, Statistical (Number NATICK / TR-15 / 007); Natick Soldier Research, Development and Engineering Center, Massachusetts Army, June 2017, retrieved from http: / / www.dtic.mil / dtic / tr / fulltext / u2 / a611869.pdf; and B. Winn et al.'s "Factors Affecting Light-Adapted Pupil Size in Normal Human Subjects [Factors Affecting the Size of the Photopic Pupil in Normal Individuals], Investigative Ophthalmology & Visual Science, 1994; or NA Dodgson, “Variation and extrema of human interpupillary distance,” Stereoscopic Displays and Virtual Reality Systems XI 5291, 36-46, 2002.

[0025] In other embodiments, a body-motivated mathematical model can be used as a probability distribution function. For example, if the size distribution of a known object (e.g., ears) can be well described by a Gaussian normal distribution, then that mathematical function (in this case, a Gaussian normal distribution) can be used as a probability distribution without modeling the full probability distribution based on available data, or the distribution can be pre-measured by the provider of the computer program implementing the above-described computer implementation method, in which case the covariance between the distributions of different sizes of different features can also be measured. These covariances can then be used later, as will be further described below. In other words, if the provider of the program obtains the probability distribution itself, for example, by measuring a large number of features on the head or face, the covariance between the sizes of different features can also be determined (e.g., whether a larger head is associated with larger ears). It should be noted that when different methods are used to obtain the probability distributions described above, the obtained probability distributions will also vary depending on the data they are based on. This may lead to corresponding changes in the estimates of true size or true distance.

[0026] By using a single probability distribution of a single feature, i.e., if at least one of the multiple features is a single feature, the true size of the target feature and thus the true distance can be roughly estimated (essentially, one could say that the maximum value of the probability distribution of a single feature is the most likely true size of that feature).

[0027] Therefore, preferably, the probability distribution of the corresponding true dimensions of at least two features is used, i.e., at least one of the multiple features is one of at least two features. In this way, the estimation can be refined, thereby achieving high accuracy.

[0028] In this way, at least one of the true size or true distance can be estimated using only images and available data. In this respect, as explained in the introduction, the true size of features in an image and the distance between the features and a specific camera device with specific optics have a fixed relationship.

[0029] Identifying multiple features in an image can be done in conventional ways, such as using Dlib or OpenCV software to detect facial features (see, for example, Adrian Rosebrock's "Facial Landmarks with Dlib, OpenCV and Python" published on April 3, 2017 at www.pyimageresearch.com, G. Bradski's "The OpenCV Library," Dr. Dobb's *Journal of Software Tools*, 2000, or DEKing's "Dlib-ml: Amachine Learning Toolkit," *Journal of Machine Learning Research*), or by using other conventional face segmentation software, or by using more sophisticated facial compression representations such as machine learning. In such compressed versions, the dimensions of the measurements are not immediately apparent in terms of manually defined measurements, but can be processed through machine learning techniques.

[0030] Another approach for identifying multiple features begins with detecting human faces in an image. This can be done using any conventional face detection method, such as the method described in "Histograms of oriented gradients for human detection" by Dalal, N., and Triggs, B., in: IEEE Conference on Computer Vision and Pattern Recognition 2005 (CVPR'05), pp. 886-893, Vol. 1, doi:10:1109 / CVPR:2005:177.

[0031] In this face detection method, an image pyramid is constructed, and features are extracted from windows sliding across the image using histograms of oriented gradients. Based on these features, a linear support vector machine classifier is trained, which classifies each window as either containing a face or not.

[0032] Then, in the method discussed here, once a face is detected, an Active Appearance Model (AAM), as described in Cootes, TF, Edwards, GJ, Taylor, CJ, 2001, “Active Appearance Models” (IEEE Transactions on Pattern Analysis and Machine Intelligence 23, 681-685. doi:10:1109 / 34:927467), is applied to the image to detect so-called landmarks. Landmark detection identifies individual points on the face that can be labeled, such as the ears, the top and bottom of the ears, the lips, the eyes, and the boundaries of the iris. AAM is a generative model built on the statistical, parameterized concept of deformable objects and consists of two parts: a shape model and an appearance model. The shape model (also called a point distribution model) defines the object as an array of landmark coordinates.

[0033] During training, the shapes of all objects of interest in the training set are aligned and normalized, for example, by generalized Procrustes analysis as described in Gower, JC, 1975, “Generalized Procrustes Analysis” (Psychometrics 40, 33-51). In this application, the objects of interest are the detected faces. Principal component analysis (PCA) can then be used to compute the orthonormal basis of this set of shapes. Therefore, the objects s ​​can be defined by the shape model as:

[0034]

[0035] in, The shape represents the average shape, p is the shape parameter specific to this object, and S is the principal component of the shape.

[0036] The appearance model includes information about the image texture around the defined logo, and its training process is very similar to that of the shape model. The first step is to extract information from the training images; in this specific application, this information includes image gradient orientation. Next, all images are warped to an average shape to align the corresponding logo pixels across all images. After image alignment, principal components of the appearance can be computed using PCA. Finally, blobs around the logo are selected, and the remaining information is discarded.

[0037] Then, similar to the shape model, the facial appearance 'a' can be defined using vector representation as:

[0038]

[0039] in, The average appearance of all training images is defined, where c is the appearance parameter of the image, and A is the principal component of the appearance.

[0040] The next step in the feature detection method described above is to fit the model to the new image. This is equivalent to finding an optimal set of parameters p and c that minimizes the difference between the image texture sampled from the shape model and the appearance model. In other words, given an image I containing the object of interest and an initial guess of the object's shape s, then the current appearance... Image texture A sampled from points in s s The difference r between them can be defined as Therefore, the optimal model parameters are given by the following formula:

[0041]

[0042] This cost function can be optimized using any optimization algorithm, such as those in the gradient descent family.

[0043] Based on the obtained markers, features can then be determined. Specifically, features can be the pixel dimensions between the identified markers. For example, once the facial marker model described above is applied to an image, a set of N marker points is generated. From this set of N marker points, a set of M = N(N-1) / 2-1 unique but related features on the face with pixel size measurements can be constructed.

[0044] Preferably, two or more of the following features are identified, and the pixel size is determined:

[0045] -Pupillary distance (PD),

[0046] - The diameter of one or both irises of the eye,

[0047] - The diameter of the pupil

[0048] -Ears are long and upright.

[0049] - The width of one or two eyes,

[0050] -Head high,

[0051] - The distance between the bottom of the chin and the midpoint of the nose between the eyes is called the Menton-Sellion distance.

[0052] -Face width, which is the widest point of the entire face measured at the top of the jaw, and / or

[0053] - Skull width, which is the maximum width of the forehead above the ears.

[0054] The aforementioned features can be easily obtained using the image analysis techniques described above, and these features are well-documented, making the probability distribution available or attainable.

[0055] In some embodiments, the pixel size can be normalized to the pixel size of one of these features. For example, some embodiments aim to ultimately determine the interpupillary distance of a person based on an image of a person, and in this case, the pixel size can be normalized by the interpupillary distance measured in pixel space, so that the features used from here on are size-free and independent of the image size. The normalized pixel size can facilitate subsequent processing and can be identified as input features to be used by subsequent machine learning algorithms.

[0056] In some embodiments, features can be selected automatically or manually. As described above, for N marker points, M = N(N-1) / 2-1 features (distances between points) can be determined. For the case of N = 88, M = 3827 potential input features have been generated. To reduce this number, feature selection techniques can be applied to identify a subset containing S features from these M features.

[0057] For example, the Minimum Redundancy-Maximum Correlation (MRMR) algorithm can be used to simplify the feature set. For example, the algorithm is described in the following articles: Ding, C., Peng, H., “Minimum redundancy feature selection from microarray gene expression data”, 2003, in: Computational Systems Bioinformatics, Proceedings of the IEEE Bioinformatics Conference, CSB2003, pp. 523-528, doi:10:1109 / CSB:2003:1227396; or Peng, H., Long, F., Ding, C., “Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy”, 2005, IEEE Transactions on Pattern Analysis and Machine Intelligence, 27, 1226-1238, doi:10:1109 / TPAMI:2005:159. The MRMR algorithm selects features by identifying those that provide the most information about the target feature (and are therefore the most relevant), while discarding features that are redundant with each other (i.e., those that are most relevant to other features). Mutual information (MI), described in Cover, TM, Thomas, JA, 1991, "Elements of Information Theory" (Williams Publishing, Inc.), can be used as a relevant metric for measuring relevance and redundancy. The target feature is the feature ultimately to be determined from an image of a person, such as the interpupillary distance (PD) to be used for fitting eyeglass frames.

[0058] Exhaustive search of the entire feature combination space is computationally expensive. As described by Peng et al. above, a greedy forward search algorithm can be implemented. The initial feature set is sorted according to the MI value and interpupillary distance, and the top P features are selected, where P > S and P << M. This set containing P features is then analyzed to identify the final feature set S. If a feature maximizes the total relevance to the target feature while minimizing the total redundancy with other features already present in S, then the feature is added from P to S. In practice, at each step of increasing the feature set to size S, the feature added results in the maximum mutual information difference d. MI Characteristics. Quantity d MI Defined as

[0059]

[0060] Where S is the feature set, PD is the interpupillary distance (in mm, used as an example of target features), and I(_;_) is the mutual information operation.

[0061] A final eigenvector of a certain size can be selected to allow estimation of the covariance matrix from a relatively small amount of labeled data (e.g., S=20).

[0062] In one aspect of the invention, based on the pixel size and the probability distribution of the identified features, the probability distribution P(pix per mm|d, θ) of the number of pixels per millimeter estimated from the features in the image can be obtained according to the following equation (5):

[0063]

[0064] In equation (5), d i It is the number of pixels spanning at least one of multiple features, specifically the i-th feature (i = 1, 2, ..., N). In other words, it is the pixel size of the i-th feature. π(θ) i θ represents the true size of feature i. i The probability distribution of and / or its covariance with other measured pixel sizes. P(d i |pix per mm, θ i ) is generated in a given π(θ) i In the case of pixels per mm and θ i To measure d i The likelihood operator. To further explain, here, d i θ is the number of pixels of the corresponding size of feature i in the face plane (which is perpendicular to the camera device and is physically equidistant from it, which is at least approximately true for face images). Pixels per mm is a variable that can be considered as a current or initial estimate of how many pixels are per millimeter, and θ i π(θ) represents the magnitude of characteristic i, expressed in physical units. i The probability distribution represents prior knowledge about the size distribution of feature i before measurement, and optionally the covariance of that distribution with the pixel sizes of other measurements (e.g., as described above). θ i By using d i Multiply by pixels per mm to calculate the actual size (e.g., in millimeters).

[0065] In other words, P(d) i |pix per mm, θ iThis takes into account measurements in pixels, given π(θ) i Given prior knowledge, feature i is considered to be either large or small. π(θ) i This can be obtained based on public databases, measurement results, etc., as described above. The basic idea is to observe different values ​​of pixels per mm to determine which values ​​are reasonable, thus ensuring that the true size of all features is reasonably sized. To give a very simplified numerical example, π(θ) i θ can represent the true size of feature i. i The probability of being 2mm is 0.25, the probability of being 3mm is 0.5, and the probability of being 4mm is 0.25. If d i If there are 6 pixels, then the probability of 2 pixels per mm is estimated to be 0.5, the probability of 3 pixels per mm is 0.25, and the probability of 1.5 pixels per mm is 0.25 based on feature i. This is just a very simple numerical example. Equation (5) now combines such estimates for multiple features i.

[0066] In another aspect of the invention, P(pix per mm|d, θ) is calculated, which is the probability distribution of the number of pixels per millimeter measured using features in the image based on d and θ of the features to be measured. P(pix per mm|d, θ) can be used for any target feature in the image and can be estimated based on a mathematical combination of probability distributions in MCMC (Monte Carlo Markov Chain) probability space exploration or other embodiments. Equation (5) gives such a mathematical combination by multiplication. In particular, when multiple probability distributions are involved, this calculation is a high-dimensional calculation that requires a large amount of computing power, memory, or both. In such cases, the probability distributions can be sampled in some way, for example, using the MCMC method described above, which is a standard method for solving such problems in statistics. This method is described, for example, in W. Hastings, “Monte Carlo sampling methods using Markov chains and their applications,” Biostatistics, Vol. 57, pp. 97-109.

[0067] For a single measurement of a single feature (e.g., i = 1), P(pix per mm|d, θ) may be based solely on π(θ). i (i=1), but as the number of measurements increases (i>1), P(pix per mm|d, θ) becomes more accurate (e.g., by using the dimensions of multiple features). ∝ Indicates the proportional relationship.

[0068] Essentially, equation (5) represents that, given some prior information π(θ) i Then, we can obtain the estimate P(pix per mm|d, θ).

[0069] Another approach to determining P(pix per mm|d, θ) and ultimately the magnitude of target features such as interpupillary distance is to use machine learning techniques, including scikit-learn as described in the following articles: Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grissel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Coumapeau, D., Brucher, M., Perrot, M., Duchesnay, E., “Scikit-learn: Machine Learning”, 2011. "Machine Learning in Python [Scikit-learn: Machine Learning in Python]", *Journal of Machine Learning Research*, 12, 2825-2830; Implementations of K-nearest neighbors and Gaussian processes are described in the following articles: Rasmussen, C., and Williams, C., "Gaussian Processes for Machine Learning" (2006). "Adaptive Computation and Machine Learning" can be used for regression, MIT Press, Cambridge, MIT.

[0070] When interpupillary distance is used as the target feature, p(PD|X(face), S), i.e., the probability distribution of the true size of the interpupillary distance, can be calculated for a given face shape X(face). p(PD|X(face)) can be considered a special case of the above equation (5), specific to the target feature, in this case, the interpupillary distance PD. An example of determining a Gaussian process will be given below. A Gaussian process allows the use of prior information in data space regions where there is no training data. It also provides a full probability distribution function for the predicted values, not just point predictions. The features and information about the true size of the features are used as input data.

[0071] The goal of a Gaussian process is to model the properties of a function by generalizing to a Gaussian distribution. That is, assuming the data consists of features X measured from facial images, then {X; PD} can be modeled by some unknown function f, with some noise added such that PDi = f(xi) + ε (where the input X will be the facial shape vector and the output PD will be the interpupillary distance). Then f itself can be considered a random variable. Here, this variable will follow a multivariate Gaussian distribution, with its mean and covariance defined as:

[0072]

[0073] Here, m(x) defines the mean function, and κ(x; x') is the kernel function or covariance function.

[0074] m(x)=E[f(x)] (7)

[0075] k(x;x)=E[(f(x)-m(x))(f(x′)-m(x′)) T (8)

[0076] Typically, the mean is set to zero, and the covariance should be chosen flexibly enough to capture the data. Here, if the covariance is written as K in matrix form, the likelihood of the random variable f will be... In practice, this means that the covariance between two input vectors x1 and x2 should be very similar to the covariance between their corresponding output values ​​PD1 and PD2—in other words, two people with similar facial shapes should have similar interpupillary distances. The function f is initialized using a prior Gaussian distribution describing the interpupillary distance, such as prior information based on interpupillary distance statistics. Data points are added to Gaussian processes, which update the distribution on f according to Bayes' theorem, as shown in equation (9) below. This corresponds to the process of training the model with a labeled dataset.

[0077]

[0078] Now that the distribution f is constrained by the training data, we can obtain an estimate of the target value PD or the new input point x. For this, we need to predict the distribution. In the context of Gaussian processes, this would be a Gaussian distribution defined as follows:

[0079]

[0080]

[0081]

[0082] in, The noise variance is defined, assuming that the noise at each data point is independent of the noise at other data points.

[0083] For kernel K, for example, an exponential kernel, a linear kernel, a MateM32 kernel, a MateM52 kernel, or a linear combination thereof can be used. Preferably, an exponential quadratic kernel is used, which yields good results, and its definition is as follows:

[0084]

[0085] Where 1 is the characteristic length scale, and This refers to the signal variance. The values ​​of these parameters and σ... n The values ​​are set to their maximum likelihood estimates, which can be obtained using optimization methods provided by scikit-learn.

[0086] Gaussian processes or any other machine learning method can be trained on training data where the true size is known (e.g., measured by other means), such that the true size of the image with pixel dimensions and the features from the training data are used to approximate the function of the probability in equation (5).

[0087] Therefore, returning to equation (5), the scale of pixels per millimeter in the given image can be determined, and consequently, the true size of any feature, such as interpupillary distance, can be determined. The distance to the camera device can also be determined by the relationship between the pixel size, the true size, and the camera device information described above.

[0088] In some embodiments, the probability distribution is selected based on additional information about the person (e.g., π(θ) in equation (5)). i In this way, a more specific probability distribution than a general probability distribution for all humans can be used. Such additional information can include, for example, gender (male or female, or other gender scores), race (Caucasian, Asian, etc.), body type, or age. For example, when gender, age, and race are known, a probability distribution specific to that combination can be used. Similar considerations apply to body type. Such additional information can be, for example, input by a person or another user. In some embodiments, additional information can also be derived from images. For example, estimates of race or gender can be obtained through image analysis.

[0089] In some embodiments, multiple images may be used instead of a single image, and accuracy may be improved by recognizing features in multiple images.

[0090] The estimated true size or distance can then be used for, for example, fitting eyeglass frames, manufacturing eyeglass lenses, or eye examinations such as the refraction test mentioned earlier.

[0091] In addition, a computer program is provided that contains instructions that, when executed on a processor, cause any of the above methods to be performed.

[0092] A corresponding storage medium is also provided, particularly a tangible storage medium, such as a memory device for storing such a computer program, a hard disk, a DVD or CD, and a data carrier signal for transmitting such a computer program.

[0093] In addition, a corresponding device is provided, which includes:

[0094] Device for providing images of human heads

[0095] A device for identifying multiple features in an image, and

[0096] A means for estimating at least one of the true size of at least one of the plurality of features and the true distance between at least one of the features and a camera device for capturing the image, based on a probability distribution of the true size of at least one of the plurality of features and the pixel size of the at least one feature.

[0097] The device can be configured to perform any of the above methods.

[0098] The techniques discussed above are not limited to applications that determine dimensions and distances associated with head features, but can be generally used to determine the true dimensions of features in an image. In such applications, the image of a human head is replaced by a general image, and the features in the image can be any object in the image, such as a tree, a person, a car, etc. Similarly, for such objects, the probability distribution regarding size is available or measurable. In this case, additional information could be, for example, the brand of the car, the type of tree, etc. Otherwise, the techniques discussed above can be applied.

[0099] Further embodiments will be described with reference to the accompanying drawings, in which:

[0100] Figure 1 This is a block diagram of the device according to an embodiment.

[0101] Figure 2 This is a flowchart illustrating a method according to an embodiment.

[0102] Figure 3 It is a simplified diagram showing various head features of a person, and

[0103] Figure 4 It showcases various features in the landscape images.

[0104] Figure 1 This is a block diagram of the device according to an embodiment. Figure 1The device includes a camera device 10 and a computing device 11. The camera device 10 includes one or more optical elements and an image sensor for capturing images of a human head. The images are provided to the computing device 11. The computing device 11 includes one or more processors. The computing device 11 may be a personal computer, or it may include multiple separate entities communicating with each other to perform tasks as described below. Figure 2 The method is described further. In some embodiments, the camera device 10 and the computing device 11 may be integrated into a single device, such as a smartphone or tablet computer. To perform the method described below... Figure 2 The method is used to program the computing device 11 accordingly.

[0105] Figure 2 This is a schematic block diagram of one embodiment of the method of the present invention. At position 20, an image of a human head is provided. Figure 1 In this embodiment, the image is provided by camera device 10.

[0106] At point 21, multiple features in the image are identified. Figure 3 Some examples of the features are shown in the figure. Figure 3 The features shown include interpupillary distance 30, iris diameter 33, eye width 31, vertical ear length 36, facial length 34, facial width 35, skull width 38, or head height 37.

[0107] Back Figure 2 At point 22, additional information about the person is provided. As explained above, examples of such additional information include body type, gender, age, race, etc.

[0108] At point 23, as described above, the true size of at least one of the multiple features or the true distance between at least one of the multiple features and the camera device 10 is estimated based on the probability distribution of the true size of at least one of the multiple features, preferably at least two of the features, and the pixel size of at least one of the multiple features, preferably at least two of the features.

[0109] As mentioned above, the techniques discussed in this article can be derived from, for example... Figure 3 The head features shown are extended to the size estimation of other features in the image. Figure 4An example is shown. Here, a scene including a tree 40, a car 41, a traffic sign 42, a person 45, a dog 46, and a street 47 is provided. For all these types of objects, there exists a typical size or size probability distribution. The pixel size of these features or objects in image 43 depends on their true size and their distance in the z-direction from the camera device 44 that captured image 43. Using the probability distribution and additional information, such as the species of tree 40, the sex, race, or age of person 45, the breed of dog 46, the brand or type of car 41, etc., such true size and / or distance in the z-direction can be determined.

[0110] Some embodiments are defined by the following examples:

[0111] Example 1. A computer-implemented method for estimating or determining the size or distance of head features, the method comprising:

[0112] Provide head images of (20) people,

[0113] Identify (21) multiple features (30-38) in the image,

[0114] Its features are,

[0115] Based on the probability distribution of the true size of at least one of the plurality of features (30-38) and the pixel size of the at least one of the plurality of features (30-38), at least one of the true size of the at least one target feature among the plurality of features (30-38) or the true distance between the at least one target feature among the plurality of features (30-38) and the camera device (10) used to capture the image is estimated (23).

[0116] Example 2. The method as described in Example 1, wherein the at least one of the plurality of features includes at least two of the plurality of features.

[0117] Example 3. The method as described in Example 1 or 2, characterized in that the features (30-38) include one or more features taken from the group consisting of:

[0118] -Interpupillary distance (30),

[0119] -Iris diameter (33),

[0120] -Pupil diameter,

[0121] -Vertical ear length (36),

[0122] - morphological surface length (34),

[0123] - Width (35),

[0124] - Cranial width (38),

[0125] - Eye width (31), and

[0126] -Head height (37).

[0127] Example 4. The method as described in any one of Examples 1 to 3, characterized in that the estimation includes calculating the probability distribution P(pix per mm|d, θ) of the number of pixels per millimeter of the image according to the following formula.

[0128]

[0129] Where, d i It is the number of pixels spanning the i = 1, 2, ..., N features, π(θ) i θ represents the true size of feature i. i The probability distribution of and / or its covariance with other measured pixel sizes, and P(d i |pixel per mm, θ i ) is given by the probability distribution π(θ) i In the case of ), the value is d per pixel per mm. i The true size θ i The probability calculation.

[0130] Example 5. The method as described in any one of Examples 1 to 4 further includes providing additional information about the person and selecting the probability distributions based on that additional information.

[0131] Example 6. The method as described in Example 5, characterized in that providing the additional information includes receiving the additional information as user input, and / or

[0132] This includes determining this additional information based on the image.

[0133] Example 7. The method as described in Example 5 or 6, characterized in that the additional information includes one or more of the person's gender, age, ethnicity, or body type.

[0134] Example 8. The method as described in any one of Examples 1 to 7, wherein estimating at least one true size of at least one of the features (30-38) includes estimating the interpupillary distance of the person.

[0135] Example 9. The method as described in any one of Examples 1 to 8, characterized in that providing an image includes providing a plurality of images, wherein the estimation (23) is based on the plurality of images.

[0136] Example 10. The method as described in any one of Examples 1 to 9, characterized by one or more of the following:

[0137] -Based on this estimate (23), the eyeglasses frame will be fitted to fit the person's head.

[0138] -Based on this estimate (23), eyeglass lenses can be manufactured, or

[0139] - Perform an eye examination based on this estimate (23).

[0140] Example 11. A device comprising:

[0141] Device (10) for providing images of human heads,

[0142] A device for identifying multiple features (30-38) in the image.

[0143] Its features are,

[0144] A means for estimating at least one of the true size of at least one of the plurality of features (30-38) and the true distance between at least one of the features (30-38) and the means (10) for capturing the image, based on the probability distribution of the true size of at least one of the plurality of features (30-38) and the pixel size of the plurality of features (30-38).

[0145] Example 12. A computer program comprising instructions that, when executed on one or more processors, cause to perform the method as described in any one of Examples 1 to 10.

[0146] Example 13. A data carrier comprising a computer program as described in Example 12.

[0147] Example 14. A data signal comprising a computer program as described in Example 12.

[0148] Example 15. An apparatus (11) comprising at least one processor and a computer program as described in Example 12, the computer program being stored for execution on the at least one processor.

Claims

1. A computer-implemented method for estimating or determining the size or distance of head features, the method comprising: Provide head images of (20) people, Identify (21) multiple features (30-38) in the image, Its features are, Based on the probability distribution of the true size of at least one of the plurality of features (30-38) and the pixel size of at least one of the plurality of features (30-38), at least one of the true size of at least one target feature among the plurality of features (30-38) or the true distance between at least one target feature among the plurality of features (30-38) and the camera device (10) used to capture the image is estimated (23). This estimation involves calculating the probability distribution P(pix per mm|d, θ) of the number of pixels per millimeter of the image according to the following formula. Where, d i It is the number of pixels spanning the i = 1, 2, ..., N features, π(θ) i θ represents the true size of feature i. i The probability distribution of and / or its covariance with other measured pixel sizes, and P(d i |pixel per mm, θ i ) is generated in a given π(θ) i In the case of pixels per mm and θ i To measure d i Likelihood operators.

2. The method as described in claim 1, characterized in that, The at least one of the plurality of features includes at least two of the plurality of features.

3. The method as described in claim 1 or 2, wherein, The probability distribution of the number of pixels per millimeter of the image is calculated based on Monte Carlo Markov chain probability space exploration.

4. The method according to any one of claims 1 to 3, characterized in that, These features (30-38) include one or more features taken from a group consisting of the following: -Interpupillary distance (30), -Iris diameter (33), -Pupil diameter, -Vertical ear length (36), - morphological surface length (34), - Width (35), - Cranial width (38), - Eye width (31), and -Head height (37).

5. The method of any one of claims 1 to 4, further comprising providing additional information about the person and selecting the probability distributions based on the additional information.

6. The method as described in claim 5, characterized in that, Providing this additional information includes receiving it as user input, and / or This includes determining this additional information based on the image.

7. The method as described in claim 5 or 6, characterized in that, This additional information includes one or more of the person's gender, age, ethnicity, or body type.

8. The method according to any one of claims 1 to 7, wherein, Estimating the true size of at least one of these features (30-38) includes estimating the interpupillary distance of the person.

9. The method according to any one of claims 1 to 8, characterized in that, Providing images includes providing multiple images, wherein the estimation (23) is based on the multiple images.

10. The method according to any one of claims 1 to 9, characterized in that... One or more of the following: -Based on this estimate (23), the eyeglasses frame will be fitted to fit the person's head. -Based on this estimate (23), eyeglass lenses can be manufactured, or - Perform an eye examination based on this estimate (23).

11. An apparatus comprising: Device (10) for providing images of human heads, A device for identifying multiple features (30-38) in the image. Its features are, A means for estimating at least one of the true size of at least one of the plurality of features (30-38) or the true distance between at least one of the features (30-38) and the means (10) for capturing the image, based on a probability distribution of the true size of at least one of the features (30-38) and the pixel size of the plurality of features (30-38), the estimation comprising calculating a probability distribution P(pix per mm|d, θ) of the number of pixels per millimeter of the image according to the following formula. Where, d i It is the number of pixels spanning the i = 1, 2, ..., N features, π(θ) i θ represents the true size of feature i. i The probability distribution of and / or its covariance with other measured pixel sizes, and P(d i |pixel per mm, θ i ) is generated in a given π(θ) i In the case of pixels per mm and θ i To measure d i Likelihood operators.

12. The device as claimed in claim 11, wherein, The device is configured to perform the method as described in any one of claims 1 to 10.

13. A computer-readable medium comprising instructions that, when executed on one or more processors, cause to perform the method as claimed in any one of claims 1 to 10.

14. A data carrier comprising the computer-readable medium as described in claim 13.

15. An apparatus (11) comprising at least one processor and a computer-readable medium as claimed in claim 13, the computer-readable medium being stored for execution on the at least one processor.

Citation Information

Patent Citations

  • Systems and methods for obtaining accurate body size measurements from two-dimensional image sequences

    JP6392756B2

  • Fitting of spectacles

    US20100220285A1

  • Method for estimating volumetric distance maps from 2D depth images

    US6262738B1

  • Method and apparatus for adjusting the display scale of an image

    US6535233B1

  • Measurement of dimensions of solid objects from two-dimensional image(s)

    US6995762B1