Methods and mobile devices for determining a person's viewpoint
By using readily available mobile devices to generate viewpoint avatars, the problem of hardware dependence in existing technologies is solved, enabling flexible and rapid viewpoint measurement and supporting personalized eyeglass lens design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies require specific hardware devices such as eye trackers or head tracking devices to determine the viewpoint, which leads to inconvenient measurements and laboratory dependence, and manual measurements are time-consuming.
Using readily available mobile devices such as smartphones or tablets, combined with image sensors and optional depth sensors, the viewpoint is determined by generating a surrogate, avoiding dependence on specific hardware. Machine learning algorithms are used to detect head and pupil positions, enabling flexible measurement of the viewpoint.
It enables the rapid and convenient determination of viewpoint distribution without the need for specific hardware, is applicable to multiple target locations, adapts to individual visual behavior, and supports personalized eyeglass lens design.
Smart Images

Figure CN118696260B_ABST
Abstract
Description
[0001] This application relates to methods and mobile devices for determining a person's viewpoint, and to corresponding computer programs and methods for manufacturing spectacle lenses based on the determined viewpoint.
[0002] To fit a particular eyeglass lens to a particular frame for a particular person, various parameters are traditionally determined; these parameters are called centering parameters. Centering parameters are the parameters required to properly position the lens within the frame (i.e., to center it) so that the lens is worn in the correct position relative to the human eye.
[0003] Examples of such centering parameters include interpupillary distance, vertex distance, face lens radius, y-coordinates of the left and right centering points (also specified as fitting point height), distance point, tilt angle, and other parameters defined in Section 5 of DIN EN ISO 13666:2012, as well as the tilt of the frame.
[0004] To further design and manufacture eyeglasses tailored to specific individuals, it is important to understand individual eye behavior at different distances. This eye behavior depends on the viewing distance itself, the scene, and the convergence of the eyes. Eye behavior determines where the eyeglass wearer looks through the lenses. According to DIN EN ISO 13666:2012 5.11, the point where a person's line of sight intersects with the rear surface of the lens when looking at a specific target is called the viewpoint, which corresponds to the aperture on the lens plane. In the context of this application, the thickness and / or curvature of the lens can sometimes be neglected, such that the viewpoint is considered as the intersection of the line of sight with a plane approximating the lens position.
[0005] Knowing the viewpoints at different distances (e.g., the distance point according to DIN EN ISO 13666:2012 5.16 or the near point according to DIN EN ISO 13666:2012 5.17) is particularly important for the design and centering of progressive multifocal lenses (PAL), where different parts of the lens provide optical correction for different distances (e.g., for myopia or hyperopia). The distance point is the assumed location of the viewpoint on the lens, used for distance vision under given conditions, and the near point is used for near vision in the same way. The myopic part of the lens should, for example, be at or around the near point, while the hyperopic part of the lens should be at or around the distance point. Similarly, knowing the (multiple) viewpoints can be helpful for the design and centering of single-vision lenses (SVL).
[0006] To illustrate, Figure 8An example viewpoint distribution for a single-vision lens within the lens's field of vision, monitored over a specific time period, is illustrated in the form of a heatmap. This heatmap represents the location of the viewpoint during typical office tasks, including writing, reading, and switching between two screens. In other words, the viewpoint (i.e., the point of intersection between the person's line of sight and the lens) is determined over time and... Figure 8 These viewpoints are illustrated in a simplified diagram. Here, when a person looks at different targets, they can move their eyes, resulting in different viewpoints, or move their head so that the viewpoint remains unchanged while the head moves, or a combination of both (e.g., head movement combined with slight eye movement). In the example shown, there are preferred viewpoint areas at approximately 20 and 40 mm horizontally and -20 and -35 mm vertically, and viewpoints are also detected in some other areas within the field of vision. This means that a particular person wearing these particular glasses mostly views through the preferred viewpoint areas, but may sometimes move their eyes to view through other parts of the lens.
[0007] Figure 9 This demonstrates the viewpoint of progressive multifocal lenses relative to... Figure 8 The example distribution of the same task is shown. Here, two main viewpoint regions can be seen, one for farsightedness and the other for nearsightedness. The regions for farsightedness and nearsightedness in progressive multifocal lenses are generally limited compared to the entire region in single-vision lenses, requiring more head movement in addition to the transitions between the far and near fields that are normally accomplished through eye movements. Therefore, the study found that people who typically wear single-vision lenses use more eye movements, i.e., they "move their eyes" more than people who wear progressive multifocal lenses, while the latter "move their heads" more.
[0008] The position and distribution of the viewpoint depend not only on the type of lens (PAL or SVL) but also on the person wearing the lens. For example, a person may move their eyes more or their head more. Therefore, knowing a person's viewpoint or viewpoint distribution for a specific task (such as reading and writing) is helpful for producing progressive multifocal lenses specifically adapted to eyeglass wearers.
[0009] Various systems are known for measuring individual visual parameters, including a person's viewpoint or viewpoint distribution, and for designing lenses accordingly, such as Essilor Visioffice 2 (https: / / www.pointsdevue.com / article / effect-multifocal-lenses-eye-and-head-movementspresbyopic-vdu-users-neck-and-shoulder), Essilor Visioffice X (https: / / www.pointsdevue.com / white-paper / varilux-x-seriestm-lenses-ne ar-vision-behaviorpersonalization), or EssilorVisiostaff (https: / / www.essilorpro.de / Produkte / Messen-und-Demonstrieren / visioffi ce / Documents / Visiostaff_Broschuere.pdf).
[0010] Some of the techniques used in Visioffice 2 are described in EP 3145386 B1. This document generally describes a method for determining at least one parameter of human visual behavior. The rotation center of at least one eye is determined in a first reference frame associated with the human head. Then, in a real-life scenario such as reading, an image of the person (including body parts) is captured, and the position and orientation of the body parts are determined in a second reference frame. Finally, the rotation center of the eye is then determined in the second reference frame. Based on the knowledge of the rotation center, eye parameters are determined. The viewpoint can be determined by mounting an eye tracker to a pair of glasses worn by the individual. This application requires an eye tracker mounted on a pair of glasses frames.
[0011] A somewhat similar system is disclosed in WO 2015 / 124574 A1, which also uses an eye tracker mounted on an eyeglass frame. Using this frame-mounted eye tracker, eye movements can be observed over longer periods. However, this requires the person to wear corresponding eyeglasses with the eye tracker attached.
[0012] Therefore, the above solution requires an eye tracker that is installed in the eyeglass frame, i.e., specific hardware.
[0013] EP 1815289 B9 discloses a method for designing spectacle lenses that takes into account the movement of a person's head and eyes. Modifications are made to commercially available head and eye movement measuring devices to determine the viewpoint, requiring specific hardware in this case.
[0014] EP 1747750 A1 discloses another method and apparatus for determining a person's visual behavior (e.g., viewpoint) for the purpose of customizing eyeglass lenses. The person must wear a specific headband with LEDs to track head movements, which can be inconvenient for that person.
[0015] EP 2342599 B1 and EP 1960826 B1 each disclose methods for determining human visual behavior based on manual measurements, with head tilt in the former via convergent data and via wires in the latter. Therefore, the various solutions discussed above all require specific equipment. Some applications using eye- and head-tracking devices may only be reliably used in laboratory settings or require time-consuming manual measurements.
[0016] US2020 / 349 754A1 discloses a method for generating a 3D model of a head using a smartphone, in which a depth map is used.
[0017] US2021 / 142 566A1 disclosed virtual try-on of eyeglass frames.
[0018] US 9 645 413 B2 and US 9 841 615 B2 discuss methods for determining viewpoints.
[0019] EP 3413122 A1 discloses a method and apparatus for determining a near-field viewpoint, wherein, in some embodiments, a device such as a tablet PC can be used to measure the near-field viewpoint. The apparatus can be used in conjunction with a centering device. The position and / or orientation of a near target, such as a tablet PC, as mentioned above, are determined, and an image of a person is acquired, such as one captured while the person is looking at the target. The near-field viewpoint is then determined based on the image, position, and / or orientation. This method essentially uses static measurements for a single position. In some embodiments, a head surrogate, i.e., a model, or a virtual eyeglasses frame adapted to the surrogate, can be used, and the viewpoint can be determined using the surrogate adapted to the virtual eyeglasses frame. To generate the surrogate, a fixed apparatus with multiple cameras generally set in a semicircle is used. Therefore, although a mobile device such as a tablet PC is involved, a specific fixed apparatus is still required to utilize the surrogate. Furthermore, only a single near-field viewpoint is determined. Additionally, a far-field viewpoint can be determined.
[0020] US2021 / 181 538A1 discloses another method that uses a double to determine the viewpoint, in which only a single viewpoint is detected.
[0021] For example, starting with EP 3413122 A1, the objective of this invention is to provide a flexible method for determining the viewpoint, which can use a virtual eyeglass frame but does not require specific hardware, and enables the measurement of the viewpoint distribution. Compared to other references cited above, this would eliminate the need for specific hardware and manual measurements (which would require trained personnel like opticians).
[0022] According to a first aspect of the invention, a computer-implemented method is provided for a mobile device including at least one sensor to determine a person's viewpoint, the at least one sensor including an image sensor and optionally a depth sensor, the method comprising:
[0023] At least one sensor is used to determine the person's head position and pupil position when the person is looking at the target, and
[0024] The viewpoint is determined based on a stand-in of at least the eye portion of the person's head, the stand-in being set to the determined head and pupil position.
[0025] The method is characterized by moving the target within a person's field of vision and determining the viewpoint for multiple locations of the target. In this way, viewpoints can be determined for multiple target locations. It is also possible to detect from the determined viewpoint whether the person is a head-moving or eye-moving operator, as initially explained, i.e., whether the person prefers to move their head to follow the target while the head moves or only moves their eyes. The person or another person can, for example, perform the movement of the target based on instructions output from a mobile device.
[0026] As used herein, a mobile device refers to a device designed to be easily carried. Typically, such a mobile device weighs less than 2 kg, usually less than 1 kg, or even less. Examples of such mobile devices include smartphones or tablet PCs. Nowadays, in addition to processors and memory, smartphones or tablet PCs include several additional components for implementing the method, particularly sensors. For example, smartphones or tablet PCs nowadays almost always include a camera, and in many cases also include depth sensors, such as time-of-flight sensors, which can be used to determine head and pupil positions and generate avatars. Furthermore, such depth sensors enable smartphone software to detect facial features and / or landmarks in an image and output the 3D coordinates of these landmarks. For example, head position can be determined using a depth sensor that measures the distance of different parts of the head and / or facial landmarks from the mobile device, and pupil position can be determined using the mobile device's camera, as described, for example, in EP 3413122A1 cited above, or pupil position can also be detected as a facial landmark.Facial landmark detection can be performed using computer vision or machine learning techniques, as described in the following literature: Wu, Y., Hassner, T., Kim, K., Medioni, G., & Natarajan, P. (2017), Facial landmark detection with tweaked convolutional neural networks, IEEE transactions on pattern analysis and machine intelligence, 40(12), 3067-3074; Perakis, P., Passalis, G., Theoharis, T., & Kakadiaris, IA (2012), “3D facial landmark detection under large yaw and expression variations”, IEEE transactions on pattern analysis and machine intelligence, 35(7), 1552-1564; or Wu, Y., & Ji, Q. (2019), Facial landmark detection: A literature survey. International Journal of Computer Vision, 127(2), 115-142.
[0027] Alternatively, only an image sensor can be provided, and signs can be detected in the image. Head position and pupil position can then be detected based on the image and the signs detected therein. Various mobile operating systems offer built-in solutions for this, such as ARKit for iOS devices and ARCore for Android devices. These solutions use trained machine learning algorithms capable of determining 3D coordinates from the 2D contours of a face in an image. While the use of a depth sensor is preferred and can produce more accurate measurements, it is an alternative for mobile devices that do not have a depth sensor.
[0028] Therefore, this method can be implemented by programming an off-the-shelf mobile device accordingly (using a so-called "app"), and no specific hardware is required. After generating a avatar using at least one camera and / or at least one depth sensor, the relative positional changes of the head, pupils, and facial landmarks are used to calculate the user's movement relative to the mobile device.
[0029] As commonly understood in the field of computing, a surrogate is a graphical representation of a person. In the context of this application, the term surrogate should be understood as a 3D surrogate, in other words, a three-dimensional model representation of a person or a part of their body. Such a 3D surrogate is also referred to as a 3D model.
[0030] A model, especially a 3D model, should be understood as a representation of a real object existing as a dataset in a storage medium (e.g., computer memory or data carrier), which in the case of a 3D model is a three-dimensional representation. For example, this three-dimensional representation could be a 3D mesh, which consists of a set of 3D points (also called vertices) and the connections between these points (also called edges). In the simplest case, these connections form a mesh of triangles. This 3D mesh representation describes only the surface of the object, not its volume. The mesh does not necessarily have to be closed. Therefore, for example, if a head is described as a mesh, it would look like a mask. Detailed information about this 3D model can be found in the following literature: Rau JY, Yeh PC, “A Semi-Automatic Image-Based Close Range 3D Modeling Pipeline Using a Multi-Camera Configuration”, Sensors (Basle, Switzerland), 2012; 12(8):11271-11293. doi:10.3390 / s120811271, specifically page 11289, “Figure 16”.
[0031] A voxel grid representing volumetric representation is a further option for representing 3D models. Here, space is divided into small cubes or cuboids, which are called voxels. In the simplest case, the presence or absence of the object to be represented is stored as a binary value (1 or 0) for each voxel. With a voxel side length of 1 mm and a volume of 300 mm × 300 mm × 300 mm, it represents a typical volume of a head, thus yielding a total of 27 million voxels. Such a voxel grid is described, for example, in the following literature: M. Nieβner, M. S. Izadi and M. Stamminger, “Real-time 3D reconstruction at scale using voxel hashing”, ACM Trans. Graph. 32, 6, Article 169 (November 2013), at doi.org / 10.1145 / 2508363.2508374.
[0032] A substitute for at least the eyes of a head means that the substitute includes at least the eyes of a person and their surrounding environment. In some embodiments, the substitute may be, for example, a substitute for the entire face or the entire head of a person, and may include other parts of the person, as long as the eyes are included.
[0033] Using a stand-in, only a few landmarks of the head must be determined during head positioning, allowing the stand-in to be set up based on the determined head position and / or facial landmarks. Similarly, since the stand-in includes the eyes, pupil position can be determined more easily. Ultimately, by using a stand-in, the method can be implemented without the person wearing specific equipment and without the need for specific eye-tracking devices, such as those in eyeglass frames.
[0034] In one alternative, a person may wear eyeglass frames. Here, according to DIN EN ISO 7998 and DIN EN ISO 8624, eyeglass frames should be understood to mean frames or retainers through which lenses can be placed on the head. In particular, the terminology used herein also includes rimless eyeglass frames. In this case, the method may include detecting eyeglass frames and determining the viewpoint based on the detected eyeglass frames. Detecting eyeglass frames is described, for example, in EP 3 542 211 A1. In another alternative, virtual eyeglass frames may be provided to the substitute in a so-called virtual fitting, known as virtual fitting. Such virtual fittings are described, for example, in US2003 / 0123026 A1, US2002 / 105530A1, or US2016 / 0327811A1. In this case, the method may include determining the viewpoint based on the corresponding virtual eyeglass frames fitted to the substitute. The advantage of using such virtual eyeglass frames is that the viewpoint can be determined for multiple eyeglass frames without the person actually wearing eyeglass frames.
[0035] Determining the position of a person's head and pupils can be performed relative to the target. In this way, based on the determined head and pupil positions of the stand-in relative to the target, the person's line of sight can be determined, and the intersection of the line of sight and the lens plane can be regarded as the viewpoint.
[0036] Setting the Stand to a specific eye and pupil position means that the Stand's eye and pupil position is set to match the specific eye and pupil position.
[0037] This method can preferably be further characterized by generating an avatar from a mobile device.
[0038] Therefore, unlike the method in EP 3413122 A1, a mobile device is used to generate the avatar. This has the advantage of not requiring a specific fixed device. Furthermore, compared to the other methods discussed above, it does not require specific hardware.
[0039] In some embodiments, in order to provide a substitute, a person can rotate their head about the longitudinal axis of the head (vertical direction) while the mobile device measures the head using a depth sensor and optionally also an image sensor of a camera.
[0040] The combined depth and 2D camera images captured in this manner will be referred to hereinafter simply as combined images, and in the case of color camera images, also as RGB-D images. In other embodiments, additionally or alternatively, the head can be rotated about another axis, such as the horizontal axis. To identify the double, the head can be held in an upright position, and the gaze can be directed horizontally into the distance (normal gaze direction). In this way, an individual double can be identified without additional equipment.
[0041] A major problem with using smartphones to generate avatars is that, compared to using fixed equipment like in EP 3413122A1, the camera pose relative to the head is unknown when capturing the head with a depth sensor and camera, whereas in a fixed arrangement, the camera pose is known through the design of the apparatus. Generally, the term pose refers to a combination of position and orientation, as defined, for example, in DIN EN ISO 8373:2012-03.
[0042] Since the way 3D objects such as heads are imaged onto image sensors or captured by depth sensors is determined by the characteristics of known devices (such as the focal length of lenses or sensor resolution), the problem of determining the camera pose is equivalent to the problem of registering the combined images mentioned above.
[0043] Typically, a composite image refers to a 2D image and a depth image acquired essentially simultaneously from corresponding locations. The 2D image can be a color image, such as an RGB image (red, green, blue), or it can be a grayscale image. The depth image provides a map of the distance between the camera and the object (in this case, the head). To capture the 2D portion of the composite image, any conventional image sensor can be used in conjunction with the corresponding camera optics. To capture the depth image, any conventional depth sensor, such as a time-of-flight sensor, can also be used. A composite image can include two separate files or other data entities, where one data entity (e.g., pixels) provides a grayscale or color value for each 2D coordinate, and the other data entity provides a depth value for each 2D coordinate. A composite image can also include only a single data entity, where both grayscale / color information and depth information are provided for each 2D coordinate. In other words, as long as both grayscale / color information and depth information of the scene captured in the image are available, the way the information is stored in a data entity such as a file is irrelevant. Cameras suitable for capturing composite images (in this case, color images) are also called RGBD cameras (red, green, blue, depth). Some modern smartphones or other mobile devices are equipped with such RGBD cameras. In other cases, an RGBD camera or depth sensor (which is then used in conjunction with the smartphone's built-in camera) can also be attached to the smartphone. It should be noted that the depth image does not need to have the same resolution as the 2D image. In such cases, scaling operations (reducing or enlarging) can be performed to adapt the resolution of the 2D image and the depth image to each other. The result is essentially a point cloud, where each point has a 3D coordinate based on its 2D coordinates in the image and its depth coordinates from the depth sensor, along with a pixel value (color or grayscale value).
[0044] Image registration typically involves finding a transform that alters one of the combined images to the other. Such transforms can include rotational, translational, and magnification components (magnification greater than or less than one) and can be written in matrix form. As mentioned above, performing registration by determining the aforementioned transforms is essentially equivalent to determining the (relative) camera pose (position and orientation) of the captured combined images.
[0045] In some embodiments, such registration may include pairwise coarse registration based on 3D markers obtained on the combined image and fine registration based on the complete point cloud represented by the combined image.
[0046] Landmarks are predefined points on the head. Such landmarks may include, for example, the tip of the nose, points on the bridge of the nose, corners of the mouth or eyes, points forming the shape of eyebrows, etc. These landmarks in combined 2D and depth images can be determined using various conventional methods. For example, trained machine learning logic, such as a neural network, can be used to determine the landmarks. In this case, for training, multiple combined 2D and depth images from different locations and for multiple different heads are used as training data, where landmarks can be manually labeled. After training, the trained machine learning logic determines the landmarks. Details can be found, for example, in the following literature: Wu, Y., Hassner, T., Kim, K., Medioni, G., & Natarajan, P. (2017), Facial landmark detection with tweaked convolutional neural networks, IEEE transactions on pattern analysis and machine intelligence, 40(12), 3067-3074; Perakis, P., Passalis, G., Theoharis, T., & Kakadiaris, IA (2012), “3D facial landmark detection under large yaw and expression variations”, IEEE transactions on pattern analysis and machine intelligence, 35(7), 1552-1564; or Wu, Y., & Ji, Q. (2019), Facial landmark detection: A literature survey. International Journal of Computer Science Vision [Face Marker Detection: Literature Review. International Journal of Computer Vision], 127(2), 115-142 (2018).
[0047] Pairwise coarse registration provides coarse alignment between pairs of 3D markers. In an embodiment, this pairwise coarse registration estimates a transformation matrix between the markers of the two combined images that aligns the markers in a least-squares sense, i.e., minimizes the error. Minimize, where, It is the i-th label of the j-th image, and This is the transformation matrix from the second image (j=2) of the corresponding pair to the first image (j=1) of the corresponding pair. This coarse registration can be performed by a method called point-to-point ICP (“iterative nearest point”), which is described, for example, in the following literature: Besl, Pazl J. and McKay, Neil D., “Method for registration of 3D shapes”, Sensor Fusion IV: control paradigms and data structures, Vol. 1611, International Society for Optics and Photonics, 1992. Preferably, to eliminate potential outliers that may arise in the label determination step, a random sampling consensus procedure can be used, as described in the following literature: Fischler, Martin A. and Bolles, Robert C., “Random sampling consensus: a paradigm for model fitting with applications to image analysis and automated cartography”, Communications of the ACM 24.6(1981):381-2395. The above and other formulas presented herein use so-called homogeneous coordinates, as is often used in computer vision applications. This means that the transformation T is represented as a 4×4 matrix [R t; 0 0 0 1], where R is a 3×3 rotation matrix, t is a translation vector, and the last row is 0 0 0 1. The 3D point (x,y,z) is augmented using a homogeneous component w, i.e., (x,y,z,w), where w = 1 is typically used. This allows translation and rotation to be included in a single matrix multiplication; that is, the matrix multiplication can be written as x2w = Tx1w (where x1w and x1w are corresponding vectors in homogeneous coordinates) instead of x2 = Rx1 + t (where the vectors x2 and x1 are in Cartesian coordinates). However, this is merely a matter of sign, and the same computation can also be performed in Cartesian or other coordinate systems.
[0048] Fine registration refines the transformations mentioned above, making them more accurate. To perform this fine registration, a complete point cloud (i.e., a complete RGBD image or a point cloud derived from such images) representing a combination of a 2D image and a depth image can be used. In particular, color information can also be used. Since coarse registration has already been performed, fine registration can be performed more efficiently than registration using only the point cloud or a corresponding grid.
[0049] Different methods can be used for fine registration. In some embodiments, the method chosen may depend on the residual error after coarse registration, such as the error e mentioned above, or other error quantities (e.g., the angular inverse difference between markers in the corresponding combined 2D image and depth image pairs based on the transformation determined after coarse registration). For example, angular difference or inverse difference can also be used. If such deviations occur (e.g., error e is less than a certain value, the angular difference between markers is less than a certain threshold angle (e.g., 5°), or the positional difference is less than a certain position (e.g., 5 cm, e.g., as the average value of the markers)), then RGBD ranging can be used for fine registration, taking into account not only the depth coordinates but also the color of each point in the point cloud. RGBD ranging is described, for example, in the following literature: Park, Jaesik, Zhou, Quan-Yi and Koltun, Vladlen, “Coloredpoint cloud registration revisited”, Proceedings of the IEEE International Conference on Computer Vision, 2017. For cases with significant differences, point-to-plane ICP can be used to register the image on the point cloud, as described in, for example, the following literature: Rusinkiewicz, Szymon and Levoy, Marc, “Efficient variants of the ICP algorithm”, Proceedings of the third international conference on 3D digital imaging and modeling, IEEE 2001.
[0050] For both alternative schemes, it is preferable to perform two registrations: one to estimate the transformation from the first combined image of the corresponding pair to the second combined image of the corresponding pair, and one to estimate the transformation from the second combined image to the first combined image, wherein the initial values of the algorithm are slightly different. This can help improve the overall accuracy. In other words, determine... and Determine the error between the two registrations If the registration is stable, the error should be close to the identity matrix I4 (i.e., a diagonal matrix with only values of 1). If the error (i.e., the deviation from the identity matrix) is less than a certain threshold, the corresponding transformation can be added as an adjective between the corresponding combined images to the so-called pose diagram. In some embodiments, the covariance of the transformation can also be determined.
[0051] Preferably, for both coarse and fine registration, instead of using all possible pairs of combined 2D and depth images, the pairs to be registered can be determined based on the classification of the combined 2D and depth images with respect to the orientation of the head from which these images were captured (e.g., based on a so-called presumed matching map indicating pairs from which registration can be performed). This can be accomplished using approximate pose data or other approximate information to determine the combined images, where the poses of the captured combined images are sufficiently similar to allow for reasonable registration. For example, if a combined image is acquired from the left side of the head and another combined image is captured from the right side of the head, it is difficult to obtain any common landmarks from the two images (e.g., the left eye is only visible on the left, the right eye is only visible on the right, and so are the left and right corners of the mouth). Therefore, classification can begin with a frontal image, categorizing the combined images into classes such as left, right, top, and bottom. Within each class, pairwise coarse and fine registration as described above is performed.
[0052] In some embodiments, the classification may be based on metric data from the image recording device itself used to capture the combined image (e.g., from the ARKit tool in the case of iOS-based devices or from the ARCore tool in the case of Android-based devices). In other cases, the presumed matching map may be obtained from the aforementioned flags via 2D / 3D correspondences and a perspective n-point solver, as described, for example, in the following literature: Urban, Steffen, Leitloff, Jens and Hinz, Stefan, “MIpnp - a real-time maximum likelihood solution to the perspective n-point problem”, arXiv preprint, arXiv:1607.08112 (2016).
[0053] In this way, the method can avoid attempting to register combined images that are difficult or impossible to do so due to the lack of common markers that improve robustness.
[0054] Based on the registration, i.e., the transformation described above, the pose can then be estimated for each composite image in the global reference frame. These poses can be head poses represented by the composite images or camera poses. As explained above, the camera pose when capturing the composite images is directly related to the image registration, and therefore to the head pose in the composite images, such that if the camera pose is known, the head pose can be determined, and vice versa. The pose map is then preferably optimized. This optimization is used for pose map optimization based on registration (including for each generated pose M in the composite images). j Possible methods for this are described in the following literature: Sungjoon, Zhou, Qian-Yi and Vladlen, Koltun, “Robust reconstruction of indoor scenes”, proceedings of the IEEE conference on computer vision and pattern recognition, 2015.
[0055] Based on these postures M j Provides an attitude graph p = {M, e}, which is composed of nodes M (i.e., attitude M). j ) and edge E (which is a transformation when determined) and possible covariance )composition.
[0056] Based on this pose graph, for further optimization, so-called edge pruning can then be performed. This can be used to remove edges... The error estimation is used to determine whether the pose is estimated from valid edges (i.e., valid transformations obtained in registration). For this, one pose is considered as a reference. Then, the edges of the pose graph are the shortest paths connecting the unoptimized pose graph to another node. This yields a further pose estimate for the node to be tested. This further estimate is then compared to the pose from the optimized pose graph. If the deviation is high in this comparison (i.e., the deviation is above a threshold), the corresponding edge can be identified as erroneous. These edges can then be removed. After this, the pose graph optimization process mentioned above is repeated without the removed edges until there are no more erroneous edges in the graph.
[0057] Based on the registration and thus determined camera pose, a surrogate can then be created from the image in a manner substantially similar to that of a fixed camera setup. This can include: fusing point clouds obtained from separate composite images based on the registration; generating meshes from the fused point clouds, for example using Poisson reconstruction, as described in Kazhdan, Michael, Bolitho, Matthew and Hoppe, Hugues, “Poisson surface reconstruction”, Proceedings of the fourth Eurographics symposium on Geometry processing, Vol. 7, 2006; and texturing the meshes from the image, for example as described in Waechter, Michael, Moehrle, Nils and Goesele, Michael, “Let there be color! Large-scale texturing of 3D reconstructions”, European Conference on Computer Vision, Springer, Cham, 2014.
[0058] Another method for generating 3D models of heads using smartphones is disclosed in WO2019 / 164502A1.
[0059] The target can be in the near field. In this way, the near point can be determined. The near field is the area where a person uses myopia, which includes the typical reading distance or distance for performing similar tasks, such as a distance of up to or approximately 40 cm from the eyes, as given under 5.28.1 in DIN DIN EN ISO 13666:2012.
[0060] The target can be a mobile device, for example, a feature of the mobile device such as a camera, or a target displayed on the mobile device. The relative position and orientation (i.e., pose) of the mobile device (target) relative to the head can then be determined using the mobile device's sensors, such as a 2D camera and depth sensor as mentioned above, when the target is moved within the field of view. For example, facial landmarks can be detected for each location as described above, and pose can be detected based on the facial landmarks, corresponding to the coarse registration process described above. Optionally, to refine the pose, fine registration as described above and further, optionally, pose map registration as described above, can be used. In this way, the pose of the mobile device can be accurately determined.
[0061] The target can also be or include 3D landscapes.
[0062] Multiple viewpoints can be provided to a database and displayed as graphs, for example, on a user interface on a mobile device screen or computer screen (as a platform that also provides other information about human vision, such as refraction), enabling virtual try-on of eyeglass frames, etc. Furthermore, such a platform can provide information on recommended ophthalmic lens optical designs based on individual points and explain the optical designs to eye care professionals or eyeglass wearers. Multiple viewpoints recorded at different time points or using different ophthalmic lens designs and / or eyeglasses are stored in the database. The platform may include the ability to directly compare the displayed graphs from different time points or using different ophthalmic lens designs and / or eyeglasses from the database.
[0063] This information can be used to custom-make eyeglasses for individuals. Examples of such customization are described in the following literature: Sheedy, JE (2004), “Progressive addition lenses—matching the specific lens to patient needs”, Optometry-Journal of the American Optometric Association, 75(2), 83-102; or Sheedy, J., Hardy, RF, & Hayes, JR (2006), “Progressive addition lenses—measurements and ratings”, Optometry-Journal of the American Optometric Association, 77(1), 23-39.
[0064] Therefore, according to another aspect, a method for manufacturing spectacle lenses is provided, comprising:
[0065] Determine the viewpoints discussed above, and
[0066] Eyeglass lenses are manufactured based on a defined viewpoint, such as PAL lenses.
[0067] In this way, eyeglass lenses can be better customized for individuals.
[0068] As mentioned, this method can be implemented using readily available mobile devices (such as smartphones or tablet PCs) equipped with corresponding sensors. In this case, computer programs for such mobile devices can be provided to implement any of the methods described above. In the context of mobile devices, such computer programs are generally referred to as applications or "apps".
[0069] According to another aspect, a mobile device is provided, comprising:
[0070] The sensor is configured to determine the position of a person's head and pupils when the person is looking at the target, and
[0071] A processor configured to determine a person's viewpoint based on a stand-in provided at least for the person's eyes, the stand-in being set to the determined head and pupil positions.
[0072] Its features are,
[0073] The mobile device is further configured to determine the person's viewpoint from multiple locations when the target is moved within the person's field of vision.
[0074] The mobile device can be configured, for example, programmed for any of the methods discussed above.
[0075] In this regard, a processor refers to any entity capable of performing corresponding calculations, such as a microprocessor or microcontroller that is programmed accordingly. As previously mentioned, a mobile device can be an off-the-shelf mobile device, such as a smartphone or tablet, which typically provides corresponding sensors or processors. Sensors can include, for example, depth sensors, time-of-flight sensors, cameras, or both.
[0076] The above concepts will be further explained using specific embodiments, wherein:
[0077] Figure 1 This is a block diagram of a mobile device that can be used to implement the embodiments.
[0078] Figure 2 This is a flowchart of the method according to an embodiment.
[0079] Figures 3 to 7 It is used for explanation Figure 2 A simplified diagram of the method, and
[0080] Figure 8 and Figure 9 This is a heatmap showing the distribution of example viewpoints.
[0081] Figure 1 A mobile device 10 according to an embodiment is shown. The mobile device 10 is a smartphone or tablet PC. Figure 1Some components of a mobile device 10 that can be used to implement the methods described herein are shown. The mobile device 10 may include other components conventionally used in mobile devices such as smartphones, which are not required for explaining the embodiments and are therefore not included in this description. Figure 1 As shown in the image.
[0082] As sensors, the mobile device 10 includes a camera 11 and a depth sensor 12. In many cases, the mobile device includes more than one camera, such as a so-called front-side camera and a so-called rear-side camera. Both types of cameras can be used in embodiments. The depth sensor 12 can be, for example, a time-of-flight based depth sensor as typically provided in mobile devices.
[0083] For inputting and outputting information, the mobile device 10 includes a speaker 13 for outputting audio signals, a microphone 14 for receiving audio signals, and a touchscreen 19. By outputting signals to the speaker 13 or displaying information on the touchscreen 19, the mobile device 10 can output instructions to a person to perform the methods described herein. The mobile device 10 can receive feedback from a person via the microphone 14 and the touchscreen 19.
[0084] Furthermore, the mobile device 10 includes a processor that controls various components of the device 10 and executes programs stored in a storage device 15. The storage device 15 in mobile devices such as smartphones or tablet PCs is typically a flash memory, etc. The storage device 15 stores computer programs that are executed by the processor 16 to perform the following... Figure 2 The method is described further.
[0085] The mobile device 10 further includes an accelerometer 17 and an orientation sensor 110. Using the accelerometer 17, the acceleration of the mobile device 10 can be determined, and the position of the mobile device 10 compared to its initial position is determined by integrating the acceleration twice. The orientation (e.g., tilt angle) of the mobile device 10 can be determined via the orientation sensor 110. The combination of position and orientation is also referred to as attitude, see DIN EN ISO 8373:2012-03.
[0086] Figure 2 A method according to an embodiment is shown, which can be achieved by... Figure 1 The storage device 15 of the mobile device 10 provides a corresponding computer program to be executed by the processor 16.
[0087] In step 20, Figure 2 The method involves identifying a double for at least the eye area of the person's head to be examined. This further... Figure 3The following is illustrated. Here, a human head 30 with eyes 31 is schematically shown. While the person (or another person) is looking directly at a distant target as shown by line 32, the person holds a mobile device 10 (e.g., a smartphone in this case) in front of the head 30. The person then rotates the head about a vertical axis 33 while the mobile device uses a depth sensor 12 and a camera 11 to measure the head at multiple rotational positions. The processor 16 then calculates at least the eye portion of the head 30 (i.e., the portion including eye 31 as shown and another eye (not shown)) and preferably a substitute for the entire head from these measurements. Alternatively, the mobile device 10 can move around the head 30, and the head 30 can be measured at different positions on the mobile device 10. The position of the mobile device 10 can be determined based on an accelerometer 17.
[0088] Return to Figure 2 At point 21, the method includes measuring head and pupil positions while a person is looking at the target. This is in Figure 4 It was showcased in [the document / platform]. Figure 4 In this context, mobile device 10 serves as a target, for example, by displaying a target on touchscreen 19 that a person should be looking at. For example, to simultaneously issue instructions to a person, mobile device 10 may display text such as "Look here now" on touchscreen 19. Mobile device 10 is held, for example, in a reading position or another typical near-field position. To look at the target, the person tilts their head 30 at an angle between line 32 and line 41. Furthermore, the person may rotate their eyes 13 to form a line of sight. Using depth sensor 12 and camera 11, smartphone 10 detects pupil position and head position. Additionally, in some embodiments, the person may wear eyeglasses 40, wherein mobile device 10 also detects the position of the eyeglasses, as discussed above. In other embodiments, as explained above, the person avatar determined in step 20 may be fitted with virtual eyeglasses. Although in Figure 4 The target is a mobile device, but people can also look at other targets or 3D landscapes or scenes, and the mobile device 10 is only used for measurement.
[0089] At point 22, the method includes setting the stand-in's head and pupil positions as measured at point 21. In other words, the head is set to the posture (position and tilt) as measured at point 21, and the stand-in's eyes are rotated so that they have the pupil positions measured at point 21.
[0090] Then, at position 23, Figure 2The method involves determining the viewpoint based on a stand-in positioned at the measured head and pupil locations. Therefore, unlike some traditional methods, instead of using direct measurements, a stand-in set to the corresponding position is used as the basis for the final determination. As mentioned above, this can be done using a real eyeglass frame or a virtual fitted eyeglass frame. (See reference...) Figures 5 to 7 Explain this determination.
[0091] As indicated by box 24, repeat steps 21 to 23 for different target locations, for example by making the mobile device like Figure 4 As indicated by arrow 43, it moves within a person's field of vision. Through repetition, this can be achieved using a smartphone or similar mobile device. Figure 8 and Figure 9 A similar viewpoint diagram. This diagram can be displayed or transmitted to the lens manufacturer, and then adapted to the viewpoint of the corresponding person.
[0092] To determine the viewpoint at 23 locations, the mobile device 10, the avatar of the head 30, the avatar of the eyes 31, and the eyeglass frame 40 (as a virtual frame or as a detected real frame) can each be represented by a six-dimensional vector, which includes at least three-dimensional position (x, y, z coordinates) and three angles (α, β, γ). That is, the mobile device can be represented by a position {x...} 移动 ,y 移动 ,z 移动} and orientation {α 移动 ,β 移动 ,γ 移动} indicates that head and / or facial features can be determined by location {x} 头部 ,y 头部 ,z 头部} and orientation {α 头部 ,β 头部 ,γ 头部} indicates that at least one eye can be determined by position {x}. 眼睛 ,y 眼睛 ,z 眼睛} and orientation {α 眼睛 ,β 眼睛 ,γ 眼睛 The orientation can be represented in Cartesian space, polar space, or a similar space. Orientation can be represented using Euler angles, rotation matrices, quaternions, or similar methods. Orientation is measured by orientation sensor 110, and the position of mobile device 10 can be set to 0 to determine other positions as explained below relative to mobile device 10.
[0093] By measuring, for example, several landmarks at the head (30°), the position of a person's head can be determined. 头部 ,y 头部 ,z 头部} and the corresponding orientation {α 头部 ,β 头部 ,γ 头部}. The position and orientation of the eye within the Stand corresponding to eye 31 can be determined as {x}. 眼睛 ,y 眼睛 ,z 眼睛} and the orientation can be determined as {α} 眼睛 ,β 眼睛 ,γ 眼睛 Then set the stand-in to the corresponding position and orientation.
[0094] Using the mobile device 10, the center of rotation of the eye can also be determined, as described, for example, in European patent application EP20202529.2, and the position and orientation of the eye can be determined based on the pupil position and the center of rotation. Furthermore, similarly for eyeglass frames, the corresponding position {x} can be determined. 镜架 ,y 镜架 ,z 镜架} and orientation {α 镜架 ,β 镜架 ,γ 镜架 As explained above.
[0095] exist Figures 5 to 7 In the figure, reference numeral 50 denotes a stand-in for the head 30, set to the position and orientation measured in step 21; reference numerals 53A and 53B specify the left and right eyes of the stand-in, respectively; and reference numeral 51 specifies the position of the mobile device 10, i.e., essentially the measurement plane. Position 51 is set to a vertical position for calculation. This means that the orientation of the eyes, head, and glasses is transformed to a coordinate system in which the mobile device 10 is in the xz plane based on the measurement orientation of the mobile device 10. Figure 5 In the diagram, 54A and 54B represent the lines of sight of eyes 53A and 53B, respectively, and... Figure 6 In the diagram, lines 64A and 64B indicate the corresponding lines of sight. When using at least one camera and / or at least one depth sensor to record the relative changes in the position and orientation of the double, no additional equipment is required to determine the viewpoint.
[0096] Figures 5 to 7 Showing heads wrapped around Figure 3 and Figure 4 The x-axis rotation, or in other words, represents the yz plane. Similar calculations can be performed on other planes.
[0097] exist Figure 5 In the middle, the eyes are basically looking straight ahead, so that the angles α corresponding to the left and right eyes are respectively... 眼睛,OS and α 眼睛,OD (As shown at point 51) corresponds to the head rotation angle α 头部 .exist Figure 6In the middle, the gaze of both eyes converges on a single point, making the angle α 眼睛 OS and α 眼睛 OD and angle α 头部 The difference lies in the relative eye rotation angle α between the left and right eyes relative to the head. 眼睛 . Figure 6 Δα in the figure specifies the eye rotation and the tilt of the eyeglass frame at the corresponding eye. 镜架 The difference between them.
[0098] exist Figure 5 In the attached diagram, reference numerals 55A and 55B specify the viewpoints of the left and right eyes, respectively. Figure 6 In the diagram, 65A and 65B specify the viewpoints for the left and right eyes, respectively. Figure 7 The diagram again illustrates how the viewpoint is constructed based on the angle in question for the right eye. Essentially, the difference between the eyeglass rotation angle and the eye rotation is estimated using Δα, and then the viewpoint 70 is calculated based on the position of the eyeglass frame 52 (e.g., by means of the coordinates given by the index frame mentioned above).
Claims
1. A computer-implemented method for determining a person's point of view (55A, 55B; 65A, 65B; 70) from a mobile device (10) comprising at least one sensor (11, 12), the at least one sensor comprising an image sensor (11), wherein, The eye point (55A, 55B; 65A, 65B; 70) is the intersection of the person's line of sight and a plane approximating the spectacle lens position, the method comprising: determining, using the at least one sensor (11, 12), a head position and a pupil position of the person when the person is looking at a target, and determining the eye point (55A, 55B; 65A, 65B; 70) based on an avatar (50) that is a 3D model of at least an eye portion of the person's head, the avatar being set to the determined head and pupil positions based on set head and pupil positions of the avatar (50) and based on a position of a plane approximating the spectacle lens position, wherein: - the method further comprises virtually fitting a spectacle frame to the avatar (50), and determining the eye point (55A, 55B; 65A, 65B; 70) based on the virtually fitted spectacle frame; or - the person wears a spectacle frame, and the method further comprises: identifying the spectacle frame worn by the person, wherein determining the eye point (55A, 55B; 65A, 65B; 70) is further based on the identified spectacle frame, characterized in that, determining the head position comprises determining a portion of the person's head or facial landmarks, and the method further comprises moving the target, wherein determining the eye point (55A, 55B; 65A, 65B; 70) is performed for a plurality of target positions, and the target is the mobile device. Determining the head position comprises determining at least one position selected from the group consisting of the head position and the facial landmarks with respect to the target.
2. The method of claim 1, wherein, The at least one sensor further comprises a depth sensor (12).
3. The method of claim 1 or 2, wherein, Determining the head position and the pupil position is performed by using the depth sensor.
4. The method of claim 3, wherein, 5. The method of claim 4, characterized in that, determining, using the depth sensor, respective poses of the mobile device with respect to the head at the plurality of target positions, and determining, for respective target positions of the plurality of target positions, respective eye points based on the respective poses. generating the avatar (50) with the mobile device, wherein generating the avatar (50) comprises:
6. The method of claim 1 or 2, wherein, providing a plurality of combined 2D images and depth images of the person's head captured from different positions, determining landmark points for each of the combined 2D images and depth images, and performing pairwise coarse registration between pairs of combined 2D images and depth images based on the landmark points, after the pairwise coarse registration, performing fine registration of the coarsely registered combined 2D images and depth images based on a complete point cloud represented by the combined 2D images and depth images, and generating the avatar (50) based on the registration. Providing the plurality of combined 2D images and depth images of the person's head comprises measuring the person's head in a plurality of positions by the at least one sensor.
7. The method of claim 6, wherein, Measuring the person's head in a plurality of positions comprises measuring the person's head in a plurality of upright positions of the person's head rotated around an axis of the head.
8. The method of claim 7, wherein, The target is provided in a near field.
9. The method of claim 1 or 2, wherein, 10. A mobile device (10) comprising: at least one sensor (11, 12) comprising an image sensor (11) and configured to determine a head position and a pupil position of a person when the person is looking at a target, and a processor configured to determine a viewpoint (55A, 55B; 65A, 65B; 70) of the person based on an avatar (50) of the person, the avatar being a 3D model of at least an eye part of the person, and being set to the determined head and pupil position based on a set head and pupil position of the avatar (50) and based on a position of a plane approximating an eyeglass lens position, wherein the viewpoint (55A, 55B; 65A, 65B; 70) is an intersection of a line of sight of the person and the plane approximating the eyeglass lens position, wherein: - the processor is further configured to virtually fit a spectacle frame to the avatar (50), and to determine the viewpoint (55A, 55B; 65A, 65B; 70) based on the virtually fitted spectacle frame; or - the person wears a spectacle frame, and the processor is further configured to recognize the spectacle frame worn by the person, wherein determining the viewpoint (55A, 55B; 65A, 65B; 70) is further based on the recognized spectacle frame, characterized in that determining the head position comprises determining a part of the head or a facial landmark of the person, and the mobile device (10) is further configured to determine the viewpoint of the person for a plurality of positions when moving the target in the field of view of the person, and the target is the mobile device.
11. The mobile device (10) of claim 10, characterized by The mobile device is configured to perform the method of any one of claims 1 to 9.
12. A computer program product comprising a computer program for a mobile device comprising at least one sensor (11, 12) and a processor, characterized in that, The computer program, when executed on the processor, causes performing the method of any one of claims 1 to 9.
13. A method for producing an eyeglass lens, characterized in that: a viewpoint (55A, 55B; 65A, 65B; 70) of a person is determined according to the method of any one of claims 1 to 9, and the eyeglass lens is produced based on the determined viewpoint (55A, 55B; 65A, 65B; 70).
Citation Information
Patent Citations
Method and device for determining the visual behaviour of a person and method of customising a spectacle lens
EP1747750A1
Method for designing spectacle lenses taking into account an individual's head and eye movement
EP1815289B9
Method for determining a progressive ophthalmic lens
EP1960826B1
Method of determining at least one parameter of visual behaviour of an individual
EP3145386B1
Method, device and computer program for determining a close-up viewpoint
EP3413122A1