Method for recording a head image and corresponding mobile device
By capturing head images on mobile devices and utilizing machine learning to identify predefined markers, the problem of insufficient centering features in image capture on mobile devices is solved, achieving efficient and adaptable head image capture suitable for personal use.
Patent Information
- Application Number
- CN202380021625.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-14
- Filing Date
- 2023-02-13
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Existing technologies struggle to ensure the necessary centering features are present in head images captured using mobile devices, especially when glasses are worn, and are computationally demanding, making them unsuitable for personal use.
By capturing multiple images of the head on a mobile device and searching for predefined markers in the images, machine learning is used to identify these markers to control image capture. Capture stops once a predefined subset of the markers is identified, reducing computational requirements.
It enables efficient capture of images containing the necessary centering features on mobile devices, reduces computational requirements, is suitable for personal use, and is adapted to situations where people wear glasses.
Smart Images

Figure CN118715476B_ABST
Abstract
Description
[0001] This application relates to a method for recording head images using a mobile device, which can then be used to generate head models, determine centering parameters, or both. Furthermore, this application relates to a corresponding device.
[0002] A 3D model of the human head (hereinafter referred to as the head model) can be used to represent a human head as a so-called surrogate in various applications. An example application is the virtual fitting and try-on of eyeglass frames, as described, for example, in WO 2019 / 008087A1. For such applications, it is desirable that the head model used matches the actual human head to give an accurate visual impression of how the eyeglass frames would look on a person.
[0003] Another application is determining centering parameters. Centering parameters are the parameters required to properly fit a lens into an eyeglass frame (the so-called centering process) so that the lens is worn in the correct position relative to the human eye. Examples of such centering parameters include pupillary distance, centering distance, vertex distance, fit point position, or gloss over time. These and other centering parameters are defined in Section 5 of DIN EN ISO 13666-2012 and are used herein in the sense defined according to this standard.
[0004] In many cases today, corresponding systems are used to determine such centering parameters automatically or semi-automatically. An example of a centering system that can also generate head models is the Zeiss VISUFIT 1000 centering system.
[0005] To determine the centering parameters, the VISUFIT 1000 system uses the method described in EP 3 363 346 B1. Here, an arrangement of nine cameras with fixed spatial relationships (i.e., positioned in fixed locations within the device) is used, such that the camera positions and orientations are known in advance. This combination of positions and orientations is also referred to herein as attitude, as defined in DIN EN-FR ISO 8373-2012-03. In this way, since the relative attitudes of the cameras are known, a head model can be generated based on techniques similar to triangulation. For centering, the described frontal and side images are recorded using the camera arrangement. The pupil center in the frontal image and the apex in the side image are detected using these images, and the 3D position of the apex of the cornea is calculated using an adapted triangulation method. As further described in US10,942,375B2, an illumination unit can be used to ensure good lighting for the eyes and visibility from all viewing angles. US2020 / 0057316 A1 discloses a similar method with a slightly different representation of input data, which also includes a 3D model of the eyeglasses frame, much like in the case of so-called virtual centering, where the person is not wearing the actual eyeglasses frame, but the frame is virtually fitted and adapted to the person. While these methods work well, they require specific equipment with a fixed camera setup, which may be available, for example, at an optician or doctor's office, but is virtually infeasible for personal, private use.
[0006] With the increasing processing power and image capture capabilities of mobile devices such as smartphones and tablets, various methods have been adopted to generate head models and perform centering using such mobile devices. Typically, in such methods, the mobile device's camera captures a person's head, or at least its eye portion (i.e., including part of the eyes), from different directions. The camera may include a depth sensor, enabling the provision of so-called RGBD (red, green, blue, depth) images. EP 3 913 424 A1 discloses a method for determining some centering parameters using a mobile device, wherein the head is captured from different positions on the mobile device and an accelerometer is used to measure the acceleration of the mobile device while it is moved between these positions.
[0007] Compared to fixed camera setups, there are two general problems when using mobile devices. Firstly, the relative camera pose of the captured image is not known a priori. Several methods exist for addressing this problem.
[0008] For example, Tanskanen, Petri et al., “Life metric 3D reconstruction on mobilephones”, Proceedings of the IEEE International Conference on Computer Vision, 2013, disclosed a life metric 3D reconstruction, such as that of a statue in a museum, in which accelerometers and / or gyroscopes set in the mobile phone use inertial tracking to perform attitude estimation of the mobile phone.
[0009] Kolev, Kalin et al., “Turning mobile phones into 3D scanners”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, using a camera motion tracking system to provide camera pose.
[0010] Muratov, Oleg, et al., “3D Capture: 3D Reconstruction for a Smartphone”, also used an inertial measurement unit to track camera pose.
[0011] These methods rely on the accuracy of accelerometers or similar components to provide camera pose, from which 3D models can be calculated.
[0012] In another approach, WO 2019 / 164502 A1 discloses using a smartphone for image capture and a 3D mesh representation to generate a head model. US2020 / 0 125 835A1 describes generating a 3D model of the head based on captured images or videos. Machine learning networks are used to track pose and facial expressions and refine the model.
[0013] A second problem when using a mobile device to capture images is ensuring that the captured image actually includes the necessary portions of the head as viewed from the desired direction, such as frontal and side views. In contrast, utilizing the fixed camera arrangement described above, the camera arrangement itself ensures that both frontal and side views of the head are captured. This application primarily addresses this second problem.
[0014] US10,755,438U1 and US10,157,477B1 disclose a head model capture process in which a person is instructed to turn their head in front of a camera on a mobile device until a certain angle is automatically determined by software running on the mobile device. The rotation angle is determined by matching a 3D mesh model obtained from each head rotation pose with an initial mesh model of the initial pose and calculating a 3D transformation that maps the mesh of the initial pose to the mesh of each subsequent pose. Based on this transformation, the rotation angle can be calculated, and when a threshold is exceeded, the correspondingly programmed mobile device issues a corresponding command to the person via the mobile device's display or via output audio information. Rotation to the left and right can be performed.
[0015] This method requires continuous matching of 3D meshes during the capture process. This can be computationally intensive, potentially limiting the method to mobile devices with very high processing power. Furthermore, due to the diversity of heads, hairstyles, etc., this method does not necessarily guarantee the reliable capture of all head features necessary for centering, making accurate centering difficult. Finally, this method is unsuitable for centering real eyeglasses when the person is wearing them and the centering parameters of those real eyeglasses are yet to be determined. Specifically, in such cases, the eyeglasses may obscure the cornea in some images, depending on the type of eyeglasses worn, particularly the base curve, bezel angle, or both. Therefore, in this situation, using a certain angle threshold for image capture may result in captured images that do not contain all the information necessary for centering.
[0016] This invention solves the problem of image capture and ensuring that relevant information is present in the captured image.
[0017] According to the first aspect, a method for capturing head images is provided, comprising:
[0018] Provide a first image of at least the eyes of the head to the mobile device.
[0019] Provide the mobile device with multiple second images of the head viewed from multiple directions.
[0020] Its features are,
[0021] The mobile device searches for predefined flags in each of these second images, and
[0022] In response to the mobile device identifying at least a predefined subset of these predefined flags in the second image of the plurality of second images, the provision of the plurality of second images to the mobile device is stopped.
[0023] Providing the first image and the second images to the mobile device may include providing (e.g., capturing) the first image and the second images by the mobile device's camera. In other embodiments, an external camera may be used, which may be connected to the mobile device wirelessly (e.g., via Bluetooth or WLAN) or via a wired connection (e.g., via USB, LAN, etc.). In this and the following aspects, the capture may be performed by the person to whom the head belongs, either by herself or by another person.
[0024] According to the second aspect, a method for capturing head images is provided, comprising:
[0025] Capture the first image of at least the eye portion of a person's head using a mobile device, and
[0026] While the head rotates relative to the mobile device in a first direction, a plurality of second images of the head are captured using the mobile device. The method is characterized by searching for predefined markers in each of these second images using the mobile device, and stopping the capture of these second images in response to identifying at least a predefined subset of these predefined markers in the second images of the plurality of second images.
[0027] The second aspect can be alternatively implemented as a method for head image capture, comprising:
[0028] Capture the first image of at least the eye portion of a person's head using a mobile device, and
[0029] While capturing multiple second images of the head using the mobile device, the head is rotated relative to the mobile device in a first direction. In this case, the method is characterized by searching for predefined markers in each of these second images using the mobile device, and stopping the capture of these second images using the mobile device in response to identifying at least a predefined subset of these predefined markers in the second images of the plurality of second images.
[0030] Therefore, for both the first and second aspects, contrary to the conventional methods discussed above in US10,755,438U1 and US10,157,477B1, the capture of the second image is not stopped when a predefined rotation angle is reached, but rather based on the presence of a marker in the captured second image. This ensures that the marker required for specific purposes such as centering is actually visible in the second image. Furthermore, marker detection can be performed with less computational power compared to generating a complete mesh as in the prior art.
[0031] The terms used above to define the methods and further characteristics will now be explained. These explanations apply to both the first and second aspects.
[0032] Mobile devices are devices designed to be carried by a person. Typically, mobile devices weigh less than 1 kg. Mobile devices, as used herein, include at least a processor for performing tasks such as searching and stopping, and for overall control of the mobile device, as well as an image capture device for capturing a first image and a second image. Typical examples of such mobile devices include smartphones and tablet PCs.
[0033] The term "image" can refer to a 2D color image, a 2D black-and-white image, a depth image captured by a depth sensor such as a time-of-flight (TOF) sensor, or a combination thereof. In a preferred embodiment, the image is a combination of a 2D color image and a depth image, and is also referred to as an RGBD image. Some modern mobile devices, such as smartphones or tablet PCs, are already equipped with such image capture devices, which include both an RGB digital camera and a depth sensor.
[0034] "At least the eyes" means that at least the human eyes are visible in the first image. The first image is preferably a frontal image, wherein the head is captured from the front. A frontal image is an image captured from the front and showing at least the human eyes.
[0035] Rotating the head relative to the mobile device can be performed by keeping the mobile device stationary and rotating the head, or by keeping the head stationary and rotating the mobile device about the head. The term "rotating the head relative to the mobile device" includes two alternatives. A first direction can be either left or right, such that the rotation is vertical about the head. In other embodiments, the rotation can be horizontal about the head.
[0036] A landmark is a specific point or area on the head. Generally, such landmarks may include, for example, the tip of the nose, a point on the bridge of the nose, the corner of the mouth or corner of the eye, the pupil, the cornea, or a point or feature on the ear. As will be explained below, in certain embodiments, landmarks on the eyes, such as the cornea and pupil, and landmarks on the ears may be used. Various landmarks associated with the ear include, but are not limited to, the earlobe, the antitragus, the tragus, the root of the helix, and the highest point of the helix (upper part of the helix). These can be used as landmark points.
[0037] Searching for and identifying predefined markers can be accomplished through various conventional methods. For example, trained machine learning logic, such as neural networks, can be used to search for and determine marker points. In this case, for training, multiple images from different locations and for multiple different heads are used as training data, where predefined markers can be manually labeled. After training, the trained machine learning logic then determines the marker points. Details can be found in, for example, the following literature: Y. Wu et al., Facial landmark detection with tweaked convolutional neural networks, IEEE transactions on pattern analysis and machine intelligence, 40(12), 3067-3074, 2017; P. Perakis et al., 3D facial landmark detection under large yaw and expression variations, IEEE transactions on pattern analysis and machine intelligence, 35(7), 1552-1564, 2012; or Y. Wu et al., Facial landmark detection: A literature survey. International Journal of Computer Vision, 127(2), 115-142 (2018).
[0038] Another algorithm using convolutional neural networks to predict the location of facial landmarks in an image is described in X. Guo et al., PFLD: A practical: facial landmark detector, arXiv preprint, 2019. arXiv:1902.10859. Another method based on regression trees is disclosed in V. Kazemi et al., One millisecond face alignment with an ensemble of regression trees, IEEE conference on Computer Vision and Pattern Recognition (CVPR), 1867-1874 (2014). This method is also implemented in the software library "dlib".
[0039] Another method is described in X. Zhu et al., Face detection, pose estimation and landmark localization in the wild, IEEE Conference on Computer Vision and Pattern Recognition, 2879-2886 (2012). Further methods using, for example, deep learning and convolutional neural networks are discussed in G. Amato et al., “A Comparison of Face Verification with Facial Landmarks and Deep Features”, conference paper MMEDIA 2018, The Tenth International Conference on Advances in Multimedia, Athens, Greece, April 2018. Another approach using a trained linear regressor was disclosed in XPurgos-Artizzu et al., “Robust face landmark estimation under occlusion”, 2013 IEEE International Conference on Computer Vision. Other methods do not use trained machine learning logic, but instead use conventional image analysis to detect landmarks, but this typically requires more effort than recent methods using machine learning.
[0040] Therefore, the various possibilities for searching and identifying the mark are known in the art.
[0041] In some cases, for feature detectors using statistical models or neural networks, it may be necessary to search for multiple landmarks, such as on the contours of an ear or another facial part, on which the landmarks are located. An example of a corresponding face alignment method using neural networks is described in Kowalski M., Naruniec J., and Trzcinski T., 2017, Deep alignment network: A convolutional neural network for robust face alignment, in Proceedings of the IEEE conference on computer vision and pattern recognition workshops (pp. 88-97). An example of a face alignment method based on statistical models is described here: Kazemi, Vahid & Sullivan, Josephine (2014), One Millisecond FaceAlignment with an Ensemble of Regression Trees, 10.13140 / 2.1.1212.2243. Most known methods are not explicitly trained on datasets containing feature points of the ears and / or eyes. Therefore, in such embodiments, these methods can be freshly trained using images or models manually labeled with the symbols used herein (particularly symbols related to ears and eyes) for the purposes described herein, namely, searching for and recognizing corresponding symbols.
[0042] An example of a curve rather than a point is the frontal contour of the cornea, which can be represented as an arc with a preset radius. The radius can be determined starting from an average radius of 8 millimeters based on the curvature of the human cornea and can be converted to a pixel radius, which can be calculated based on R_pixels = R_cornea_real / dist_eye * f_pixel. R_pixel is the radius of the cornea in pixels in a specific image, R_cornea_real is the approximate true radius of the cornea, for example, 8 millimeters, dist_eye is the distance of the eye from the optical center of the camera, as determined, for example, using the camera's depth sensor mentioned above, and f_pixel is the camera's focal length in pixels, which is part of the camera calibration, i.e., known for a particular camera. The conversion between the real object and the pixel-scale of the object in the image based on camera calibration data is known in itself; see, for example, camera calibration and 3D reconstruction https: / / docsopencv.org / 2.4 / modules / calib3d / doc / camera_calibration_and_3d_reconstruction.html Richard Hartley and Andrew Zisserman, “Multiple View Geometry in Computer Vision”, Cambridge University Press, 2004, Section 1.2 “Camera Projections”.
[0043] "Predefined subset" means that when performing a search of all predefined flags, only a subset needs to be identified to stop. For example, as described later, multiple ear points can be searched, and image capture can stop if most of these points have been identified. More than one predefined subset can be used.
[0044] The capture, search, and stop described above are performed by the mobile device itself, while the positioning of the head relative to the mobile device is performed by a person or another person holding the mobile device, including turning the head relative to the mobile device in a first direction. For the sake of brevity, in the following explanation, it will be assumed that the person performs this step and similar steps described below by himself or her, but it should be understood that another person may at least assist this person in performing the corresponding steps.
[0045] To facilitate human actions such as turning, mobile devices can output corresponding instructions. For example, instructions can be output as lengthy audio commands or through certain audible signals, and / or can be executed by displaying visual instructions or prompts on the mobile device's display. Typical mobile devices, such as smartphones or tablet PCs, usually include multiple speakers and a display for outputting this information.
[0046] For example, in one embodiment, the method may further include outputting an instruction to stop moving the head relative to the mobile device in a first direction in response to recognizing at least a predefined subset (i.e., when capture stops). In this way, the person is notified to stop rotating, and the capture of the second image is complete.
[0047] Choosing a predefined subset means that when the predefined subset is identified in a second image, it implies that an appropriate second image has been captured for later use. For example, the predefined subset could indicate that the image is a side view, where the predefined subset is visible.
[0048] After the rotation stops, the method may further include: capturing a third image while rotating the head relative to the mobile device in a second direction opposite to the first direction, which corresponds to capturing a third image of the head with the mobile device while rotating the head relative to the mobile device in a second direction opposite to the first direction.
[0049] Search for additional predefined markers in each of the third images; and
[0050] The capture of the third images is stopped in response to the identification of at least another predefined subset of these additional predefined markers in the third images of the plurality of third images.
[0051] In this way, for example, both sides of the head can be captured if the first direction corresponds to, for example, one of the directions left or right, and the second direction corresponds to the opposite direction of the first direction, i.e., the other of the directions left or right. Otherwise, the mechanism here is the same as the mechanism explained above for the first direction. Additional predefined markers can be substantially the same as the markers defined above, except for the other side of the head, and similarly, another predefined subset can be substantially the same as that predefined subset, except for the other side of the head. In this way, for example, a side view image from the first side of the head can be captured in a second image, and a side view image from the other side of the head can be captured in a third image.
[0052] Also here, an instruction can be output to stop turning the head in the second direction in response to the recognition of another predefined subset.
[0053] In other embodiments, capturing a third image may be omitted. For example, as an alternative, an image of only one side of the head may be captured, and for the other side of the head, symmetry may be assumed as an approximation. This is an example approximation method because the human head is not typically perfectly symmetrical, but it speeds up the capture process on the other hand.
[0054] The first image can be a frontal image. This means, for example, capturing both eyes of a person in the first image. Predefined subsets and (if used) additional predefined subsets can then be used to ensure the capture of side images, as mentioned above.
[0055] To facilitate the capture of frontal images, corresponding instructions can be output to the person.
[0056] For example, during the image capture process, a so-called front-facing camera (a camera on the same side as the mobile device's display) can typically be used. The image captured by the front-facing camera can be displayed on the mobile device's display using a mirror function. To capture a frontal image, a schematic facial contour can then be displayed over the image captured by the front-facing camera, and the person can be instructed to hold the mobile device so that their face is within the facial contour shown in the image. Additionally, further landmarks, such as the two eyes, mouth, and nose, can be detected to verify that a frontal image has been captured.
[0057] Predefined landmarks and (if used) additional predefined landmarks may include at least one landmark selected from the group consisting of points or features on the pupil, corneal apex, corneal contour, and ear. Preferably, both the pupil and corneal apex or corneal contour may be used, and multiple points or other features on the ear may be used. By using the pupil and corneal apex or corneal contour, it can be ensured that the features necessary for determining the centering parameters are visible in at least a second or third image in which these landmarks are identified. By using points or features on the ear, the side view can be ensured. Furthermore, for some virtual fitting procedures, the ear is further used, as explained below.
[0058] For example, a predefined subset or another predefined subset may then include a larger portion of the ear points or features from the predefined marker points, which ensures that at least the largest portion of the ear is visible.
[0059] Once the first, second, and optionally third images have been captured, these images are processed in the embodiments. This image processing can be performed by the mobile device itself or at another device. For example, the captured images can be sent to another device via a network such as the Internet, and optionally, the results can then be received from that other device at the mobile device. For example, the other device could be one with higher computing power. For instance, since no computational grid is required during image capture, image capture itself can be performed with lower computing power compared to existing methods using grids discussed above, and the images can then be transferred to a device with higher computing power (if the mobile device's computing power is insufficient).
[0060] In some embodiments, processing the captured image may include determining at least one geometric parameter based on the image. Geometric parameters may be, for example, distances between landmarks on the head, such as pupillary distance, head width, etc. In other embodiments, geometric parameters may include distances between landmarks on the head and landmarks on the eyeglass frame, such as the distance between the cornea and the position of the eyeglasses on the frame. The case involving eyeglass frames is described in more detail below. At least one geometric parameter may be a centering parameter. In this way, various geometric parameters can be determined, and the image capture process described above can ensure that landmarks necessary for determining at least one geometric parameter are visible, for example, in a second image or optionally also in the third image described above, or that the image is captured from a direction necessary for determining the geometric parameter. In one embodiment, the image capture process ensures that a side view image is captured.
[0061] During the image capture process, a person can be directed to look at a target. This target can be a part of the mobile device, such as the mobile device's camera or a corner of the mobile device, or it can be a target displayed on the mobile device's screen. If an image of a person is displayed during image capture, the target can also be the person's eyes as shown.
[0062] Specifically, when a person holds the mobile device, the device is relatively close to the client's eyes, causing the gaze of both eyes to converge to the presented target during image acquisition. In this case, convergence correction can be applied to correct this convergence. In the above case, the position of the target relative to the camera is known, allowing convergence correction to be performed by transforming the 3D eye point via one or more rotations around the eye's center of rotation. The location center is known based on the atomic model of the eye, for example, its corresponding position relative to the apex of the cornea, which can be detected as a marker as mentioned above. Therefore, the relative position of the eye point can be adapted to match the position of a person focusing on a gaze target at infinity in the direction known as the primary gaze direction or zero gaze direction.
[0063] In other embodiments, a person may be instructed to look at infinity, such as a target on a wall, essentially in the main gaze direction mentioned above, and another person may move the mobile device around the head. In such cases, convergence / divergence correction is unnecessary.
[0064] In some embodiments, image processing may include generating a head model or at least a model of the eye region based on the image. This can be performed in any conventional manner, such as as described in the documents referenced in the Background section.
[0065] In this head model, the aforementioned landmarks can be detected. One landmark necessary for determining certain centering parameters is the position of the cornea apex in three dimensions, referred to below as the 3D vertex position. One method for obtaining this applicable 3D vertex position is described in EP 3 363 346 B1 mentioned above. When using this method, it is necessary to obtain a side view of the human head, showing the cornea from the side. This can be ensured by the method described above based on a subset of predefined landmarks that must be detected in the second image. Alternatively, the reconstructed 3D model of the head, and here its eye portion, can be used directly to determine the 3D vertex display.
[0066] Because the human cornea is transparent, it can appear flat in a head model in some methods, such as those using image-based pattern projection and triangulation. This can be mitigated by a heuristic approach, for example, applying an offset of approximately 3 millimeters to the area surrounding the center of the pupil. In this case, when capturing a second or third image, and specifically a lateral image, the person can be instructed to focus on any point at a distance, giving the person a general primary gaze direction.
[0067] Therefore, different methods can be used to determine the position of 3D vertices.
[0068] As mentioned above, some geometric parameters, especially centering parameters, can include the distance between the markings on the head and the markings on the eyeglasses frame. Typically, two methods can be used to determine the position of the markings on the eyeglasses frame.
[0069] In the first method, the procedure involves virtually fitting a model of the eyeglasses frame to a head model. Such virtual fitting is sometimes referred to as virtual fitting and is described, for example, in EP 3 410 178A1 or EP 3 425 447 A1. Based on markers on the fitted model of the eyeglasses frame and markers on the head model, geometric parameters such as centering parameters can then be determined. In EP 3 410 178A1, multiple images of a person's head are captured from different directions. Another virtual fitting method is disclosed in EP3 631 570 B1. In the case of EP 3 649 505 B1, the parameterized frame model can also be fitted to the generated head model.
[0070] In alternative methods, a person is wearing eyeglasses when the first, second, and optionally third images are captured. The eyeglasses are then identified in the images, and geometric parameters can be determined based on the identified eyeglasses. For example, in this case, the head model mentioned above could include a model of the eyeglasses generated based on the images (the first, second, and optionally third images), and geometric parameters such as centering parameters can be determined. In another method, a generic eyeglasses model can be fitted to the eyeglasses identified in the images, and the centering parameters can be determined based on the fitted model, as described, for example, in WO 2018 / 138258 A1. Other methods for obtaining the centering parameters of a solid-frame eyeglasses frame worn by a person can be based on the methods disclosed in EP 3 363 346 B1 or EP 3 574 370 B1.
[0071] If the first method described above, i.e., virtual try-on, is to be used, but the person is wearing eyeglasses, the mobile device can issue a command to the person to remove the eyeglasses. If the person needs relatively strong corrective lenses, then in the case where the person is not wearing eyeglasses (or other vision assistive devices, such as contact lenses), the aforementioned command to perform the method can be issued as a precise audio command, because the person may not be able to recognize visual commands on the mobile device's display.
[0072] In another approach, when a person is wearing eyeglasses, the frames can be virtually removed from the image and thus omitted from the generated head model. Another use case for this functionality is a side-by-side comparison of a person's real eyeglasses with virtual frames without needing to capture a sequence of images. Automated removal can be achieved by training machine learning models, such as neural networks, to remove frames from images. As training material, pairs of real images of people with or without eyeglasses can be used. Alternatively, or alternatively, artificial training material can be generated by virtually trying on eyeglasses and generating rendered image pairs, which has the advantage that people will exhibit identical facial expressions in both images (with or without eyeglasses), thus tending to produce fewer artifacts during training. Another method for removing eyeglasses from images is disclosed in Wu Chenyu et al., "Automatic eyeglasses removal from face image," IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Computer Society, Vol. 26, No. 3, 2004, 322-366.
[0073] The centering parameters thus determined can then be transmitted to the lens manufacturer to manufacture lenses accordingly for the corresponding eyeglass frames.
[0074] As mentioned above, during the image capture process, the captured current image can be displayed on the device's screen in a mirrored manner. In other cases, when the mobile device is appropriately equipped and has sufficient computing power, a mesh or other model of the head can be displayed continuously. For example, some modern smartphones or tablets are equipped with RGB cameras as mentioned above and have corresponding libraries for generating 3D representations. ARKit for iOS devices or ARCore in the case of Android devices provides such a possibility. In this case, an image based on this 3D representation (e.g., a so-called avatar) can be displayed instead of a mirrored image. For example, in the case of some iOS devices, such a method can be applied to infrared mode projection.
[0075] In another aspect of the invention, a corresponding mobile device is provided, comprising at least a processor for implementing the method according to the second aspect, and optionally (if the mobile device captures an image) also for the first aspect, and a camera device for capturing images. The processor is configured to substantially control the mobile device to perform the method described above, i.e., control the camera for a corresponding display used for image capture, analyze the image to identify markers as explained above, and accordingly stop image capture. The mobile device may further include a display, a loudspeaker, or any other input / output device to issue corresponding instructions to a person for these steps of the method, depending on the person's actions, such as turning their head relative to the mobile device. It should be noted that, based on markers, the mobile device can check whether a person has correctly turned their head, for example, based on the visibility of eye markers such as pupils and ear markers that change in the image. For example, when turning the head from a frontal image, one of the eyes disappears during the turn, and one of the ears becomes increasingly visible. If such a change in markers is not detected, an error message can be output. In other methods, such error detection is not performed, and unwanted images are captured when the user operates the device incorrectly (e.g., without turning their head), and the results of the method (such as the calculation of centering parameters) will be incorrect.
[0076] To compensate for the risk of incorrect centering parameters, in some embodiments, the user interface of the mobile device can be configured (e.g., programmed) to present a final check to a person or another person before the centering parameters are actually used for centering. Given a reconstructed 3D corneal vertex of the eye in a head model, these 3D points can be projected back into one or more of the captured images, allowing the person or another person to check for correct positioning. Camera calibration data can be used to map the 3D points into pixels of the image, as described in the publication cited above by Richard Hartley and Andrew Zisserman. In a similar manner, the mobile device, or other devices to which the data is transmitted, can present the person or another person with a final check screen that visualizes additional markers and 3D points projected back into a 2D image. Another method of presenting centering parameters is to draw a technical sketch of the frame including the centering point and distance. The technical sketch can be an orthographic projection of the 3D model into a 2D plane and includes the possibility of showing this reality to scale on the display of the mobile device.
[0077] In the case of, for example, smartphones or tablet PCs, a corresponding computer program (often referred to as an "app" in the case of smartphones and tablet PCs) can be provided to program the mobile device accordingly to perform any of the methods described above, for example, to cause the mobile device's processor to control the mobile device to perform the method. Such a computer program can be provided on tangible storage media such as hard drives, CDs, DVDs, memory modules, or memory (e.g., the memory of a mobile device) or transmitted as a data carrier signal.
[0078] Embodiments will now be described with reference to the accompanying drawings.
[0079] Figure 1 This is a block diagram of the device according to an embodiment.
[0080] Figure 2 This is a flowchart illustrating a method according to an embodiment.
[0081] Figures 3A to 3C It is a demonstration Figure 2 A simplified diagram of the method.
[0082] Figure 4 It is a simplified diagram showing the symbol on a person's head.
[0083] Figures 5A to 5C It is a simplified diagram showing the markers detected on the eyes and ears for various angles of head rotation.
[0084] Figure 6 This is a flowchart illustrating a method according to an embodiment.
[0085] In the following sections, embodiments relating to the use of mobile devices such as smartphones or tablet PCs to capture images (which can then be used to generate 3D models of a head or parts thereof to determine centering parameters) will be described. Figure 1 A block diagram of a device 10 that may be used in some embodiments is shown. Device 10 may be a smartphone or tablet PC, but may also be a dedicated device. Device 10 is accordingly programmed to perform methods as discussed herein.
[0086] Device 10 includes a camera 12 and a depth sensor 11. The depth sensor 11 and camera 12 form an RGBD camera as discussed above and are an example of a camera device. In other embodiments, the depth sensor 11 may be omitted. Furthermore, device 10 includes a touchscreen 13, a processor 15, a memory 14, and an input / output interface 16. The touchscreen 13 can be used to control device 10 and serves as an example of a display; it can also be used to output instructions to a person, such as instructions to capture an image, as discussed above and further below. Figure 2As explained above. Processor 15 executes instructions stored in memory 14 to implement the methods discussed above and below. Input / output interface 16 can provide communication with networks such as the Internet, for example, to transmit captured images to another device without performing further processing on the images in device 10, and optionally also receive the results of such calculations, or transmit determined centering parameters to the lens manufacturer. Furthermore, input / output interface 16 may include additional components for communicating with a user, such as a loudspeaker for outputting instructions or a microphone for receiving comments from the user. Figure 1 Only some components commonly used in mobile devices are shown, and other components may also be provided, such as sensors like accelerometers or orientation sensors, which are found in many conventional smartphones or tablet PCs and can be used in some embodiments to determine the orientation of the mobile device when capturing images.
[0087] For example, in order to generate a head model of a human head or its eye portion, or to determine centering parameters, multiple images from the head are captured from multiple different positions of the device 10 relative to the head by rotating the head relative to the mobile device. Figure 2 It is a flowchart illustrating a method according to a corresponding embodiment. Figures 3A to 3C , Figure 4 and Figures 5A to 5C Further demonstrations are shown. Figure 2 Various simplified diagrams of the methods. Typically, in order to perform... Figure 2 The method is captured by device 10 in Figure 3A The image shown schematically is of the head 30. To capture the head 30 from different directions, the head 30 is rotated relative to the device 10, as shown... Figure 3A As indicated by arrow 31 in the diagram. This rotation can be achieved by rotating the head 30 about the vertical axis or by moving the device 10 about the head 30. Also, as mentioned above, the device 10 can output corresponding commands for the rotation.
[0088] Return to Figure 2 In step 20, the method includes capturing a frontal image as a first image. This is in Figure 3B The image is schematically shown in which device 10 captures a frontal image of head 30, i.e., an image in which both eyes are visible and substantially symmetrical with respect to the nose. It should be noted that while the complete head is captured in the illustrated embodiment, in other embodiments only the eye portion may be captured when only a model of the eyes is required for a particular purpose.
[0089] In step 21, the method then includes rotating the head relative to the mobile device in a first direction while capturing the second image. This rotation could be, for example, a leftward rotation. Figure 3CThis illustrates how, while the mobile device 10 is capturing an image, the head 30 is turned to the left so that the ear 32 is now substantially visible from the side. As mentioned above, the device 10 can output a corresponding command to the person in order to initiate the rotation.
[0090] While device 10 is used to capture images in steps 20 and 21, in other embodiments a camera external to device 10 and linked to device 10 in a wired or wireless manner can be used. Furthermore, while in step 21 images from different directions are generated by rotating the head, in other embodiments other methods can be used, such as moving device 10 or an external camera relative to the head.
[0091] In step 22, the method includes stopping capturing the second image when a predefined marker is detected in the second image. Along with stopping capture, a corresponding instruction can be output to the person to rotate their head relative to the mobile device 10 in a first direction.
[0092] As explained above, predefined flags can indicate that a side view image has been captured, where flags necessary for later processing are visible. Figure 4 Examples of landmarks are shown. In this case, the cornea and pupil are landmarks related to the eyes, and additionally, various landmarks related to the ears (i.e., the earlobe, antitragus, tragus, the root of the helix, and the highest point of the helix (upper part of the helix)) can be used as landmarks. Other ear-related landmarks may also be used.
[0093] Figures 5A to 5C The example image shows the eye and ear regions as the head is further gradually rotated relative to the mobile device. Figure 5A This illustrates the case with relatively small rotation, in Figure 5B Turn your head a little more, and in Figure 5C This allows for greater head movement. The eye and ear areas are shown separately to avoid the need to show other areas that are not marked in this particular embodiment, but are part of the same image.
[0094] For the eye 50, the pupil 51 and a curve 52 representing the cornea (frontal contour of the cornea) are detected. For the ear, marker points 53 to 510 are detected, some of which correspond to reference points. Figure 4 The points being explained. For example, marker 56 represents the earlobe, and marker 510 represents the upper part of the helix, that is, the very top of the helix. For feature detectors, additional features may be used or required, such as... Figure 4 Points other than those shown in the diagram, as explained above. In Figure 5A and Figure 5B In this example, the camera captures the face from a slightly angled angle, starting from the front. Figure 5A , Figure 5B and Figure 5C The portion of the ear, not shown, is outside the camera's field of view.
[0095] exist Figure 5A , Figure 5B and Figure 5C In each of these images, the three landmarks of the eye (i.e., the frontal outlines of the cornea 50, pupil 51, and cornea 52) are visible. However, in... Figure 5A and Figure 5B In the middle, only a relatively small portion of the ear is visible, and only markers 53 to 56 can be detected, while markers 57 to 510 are not detectable. Conversely, in... Figure 5C In the middle, the ears are almost completely visible, and all ear markers are visible except for marker 59. Therefore, as it is necessary to... Figure 2 The subset of markers identified in step 21 can be markers 51, 52, 53 through 58 and 510. Once all these markers are detected, the method determines that the side view has been sufficiently captured. Figure 5C As can be seen, the eyes are indeed captured primarily from the side. The second image in which these markings are visible can also be considered a side view of the head for further processing. It should be noted that when a person is wearing actual eyeglass frames, in some views, the frames may obscure some eye markings, making the eye marking detection used to stop capturing in step 22 ensure that in the image where the markings are visible, these eyeglass frames do not obscure the eyes or at least the relevant eye markings.
[0096] Return to Figure 2 Optionally, the process can then be repeated for a second direction, i.e., steps 21 and 22 to 23. For example, when in step 21, the head is turned from the frontal position to the left, and then in step 23, the head can be turned from the frontal position to the right so as to also capture the left side of the head.
[0097] The captured images can then be used for model generation and centering parameter determination. Figure 6 The diagram shows a flowchart illustrating the corresponding method according to an embodiment.
[0098] exist Figure 6 In step 60, the method includes capturing an image of a person's head using a mobile device. The image capture process in step 60 is performed as described in the reference above. Figure 2 As explained above. In step 61, the method includes generating a 3D model of the head based on the image. Additional information about the position of the eyeglasses frame relative to the head is needed to determine the centering parameters. Figure 6Two alternatives are shown. In the first alternative, the person is not wearing glasses when the image is captured, or the glasses are removed from the image or model, as explained above. In this case, in step 62, a virtual frame is added, i.e., a model of the glasses is fitted to the 3D model of the head generated in step 61. Alternatively, when the person is wearing real glasses, in step 63, the frame is extracted from the image or from the generated 3D model, and this frame is used to determine the centering parameters. In both cases, the centering parameters are determined in step 64. This is performed in more detail as explained above.
[0099] Some other embodiments are defined by the following terms:
[0100] Clause 1. A method for capturing a head image, comprising:
[0101] Capture a first image of at least the eyes of the head using a mobile device.
[0102] While capturing multiple second images of the head using the mobile device, the head is rotated relative to the mobile device in a first direction.
[0103] Its features are,
[0104] Search for predefined flags in each of these second images, and
[0105] The capture of the plurality of second images is stopped in response to the identification of at least a predefined subset of these predefined flags in the second images of the plurality of second images.
[0106] Clause 2. The method as described in Clause 1, further comprising outputting an instruction to stop moving the head relative to the mobile device in the first direction in response to identifying at least the predefined subset in the plurality of second images.
[0107] Clause 3. The method as described in Clause 1 or 2, characterized in that it further includes, after the cessation:
[0108] While capturing multiple third images of the head using the mobile device, the head is rotated relative to the mobile device in a second direction opposite to the first direction.
[0109] Search for additional predefined flags in each of these third images, and
[0110] The capture of the third images is stopped in response to the identification of at least another predefined subset of these additional predefined flags in the third images among the plurality of third images.
[0111] Clause 4. The method as described in Clause 3, further comprising outputting an instruction to stop rotating the head relative to the mobile device in the second direction in response to identifying at least the other predefined subset in the plurality of third images.
[0112] Clause 5. The method as described in any one of Clauses 1 to 4, characterized in that the first image is a frontal image.
[0113] Clause 6. The method as described in any one of Clauses 1 to 5, characterized in that at least one of the group consisting of the predefined subset and the other predefined subset indicates a side view of the head.
[0114] Clause 7. The method as described in any one of Clauses 1 to 5, characterized in that at least one of the group consisting of these predefined marks and these additional predefined marks includes at least one point selected from the group consisting of: pupil, corneal apex, corneal contour and point on ear.
[0115] Clause 8. The method as described in any one of Clauses 1 to 7, characterized in that it further comprises determining at least one geometric parameter based on the first image and the second images.
[0116] Clause 9. The method of any one of Clauses 1 to 8, characterized in that it further comprises generating a head model based on the first image and at least one of the plurality of second images or based on the first image, at least one of the plurality of second images, and at least one of the plurality of third images.
[0117] Clause 10. The method as described in Clause 9, further comprising virtually fitting a model of the eyeglasses frame to the head model.
[0118] Clause 11. The method of any one of Clauses 1 to 9, characterized in that the head is wearing an eyeglass frame, and the method further includes identifying the eyeglass frame in at least one of the first image and the plurality of second images.
[0119] Clause 12. The method as described in any one of Clauses 10 or 11, characterized in that it further comprises calculating at least one centering parameter based on at least one of the group consisting of the first image, the second images, the head model, the fitted model of the eyeglasses frame, and the recognition of the eyeglasses frame.
[0120] Clause 13. A method for manufacturing spectacle lenses based on centering parameters calculated according to the method described in Clause 12.
[0121] Clause 14. A mobile device, characterized in that it comprises:
[0122] The camera, which is used to capture images, and
[0123] A processor configured to control the mobile device to perform the following operations:
[0124] Capture the first image of at least the eye portion of the head.
[0125] Multiple second images of the head are captured while the head rotates relative to the mobile device in a first direction.
[0126] The processor is further configured to control the mobile device to perform the following operations:
[0127] Search for predefined flags in each of these second images, and
[0128] The capture of the second images is stopped in response to the identification of at least a predefined subset of these predefined flags in the second images of the plurality of second images.
[0129] Clause 15. A computer program for a mobile device including a camera and a processor, characterized in that, when executed on the processor, it causes the mobile device to perform the method as described in any one of Clauses 1 to 12.
Claims
1. A method for head image capturing, comprising: providing a mobile device (10) with a first image of at least an eye portion of a head (30), providing the mobile device (10) with a plurality of second images of the head (30) from a plurality of directions, characterized in that searching, by the mobile device (10), for predefined landmarks (51-510) in each of the second images, wherein a landmark is a specific point or area on the head, and stopping providing the mobile device (10) with the plurality of second images in response to identifying, by the mobile device, at least a predefined subset of the predefined landmarks (51-510) in a second image of the plurality of second images, to ensure that landmarks required for a specific purpose are actually visible in the second images.
2. The method of claim 1, wherein, providing the mobile device with the first image and the second images comprises providing the first image and the second images captured by a camera (11, 12) of the mobile device.
3. The method of claim 1 or 2, wherein, providing the mobile device (10) with a plurality of second images of the head (30) from a plurality of directions is providing the mobile device (10) with the plurality of second images while the head (30) is turned in a first direction relative to the mobile device (10).
4. The method of claim 1 or 2, wherein, further comprising outputting instructions to start turning the head (30) in a first direction relative to the mobile device (10) after providing the first image and before providing the plurality of second images and / or outputting stopping turning the head (30) in the first direction relative to the mobile device (10) in response to identifying at least the predefined subset in the second image of the plurality of second images.
5. The method of claim 1 or 2, wherein, further comprising after the stopping: providing the mobile device (10) with a plurality of third images of the head (30) from a plurality of further directions; searching, with the mobile device (10), for further predefined landmarks (51-510) in each of the third images, and stopping providing the mobile device (10) with the plurality of third images in response to identifying, with the mobile device (10), at least a further predefined subset of the further predefined landmarks (51-510) in a third image of the plurality of third images.
6. The method of claim 3, wherein, further comprising after the stopping: providing the mobile device (10) with a plurality of third images of the head (30) from a plurality of further directions; searching, with the mobile device (10), for further predefined landmarks (51-510) in each of the third images, and stopping providing the mobile device (10) with the plurality of third images in response to identifying, with the mobile device (10), at least a further predefined subset of the further predefined landmarks (51-510) in a third image of the plurality of third images.
7. The method of claim 4, wherein, further comprising after the stopping: providing the mobile device (10) with a plurality of third images of the head (30) from a plurality of further directions; searching, with the mobile device (10), for further predefined landmarks (51-510) in each of the third images, and stopping providing the mobile device (10) with the plurality of third images in response to identifying, with the mobile device (10), at least a further predefined subset of the further predefined landmarks (51-510) in a third image of the plurality of third images. stopping providing the plurality of third images in response to recognizing at least the further predefined subset of the further predefined markers (51-510) in the third image of the plurality of third images with the mobile device (10).
8. The method of claim 5, wherein, providing the third images to the mobile device (10) comprises providing the third images captured by a camera (11, 12) of the mobile device.
9. The method of claim 5, wherein, capturing the plurality of third images while the head is turning in a second direction relative to the mobile device (10) after stopping providing the plurality of second images and before providing the plurality of third images.
10. The method of claim 6, wherein, capturing the plurality of third images while the head is turning in a second direction relative to the mobile device (10) after stopping providing the plurality of second images and before providing the plurality of third images.
11. The method of claim 7, wherein, capturing the plurality of third images while the head is turning in a second direction relative to the mobile device (10) after stopping providing the plurality of second images and before providing the plurality of third images.
12. The method of claim 5, wherein, further comprising outputting, by the mobile device (10), instructions to start turning the head (30) in a second direction relative to the mobile device (10) after stopping providing the plurality of second images and before providing the plurality of third images.
13. The method of claim 6, wherein, further comprising outputting, by the mobile device (10), instructions to start turning the head (30) in a second direction relative to the mobile device (10) after stopping providing the plurality of second images and before providing the plurality of third images.
14. The method of claim 7, wherein, further comprising outputting, by the mobile device (10), instructions to start turning the head (30) in a second direction relative to the mobile device (10) after stopping providing the plurality of second images and before providing the plurality of third images.
15. The method of claim 10, wherein, the second direction is opposite to the first direction.
16. The method of claim 11, wherein, the second direction is opposite to the first direction.
17. The method of claim 13, wherein, the second direction is opposite to the first direction.
18. The method of claim 14, wherein, the second direction is opposite to the first direction.
19. The method of claim 12, wherein, further comprising outputting instructions to stop turning the head (30) in the second direction relative to the mobile device (10) in response to recognizing at least the further predefined subset in the third image of the plurality of third images.
20. The method of claim 5, wherein, searching for the further predefined markers (51-510) uses trained machine learning logic.
21. The method of claim 1 or 2, wherein, searching for the predefined markers (51-510) uses trained machine learning logic.
22. The method of claim 1 or 2, wherein, the first image is a frontal image.
23. The method of claim 5, wherein, at least one of the group consisting of the predefined subset and the further predefined subset is indicative of a side view of the head (30).
24. The method of claim 5, wherein, at least one of the group consisting of the predefined markers (51-510) and the further predefined markers (51-510) comprises at least one point selected from the group consisting of: a pupil, a corneal apex, a corneal profile, and a point on an ear.
25. The method of claim 1 or 2, wherein, further comprising determining at least one geometric parameter based on the first image and the second images.
26. The method of claim 5, wherein, further comprising generating a head model based on the first image, and at least one of the plurality of second images, or based on the first image, at least one of the plurality of second images, and at least one of the plurality of third images.
27. The method of claim 26, wherein, Generating the head model is performed by the mobile device (10).
28. The method of claim 26, wherein, Further comprising virtually fitting a model of a spectacle frame to the head model.
29. The method of claim 1 or 2, wherein, The head (30) wears a spectacle frame, and the method further comprises identifying the spectacle frame in the first image, and at least one of the plurality of second images.
30. The method of claim 28, wherein, Further comprising calculating at least one centering parameter based on at least one of the group consisting of the first image, the second images, the head model, the fitted model of the spectacle frame, and the identification of the spectacle frame.
31. A method for manufacturing a spectacle lens based on a centering parameter calculated according to the method of claim 30.
32. A mobile device (10) characterized by: Comprising: a processor (15) configured to control the mobile device (12) to: obtain a first image of at least an eye portion of a head (30), obtain a plurality of second images of the head (30) from a plurality of directions, characterized in that the processor (15) is further configured to control the mobile device to: search for predefined landmarks (51-510) in each of the second images, wherein a landmark is a specific point or area on the head, and stop obtaining the plurality of second images in response to identifying at least a predefined subset of the predefined landmarks (51-510) in a second image of the plurality of second images with the mobile device to ensure that the landmarks needed for a specific purpose are actually visible in the second images.
33. A computer program for a mobile device (10) comprising a processor (15), characterized in that, a computer program product comprising computer executable instructions for causing the mobile device (10) to perform the method of any one of claims 1 to 30 when executed on the processor.
Citation Information
Patent Citations
Computer-implemented method for detecting a cornea vertex
EP3363346B1
Method, device and computer program for virtual adapting of a spectacle frame
EP3410178A1
Method, device and computer program for virtual adapting of a spectacle frame
EP3425447A1
Computer-implemented method for determining a representation of a spectacle socket rim or a representation of the edges of the glasses of a pair of spectacles
EP3574370B1
Method, device and computer program for virtual adapting of a spectacle frame
EP3631570B1