Fitting of head-mounted wearable device from two-dimensional image
By capturing the user's two-dimensional image on the user-operated computing device, generating a user grid and adjusting the virtual frame in combination with the reference grid, the problem of difficulty in providing accurate adaptation and customization in existing systems is solved, and a more accurate selection, size setting and adaptation of head-mounted wearable devices is achieved to meet the specific needs of users.
Patent Information
- Application Number
- CN202280101212.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2025-06-03
AI Technical Summary
Existing systems are difficult to provide accurate adaptation and customization, especially when head-mounted wearable devices are not selected, sized and adapted to meet the specific needs of users, especially facial features, without going to a retail store.
By capturing the user's two-dimensional image on the user-operated computing device, generating a user grid and identifying facial landmarks, rigidly transforming it in combination with the reference grid and the virtual frame, adjusting the position of the virtual frame to correspond to the inode of the user grid, thereby providing more accurate selection, size setting and adaptation of the head-mounted wearable device.
It realizes that without going to a retail store, users can self-guided and unsupervisedly capture image data, and make more accurate selection, size setting and adaptation of head-mounted wearable devices to meet users' specific facial characteristics needs.
Smart Images

Figure CN120092268A_ABST
Abstract
Description
Technical Field
[0001] This specification generally relates to sizing and / or fitting of wearable devices, and in particular, to sizing and / or fitting of head-mounted wearable devices. Background Art
[0002] Wearable devices can include, for example, head-mounted wearable devices, wrist-mounted wearable devices, hand-mounted wearable devices, pendants, fitness trackers, body sensors, and other such devices. Head-mounted wearable devices can include, for example, smart glasses, headphones, goggles, earbuds, etc. Wrist-mounted / hand-mounted wearable devices can include, for example, smart watches, smart bracelets, smart rings, etc. In some cases, a user may want to select and / or customize a wearable device for fitting and / or functionality. For example, a user may wish to select and / or customize an ocular wear, including selection of a frame, incorporation of prescription lenses, and other such features. Summary of the Invention
[0003] Systems and methods are described herein for providing selection, sizing, and / or fitting of a head-mounted wearable device based on a two-dimensional image of a user captured via an application executed on a computing device operated by the user. Based on one or more facial landmarks detected in the user's image, a user grid representing the head (e.g., a portion of the head, such as the user's face) is generated. A virtual frame is positioned on a reference grid, e.g., at a location on the reference grid corresponding to the sellion. The reference grid can represent a generic face generated based on data collected from a relatively large number of subjects. A rigid transform can be performed to project the reference grid and the virtual frame onto the user grid. The position of the virtual frame can be shifted or adjusted such that the bridge portion of the virtual frame is positioned to correspond to the sellion of the user grid, such that the virtual frame is positioned as a user might wear a corresponding physical frame. Generally, the user grid and / or the reference grid can also represent the user's head.
[0004] The proposed solution particularly relates to a (computer-implemented) method, especially for the partial or full automation of the selection, sizing, and / or adaptation of a head-mounted wearable device to meet the specific needs of a user, the method comprising: capturing image data of an initial image including the face of the user via an application executed on a computing device operated by the user; generating a user mesh representing the face of the user based on the image data; identifying index nodes in the user mesh corresponding to a set portion of the face of the user captured in the image data; identifying index nodes in a reference mesh, the index nodes of the reference mesh corresponding to a set portion of the reference mesh, the set portion of the reference mesh corresponding to the set portion of the user mesh; positioning a virtual frame of the head-mounted wearable device at a position on the reference mesh corresponding to the index nodes of the reference mesh; projecting the reference mesh and the virtual frame onto the user mesh; and adjusting the position of the virtual frame to correspond to the index nodes of the user mesh. Based on the virtual frame whose position has been adjusted relative to the user mesh, components of the head-mounted wearable device and / or a model of the head-mounted wearable device are selected or manufactured for the user for whom the user mesh has been generated. For example, an image from the initial image captured by the user - including a virtual rendering of a frame positioned on the face of the user - can be presented to the user. The rendering of the frame on the face of the user can represent the actual adaptation of the corresponding physical frame on the face and head of the user, thus allowing the user to make a relatively accurate assessment of the adaptation and appearance of the frame. Thereby, the partial or full automation of the selection, sizing, and / or adaptation of a head-mounted wearable device based on a two-dimensional image of the user to meet the specific needs of the user, especially the specific facial characteristics of the user, can be facilitated and / or accelerated.
[0005] Details of one or more implementations are set forth in the accompanying drawings and the following description. Other features will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1A Shows an example head-mounted wearable device worn by a user.
[0007] Figure 1B is Figure 1A a front view of the example head-mounted wearable device shown in Figure 1C and
[0008] Figures 2A to 2C shows an example eye adaptation measurement.
[0009] Figure 3 is a block diagram of a system according to an implementation described herein.
[0010] Figure 4A andFigure 4B is a front view of a user, and Figure 4C is a side view of the user, showing example facial and / or cranial landmarks.
[0011] Figures 5A to 5E shows a process for sizing and / or fitting a frame of a head-mounted wearable device according to implementations described herein.
[0012] Figure 6 is an example sizing and / or fitting image according to implementations described herein.
[0013] Figure 7 is a flowchart of an example method according to implementations described herein. DETAILED DESCRIPTION
[0014] The selection of a wearable device—such as a head-mounted wearable device in the form of an ocular wear or glasses—can depend on a determination of physical or wearable fit to ensure that the ocular wear is comfortable for the user to wear and / or aesthetically complementary to the user. Incorporating corrective lenses into a head-mounted wearable device can depend on a determination of ocular fit to ensure that the head-mounted wearable device can provide the desired vision correction. In the case of a head-mounted wearable computing device—in the form of, for example, smart glasses that include computing / processing and display capabilities—the selection can also depend on a determination of display fit to ensure that visual content is visible to the user. Existing systems for purchasing these types of wearable devices cannot provide accurate fitting and customization, especially without a visit to a retail store. That is, accurate sizing and / or fitting typically depends on the user visiting a retail store where samples are available for physical trial, and an optician facilitates the determination of wearable fit and / or ocular fit and / or aesthetic fit based on the physical trial and measurements collected using specialized equipment. In some cases, existing virtual systems that provide an online selection of a wearable device (such as ocular wear / glasses) only superimpose an image of the selected frame on an image of the user. Virtually placing an image of the selected frame on an image of the user does not take into account the user's facial features, which can affect the fit of the physical frame to the user. For example, variations in the height of the nose bridge can affect how the physical frame is positioned on the user's face / head, thereby affecting the fit and functionality of the head-mounted wearable device when worn by the user. Thus, these types of systems can produce inaccurate results when selecting ocular wear in this manner.
[0015] Systems and methods according to implementations described herein provide virtual fitting of wearable devices based on one or more features detected in image data. Systems and methods according to implementations described herein utilize a reference grid or canonical grid, which can be a three-dimensional representation of the body part on which the wearable device is to be worn. For example, in the selection and / or sizing and / or fitting of a head-mounted wearable device, the reference grid or average grid or canonical grid can represent an average or generic face and / or head generated based on previously collected data for a relatively large number of users. Systems and methods according to implementations described herein can generate a user grid, which can be a representation of the body part of the user on which the wearable device is to be worn. In the selection and / or sizing and / or fitting of a head-mounted wearable device, the user grid can represent the user's face / head.
[0016] In some examples, the image data includes two-dimensional images captured by the user via an application executed on the user's computing device. In some examples, a selected feature (i.e., one of the one or more features detected in the image data) is mapped to a key point in the reference grid or canonical grid. A rigid transformation can be applied to the key point in the reference grid to project the key point onto the user grid. The distance between the key point in the reference grid and the key point in the user grid can be used to align or adjust the reference grid and the user grid. In some examples, the distance between the key point in the reference grid and the key point in the user grid can be used to determine, for example in pixels, the vertical and horizontal distances for projection onto the two-dimensional image captured by the user. This can provide a more accurate placement of the wearable device on the two-dimensional image captured by the user operating the computing device. In some examples, the systems and methods described herein provide fitting of a head-mounted wearable device in the form of smart glasses, which include processing / computing capabilities and display capabilities and / or corrective lenses. Systems and methods according to implementations described herein can facilitate the user capturing, in a self-guided, unsupervised or unregulated manner, image data for the detection and fitting of one or more features of the wearable device without having to visit a retail store and / or have an in-person or virtual appointment with a technician or sales agent.
[0017] In the following, for purposes of discussion and illustration only, systems and methods will be described with respect to the selection, sizing, and / or fitting of a head-mounted wearable device. For purposes of discussion and illustration only, among the features detectable within a two-dimensional image captured by a user, the nasion point will be used to coordinate a rigid transformation between a reference grid and user facial key points. The principles described herein can be applied to the sizing and / or fitting of other types of wearable devices, which other types of wearable devices include, for example, glasses that may or may not include processing / computing / display capabilities and / or corrective lenses, or other types of wearable devices. Similarly, as a supplement to or alternative to the nasion point detected within image data, the principles described herein can also utilize other features.
[0018] Figure 1A Shown is a user wearing an example head-mounted wearable device 100 in the form of smart glasses or augmented reality glasses, the example head-mounted wearable device including display capabilities, eye / gaze tracking capabilities, and computing / processing capabilities. Figure 1B is Figure 1A a front view of the example head-mounted wearable device 100 shown in Figure 1C and
[0019] is a rear view of the example head-mounted wearable device. The example head-mounted wearable device 100 includes a frame 110. The frame 110 includes a front frame portion 120 and a pair of arm portions 130 rotatably coupled to the front frame portion 120 via respective hinge portions 140. The front frame portion 120 includes a rim portion 123 surrounding a respective optical portion in the form of a lens 127, wherein a bridge portion 129 connects the rim portions 123. The arm portions 130 are coupled (e.g., pivotally or rotatably coupled) to the front frame portion 120 at a peripheral portion of the respective rim portions 123. In some examples, the lens 127 is a corrective / prescription lens. In some examples, the lens 127 is an optical material including a glass and / or plastic portion that does not necessarily include corrective / prescription parameters.
[0019] In some examples, the wearable device 100 includes a display device 104 that can output visual content, for example, at an output coupler 105 such that the visual content is visible to the user. For purposes of discussion and illustration only, in Figure 1B and Figure 1CIn the example shown, the display device 104 is disposed in one of the two temple portions 130. The display device 104 may be disposed in each of the two temple portions 130 to provide a binocular output of content. In some examples, the display device 104 may be a see-through near-eye display. In some examples, the display device 104 may be configured to project light from a display source onto a portion of a teleprompter glass that acts as a beam splitter angled (e.g., 30 degrees to 45 degrees). The beam splitter may allow reflection and transmission values that allow light from the display source to be partially reflected while the rest of the light is transmitted. Such an optical design may allow the user to see two physical items in the world, e.g., through the lens 127, beside the content (e.g., digital images, user interface elements, virtual content, etc.) output by the display device 104. In some implementations, waveguide optics may be used to depict content on the display device 104.
[0020] In some examples, the head-mounted wearable device 100 includes one or more of an audio output device 106 (such as, for example, one or more speakers), a lighting device 108, a sensing system 111, a control system 112, at least one processor 114, and an outward-facing image sensor 116 (e.g., a camera). In some examples, the sensing system 111 may include various sensing devices, and the control system 112 may include various control system devices, which include, for example, one or more processors 114 operably coupled to components of the control system 112. In some examples, the control system 112 may include a communication module that provides communication and information exchange between the wearable device 100 and other external devices. In some examples, the head-mounted wearable device 100 includes a gaze tracking device 115 for detecting and tracking the direction and movement of eye gaze. The data captured by the gaze tracking device 115 may be processed to detect and track the direction and movement of eye gaze as user input. For purposes of discussion and illustration only, in Figure 1B and Figure 1C the example shown, the gaze tracking device 115 is disposed in one of the two temple portions 130. In Figure 1B and Figure 1C the example arrangement shown, the gaze tracking device 115 is disposed in the same temple portion 130 as the display device 104 such that the user's eye gaze can be tracked not only with respect to objects in the physical environment but also with respect to the content output for display by the display device 104. In some examples, the gaze tracking device 115 may be disposed in each of the two temple portions 130 to provide gaze tracking of each of the user's two eyes. In some examples, the display device 104 may be disposed in each of the two temple portions 130 to provide a binocular display of visual content.
[0021] When selecting and / or sizing and / or fitting a wearable device (such as the example head-mounted wearable device 100 shown in Figures 1A to 1C ) for a particular user, a variety of different sizing and fitting measurements and / or parameters can be considered. This can include, for example, wearable fitting parameters or wearable fitting measurements. Wearable fitting parameters / measurements can consider how a particular frame 110 fits and / or looks and / or feels on a particular user. Wearable fitting parameters / measurements can consider a variety of factors, such as, for example, whether the frame portion 123 and the bridge portion 129 are shaped and / or sized such that the bridge portion 129 rests comfortably on the user's nose, whether the frame 110 is wide enough to be comfortable on the temples but not so wide that the frame 110 cannot remain relatively stationary when worn by the user, whether the temple portions 130 are sized to rest comfortably on the user's ears, and other such comfort-related considerations. Wearable fitting parameters / measurements can consider other on-wear considerations, including how the frame 110 is positioned based on the user's natural head posture / the position where the user tends to naturally wear his / her glasses. In some examples, aesthetic fitting measurements or parameters can consider whether the frame 110 is aesthetically pleasing to the user / is compatible with the user's facial features, etc.
[0022] In a head-mounted wearable device that includes display capabilities, display fitting parameters or display fitting measurements can be considered to select and / or size and / or fit the head-mounted wearable device 100 for a particular user. Display fitting parameters / measurements can be used to configure the display device 104 of the selected frame 110 for a particular user such that the content output by the display device 104 is visible to the user. For example, display fitting parameters / measurements can facilitate the calibration of the display device 104 such that visual content is output within at least a set portion of the user's field of view. For example, display fitting parameters / measurements can be used to configure the display device 104 to provide at least a set level of gaze ability that corresponds to the amount or portion or percentage of visual content that is visible to the user at the periphery (e.g., the least visible corners) of the user's field of view.
[0023] In an example where the head-mounted wearable device 100 will include corrective lenses, eye fitting parameters or eye fitting measurements can be considered during the selection and / or sizing and / or fitting process. Figures 2A to 2CSome example eye fit measurements are shown. The eye fit measurements can include, for example, the pupil height PH (the distance from the center of the pupil to the bottom of the corresponding lens 127). The eye fit measurements can include the interpupillary distance IPD (the distance between the pupils). The IPD can be characterized by monocular pupil distances, such as the left pupil distance LPD (the distance from the central portion of the bridge of the nose to the left pupil) and the right pupil distance RPD (the distance from the central portion of the bridge of the nose to the right pupil). The eye fit measurements can include the pantoscopic angle PA (the angle defined by the tilt of the lens 127 relative to the vertical). The eye fit measurements can include the vertex distance V (the distance from the cornea to the corresponding lens 127). The eye fit measurements can include other such parameters, or provide a metric for the selection and / or sizing and / or fitting of a head-mounted wearable device 100 (with or without the display device 104 as described above) that includes corrective lenses. In some examples, the eye fit measurements, together with display fit measurements, can provide for the output of visual content by the display device 104 within a defined three-dimensional volume such that the content is within the user's corrected field of view and is thus visible to the user.
[0024] Figure 3 is a block diagram of an example system for predicting the sizing and / or fitting of a wearable device based on at least one key point or landmark or feature detected in at least one image (e.g., a two-dimensional image) captured by a computing device operated by a user. The system can utilize at least one three-dimensional reference grid or canonical grid in determining the sizing and / or fitting of the wearable device. In an example where the wearable device is a head-mounted wearable device, the reference grid can represent a generic head generated based on data previously collected from a relatively large number of subjects. Wearable devices that can be sized and / or fitted in this manner by the system can include the various wearable computing devices described above. Hereinafter, for purposes of discussion and illustration only, the sizing and / or fitting of a head-mounted wearable device (such as the example head-mounted wearable device 100) by the system will be described.
[0025] The system can include one or more computing devices 300. The computing device 300 can be operated by a user for whom the wearable device is to be sized / fitted. The computing device 300 can be, for example, a handheld device such as a smart phone or a tablet computing device, a desktop or laptop computing device, and other such computing devices that can be operated by a user to capture an image of the user. The computing device 300 can access additional resources 302 to facilitate sizing and / or fitting of the wearable device. In some examples, the additional resources 302 can be locally available on the computing device 300. In some examples, the additional resources 302 can be available to the computing device 300 via the network 306. In some examples, some of the additional resources 302 can be locally available on the computing device 300 and some of the additional resources 302 can be available to the computing device 300 via the network 306. The additional resources 302 can include, for example, server computer systems, processors, databases, machine learning modules, memory storage, etc. In some examples, the processor 390 can provide various processing functions via, for example, an object recognition engine, a pattern recognition engine, a simulation engine, a fitting engine, and other such processors. In some examples, the additional resources 302 include machine learning models and / or algorithms that support sizing and / or fitting of the wearable device.
[0026] The computing device 300 can operate under the control of the control system 370. The computing device 300 can communicate directly (via wired and / or wireless communication) or via the network 306 with one or more external computing devices 304 (another wearable computing device, another mobile computing device, etc.). In some examples, the computing device 300 includes a communication module 380 to facilitate external communication. In some examples, the computing device 300 includes a sensing system 320 that includes various sensing system components, which include, for example, one or more image sensors 322, one or more position / orientation sensors 324 (including, for example, an inertial measurement unit, an accelerometer, a gyroscope, a magnetometer, etc.), one or more audio sensors 326 that can detect audio input, one or more touch input sensors 328 that can detect touch input, and other such sensors. The computing device 300 can include more or fewer sensing devices and / or combinations of sensing devices.
[0027] In some examples, the image sensor 322 can include, for example, a camera, such as, for example, a front-facing camera, an outward-facing or world-facing camera, etc., that can capture static and / or moving images of the environment external to the computing device 300. The static and / or moving images can be displayed by a display device of the output system 340, and / or sent externally via the communication module 380 and the network 306, and / or stored in the memory 330 of the computing device 300 and / or in a memory device available in the additional resources 302.
[0028] The computing device 300 may include one or more processors 390. The processor 390 may include various modules or engines configured to perform various functions. In some examples, the processor 390 may include an object recognition module, a pattern recognition module, a configuration recognition module, and other such processors. The processor 390 may be formed in a substrate configured to execute one or more machine-executable instructions or software fragments, firmware, or a combination thereof. The processor 390 may be semiconductor-based, including semiconductor materials that can execute digital logic. The memory 330 may include any type of storage device that stores information in a format readable and / or executable by the processor 390. The memory 330 may store applications and modules that perform certain operations when executed by the processor 390. In some examples, the applications and modules may be stored in an external storage device and loaded into the memory 330.
[0029] As described above, the systems and methods according to the implementations described herein provide for the selection and / or sizing and / or fitting of the frame of a head-mounted wearable device. In some examples, one or more key points, or one or more landmarks, or one or more features may be detected in the image data of the user's face and / or head. The image data may be a two-dimensional image captured via an application executed on a computing device operated by the user. In some examples, one or more key points, or landmarks, or features include the sellion or sellion point. The sellion point may be defined at the midline of the user's nasal root or nasal bridge. The sellion point may be located at the point of maximum curvature of the nasal profile, at the root end of the nasal bridge, at the transition point between the nasal bridge and the forehead. The sellion point may represent the deepest depression of the nasal bone.
[0030] Figure 4A is an example two-dimensional image 400 that can be captured via an application executed on a computing device (such as, for example, the example computing device 300 described above Figure 3 ). The example two-dimensional image 400 provides a basic front view of the user. Figure 4B Shows a plurality of example key points or features or landmarks 410 that can be detected in the user's two-dimensional image 400, the plurality of example key points or features or landmarks including the sellion point 420 at the root end of the user's nasal bridge. More or fewer key points or features or landmarks 410 than those shown in Figure 4B may be detected in the two-dimensional image 400 captured via an application executed on a computing device operated by the user. Figure 4C is a side view of the user, which is provided only for further showing the position of the sellion point 420 relative to the user's nasal bridge.
[0031] As described above, a two-dimensional image 400 of a user's face can be captured via an application executed on a computing device (such as computing device 300) operated by the user. An object recognition engine and / or a pattern recognition engine available via additional resource 302 can analyze the image 400 to detect one or more landmarks 410, which include the nasion point 420. In some examples, any one or a grouping of the landmarks 410 can be used to predict the positioning and / or fitting of a frame on the user's face and / or head. However, the physical positioning of the frame on the user's face / head can vary and be significantly affected based on the nasal bridge height, nose shape, etc. For example, a relatively higher or lower nasal bridge height can result in a significant shift in the vertical positioning of the frame on the user's face. This can lead to inaccurate virtual sizing and / or fitting of the frame of the head-mounted wearable device. Therefore, in a virtual fitting scenario, using the nasion point 420 to predict the virtual placement of the frame on the image 400 of the user's face / head can provide a more accurate prediction of the sizing and fitting of the frame of the head-mounted wearable device.
[0032] Hereinafter, systems and methods will be described for using the nasion point 420 detected in the two-dimensional image 400 for the purpose of virtual fitting to predict the placement of a virtual frame on the image 400 of the user. In some examples, for instance, a simulation engine available to the computing device 300 via additional resource 302 can access one or more machine learning models available via the additional resource to generate Figure 5A the user mesh 500 shown in. The user mesh 500 can be generated based on data obtained through the analysis of the two-dimensional image 400 by the object recognition engine and / or the pattern recognition engine. Based on the two-dimensional image 400 captured by the user, the user mesh 500 can represent the user's face / head. The user mesh 500 can include a plurality of interconnected nodes, some of which are labeled with reference numeral 505 in Figure 5A shown.
[0033] In some examples, as Figure 5B shown, a reference mesh 550 can be accessible to the computing device 300 via one of the databases in, for example, the additional resource 302. The reference mesh 550 can represent a generic face / head generated based on data collected from a relatively large number of subjects. The reference mesh 550 can include a plurality of interconnected nodes, some of which are labeled with reference numeral 505 in Figure 5BIt is marked with reference numeral 555 in the figure. Since the reference grid 550 represents a somewhat general grid of a general face / head generated based on data collected from a relatively large number of subjects, the reference grid 550 is not specific to the user grid 500. That is, there is no one-to-one correspondence between the nodes 555 of the reference grid 550 and the nodes 505 of the user grid 500.
[0034] In some examples, one of the multiple nodes 555 of the reference grid 550 can be identified as an index node. In the examples described herein, the index node of the reference grid 550 can be the node among the multiple nodes 555 that is closest to the mapped position of the nasion point in the reference grid 550. In this example, the index node can be identified as the nasion point node 552 as shown in Figure 5B as shown. As Figure 5C shown, the virtual frame 590 can be positioned on the reference grid 550, where the bridge portion 598 of the virtual frame 590 is positioned to correspond to the nasion point node 552 to simulate the position where a user with a face / head matching the reference grid 550 would naturally wear the corresponding physical frame.
[0035] A rigid transformation of the reference grid 550 (on which the virtual frame 590 is positioned) can be performed to project the reference grid 550 (and the virtual frame 590) onto the user grid 500, as shown in Figure 5D as shown. In some examples, the rigid transformation can include rotation, translation, and scaling of some number of nodes 555 of the reference grid 550 or a subgroup of the nodes 555 of the reference grid 550 to be related to the corresponding / corresponding subgroup of the nodes 505 of the user grid 500. This rigid transformation does not produce a node-to-node correspondence for each node of the reference grid 550 and the user grid 500. However, this method can provide an approximation that can be adjusted to predict the size setting and / or fitting of a selected frame of the head-mounted wearable device.
[0036] Figure 5D shows the initial placement position of the virtual frame 590 on the user's face / head based on the projection of the reference grid 550 (and the virtual frame 590) onto the user grid 500. As shown in Figure 5D as shown, there is a positional difference (a vertical difference in the orientation shown in Figure 5D ) between the index node of the reference grid 550 (i.e., the nasion point node 552 (and the associated initial placement position of the bridge portion 598 of the virtual frame 590 from the reference grid 550)) and the index node of the user grid 500 (i.e., the nasion point node 502 identified in the user grid 500). Figure 5E shows the adjusted virtual placement position of the virtual frame 590 on the user's head / face. InFigure 5E In [reference], the position of the virtual frame 590 has been adjusted or shifted such that the position of the bridge portion 598 corresponds to the nasion node 502 of the user mesh 500. The virtual frame 590 is moved from Figure 5D the initial virtual placement position shown in [reference] to Figure 5E the adjusted virtual position shown in [reference] can position the virtual frame 590 at a position on the user's head / face that more closely simulates how the user would wear the corresponding physical frame. This can provide a more representative indication for sizing and / or fitting the selected frame of the head-mounted wearable device on the user's face / head.
[0037] Figure 6 A two-dimensional image 600 of the virtual frame 590 positioned at the adjusted virtual position on the user's head / face is shown. The two-dimensional image 600 can be presented to the user, for example, via an application executed on a computing device operated by the user. In some examples, the virtual frame 590 can be superimposed on Figure 4A the initial image 400 shown in [reference] to generate Figure 6 the image 600 shown in [reference]. Figure 6 The position of the virtual frame 590 in the image 600 shown in [reference] represents how the user would wear the corresponding physical frame, how the corresponding physical frame would look on the user's face, and how the corresponding physical frame would fit the user. Thus, the user can evaluate the image 600 of the virtual frame 590 positioned at the adjusted virtual position on the user's head / face to confirm the sizing and / or fitting of the selected frame of the head-mounted wearable device.
[0038] As described above, the nasion can be detected relatively reliably within the two-dimensional image 400 captured by the user. Thus, the nasion 420 detected in the image 400 can provide a relatively reliable reference point for the placement of the virtual frame 590 on the user's face / head in the image 400, which in turn can provide a relatively reliable virtual sizing and / or fitting of the head-mounted wearable device for the user. In the above example, a single generic reference grid was used. As described above, the reference grid 550 was generated based on data collected from a relatively large number of subjects. The rigid transformation of a subgroup of the nodes 555 of the reference grid 550 to the corresponding subgroup of the nodes 505 of the user grid 500 can provide a relatively reliable basis for the initial virtual placement of the virtual frame 590 on the user's face / head in the two-dimensional image 400. The identification of the nasion node 502 in the user grid 500 (based on the detection of the nasion 420 in the user's image 400) and the identification of the nasion node 552 in the reference grid 550 can provide a basis for the shift of the virtual frame 590 from the initial virtual placement position to an adjusted virtual placement position that is close to the nasion node 502 in the user grid 500 (corresponding to the nasion 420 identified in the user's image 400). The adjusted virtual placement position can represent the position where the user would naturally wear the corresponding physical frame.
[0039] The relatively lower computational load associated with performing the rigid transformation and adjusting the placement position of the virtual frame 590 based on the position of the nasion in this manner can allow these processes to be performed locally on the user device, rather than relying on the use of external computing resources. This can facilitate the virtual sizing and / or fitting of the head-mounted wearable device for the user without having to visit a retail store, and / or without the assistance of an optician or sales agent (virtual or in person), and / or without the use of specialized equipment.
[0040] The above example utilized a single reference grid to determine the placement position of the virtual frame 590 on the user's face / head for virtual sizing and / or fitting of the head-mounted wearable device. In some implementations, more than one reference grid can be used to perform the sizing and / or fitting operations as described above.
[0041] For example, as described above, the nasal bridge height (which can be detected, for example, based on the identification of the nasion point 420 in the user's image 400) can affect the manner and location in which the user wears the physical frame of the head-mounted wearable device. The systems and methods described above are implemented using a single reference grid projected onto the user mesh via a rigid transformation. In some examples, multiple reference grids based on nasal bridge height can be used to facilitate sizing and / or fitting of the head-mounted wearable device. In some examples, the system can select, from among the multiple reference grids, the reference grid that is most suitable for sizing and / or fitting the frame of the head-mounted wearable device for a particular user. For example, the system can select a first reference grid in response to determining that the user has an average nasal bridge height. Similarly, the system can select a second reference grid in response to determining that the user has a relatively high nasal bridge height, and can select a third reference grid in response to determining that the user has a relatively low nasal bridge height. The first reference grid can be generated based on data collected from a relatively large number of subjects all determined to have an average nasal bridge height. Similarly, the second reference grid can be generated based on data collected from a relatively large number of subjects all determined to have a relatively high nasal bridge height, and the third reference grid can be generated based on data collected from a relatively large number of subjects all determined to have a relatively low nasal bridge height. The implementation of a reference grid that is more closely adapted to the user's facial characteristics can improve the accuracy of the virtual placement of the virtual frame on the user's face / head (compared to how the user would wear the corresponding physical frame), and / or can reduce the computational load associated with the virtual placement of the frame on the image of the user's face / head.
[0042] In some examples, based on the detected position of the nasion point 420, compared to other facial landmarks 410 detected in the two-dimensional image 400, the system can determine that the user has an average nasal bridge height, or a relatively high nasal bridge height, or a relatively low nasal bridge height. In some examples, the object / pattern recognition engine, simulation engine, and / or machine learning model described above can facilitate the detection of the facial landmarks 410 and the nasion point 420, and the determination of whether the user belongs to a first category of users having an average nasal bridge height, a second category of users having a relatively high nasal bridge height, or a third category of users having a relatively low nasal bridge height. In some examples, threshold distances between various facial landmarks 410 and / or between the facial landmarks 410 and the nasion point 420 can be used to determine whether the user belongs to the first category, the second category, or the third category. The system can select a reference grid, for example from among the multiple reference grids available to the system, for virtual sizing and / or fitting of the frame of the head-mounted wearable device based on which category is associated with the user. The selected reference grid can then be applied in a manner similar to that described above to place the virtual frame 590 on the user's face / head.
[0043] The bridge height is just one example of how the use of multiple reference grids can further facilitate sizing and / or fitting of a head-mounted wearable device for a particular user. Other reference grids can be implemented similarly based on other characteristics—e.g., characteristics and / or features that can be detected in an image 400 captured via an application executed on a computing device operated by the user. This can include, for example, characteristics associated with the width of the user's nose, such as one or more widths taken at a specified portion of the nose, a ratio of widths taken at specified portions of the nose, etc. Other characteristics can include, for example, a detected profile and / or change in profile of the nose—which can affect where the bridge portion of the frame will be positioned on the user's nose, and other such characteristics and / or features.
[0044] Figure 7FIG. 700 is a flowchart of an example method 700 that is an example method according to the implementations described herein. A user operating a computing device (such as, for example, the computing device 300 described above or other computing devices) may initiate an image capture function of the computing device (block 710). The image capture function may be accessed via an application running on the computing device operated by the user. Initiation of the image capture function may cause an image sensor (such as, for example, the image sensor of the front camera of the computing device 300) to capture two-dimensional image data including the user's face and / or head (block 715). One or more fixed features or landmarks may be detected in the image (block 720). The one or more fixed features or landmarks may include facial landmarks that remain substantially static, such as the nasion point defined at the midline of the user's radix or bridge of the nose, and other such fixed facial landmarks. A user mesh may be generated based on an analysis of the image and the one or more facial landmarks detected (block 725). In some examples, the user mesh may be generated by one or more machine learning models accessible to the computing device. The system may access or retrieve a reference mesh (block 730). The reference mesh may be retrieved from a database accessible to the user. In some examples, a single reference mesh is available. In some examples, multiple reference meshes representing multiple different facial characteristics (such as different bridge of the nose height classifications) may be available. A virtual frame associated with sizing and / or fitting a head-mounted wearable device for the user may be placed on the reference mesh (block 735). The virtual frame may be placed on the reference mesh based on nodes in the plurality of nodes of the reference mesh and in particular identification of the nasion node of the reference mesh corresponding to the nasion point region of the reference mesh. A transformation may be performed to project the reference mesh and the virtual frame onto the user mesh (block 740). The transformation may be a rigid transformation including rotation, translation, and scaling to fit at least some of the nodes of the reference mesh to the user mesh. Responsive to determining that the nasion node of the reference mesh is aligned with the corresponding nasion node of the user mesh (block 745), a sizing / fitting image may be output, for example, via an application executed on the computing device. Responsive to determining that the nasion node of the reference mesh is offset from or not aligned with the corresponding nasion node of the user mesh (block 745), the system may shift the placement location of the virtual frame from an initial placement location (where the bridge portion of the virtual frame is positioned at the nasion node of the reference mesh and is offset relative to the nasion node of the user mesh) to an adjusted location (block 755) (where the bridge portion of the virtual frame is positioned to correspond to the nasion node of the user mesh) before outputting the sizing / fitting image (block 750).
[0045] Some examples are provided below.
[0046] Example 1: A computer-implemented method includes: capturing image data of an initial image including the face of the user via an application executed on a computing device operated by the user; generating a user mesh based on the image data, the user mesh representing the face of the user; identifying index nodes in the user mesh, the index nodes corresponding to a set portion of the face of the user captured in the image data; identifying index nodes in a reference mesh, the index nodes of the reference mesh corresponding to a set portion of the reference mesh, the set portion of the reference mesh corresponding to the set portion of the user mesh; positioning a virtual frame of a head-mounted wearable device at a position on the reference mesh corresponding to the index nodes of the reference mesh; projecting the reference mesh and the virtual frame onto the user mesh; and adjusting the position of the virtual frame to correspond to the index nodes of the user mesh.
[0047] Example 2: The computer-implemented method according to Example 1, wherein identifying the index nodes in the user mesh includes identifying a nasion node in the user mesh, the nasion node corresponding to the position of the nasion portion of the face of the user captured in the image data; and identifying the index nodes in the reference mesh includes identifying a nasion node in the reference mesh, the nasion node corresponding to the position of the nasion portion of the face represented by the reference mesh.
[0048] Example 3: The computer-implemented method according to Example 1 or Example 2, wherein projecting the reference mesh and the virtual frame onto the user mesh includes performing a rigid transformation of the reference mesh and the virtual frame onto the user mesh.
[0049] Example 4: The computer-implemented method according to Example 3, wherein the reference mesh includes a plurality of nodes, and wherein performing the rigid transformation includes performing rotation operations, translation operations, and scaling operations on a subgroup of the plurality of nodes of the reference mesh to adapt the reference mesh to the user mesh.
[0050] Example 5: The computer-implemented method according to any one of the foregoing examples, wherein generating the user mesh includes: detecting one or more facial landmarks in the image data; and generating the user mesh based on the one or more facial landmarks by a machine learning model.
[0051] Example 6: The computer-implemented method according to any one of the foregoing examples, further comprising outputting an adapted image, the adapted image including a rendering of the virtual frame superimposed on the initial image of the face of the user generated based on the image data at a position corresponding to the index nodes of the user mesh.
[0052] Example 7: The computer-implemented method as described in Example 6, wherein adjusting the position of the virtual frame includes: comparing the positions of the index nodes of the reference grid with the positions of the index nodes of the user grid, including: detecting the distance between the index nodes in the reference grid and the index nodes in the user grid; determining the corresponding pixel distance between the index nodes of the reference grid and the index nodes of the user grid; and adjusting the position of the virtual frame in the adapted image based on the pixel distance.
[0053] Example 8: The computer-implemented method as described in any of the foregoing examples, wherein capturing the image data includes capturing a two-dimensional image of the face of the user; and wherein the user grid is a three-dimensional grid corresponding to the face of the user, and the reference grid is a three-dimensional grid generated based on previously collected data representing multiple subjects.
[0054] Example 9: The computer-implemented method as described in any of the foregoing examples, further comprising selecting a reference grid from a plurality of reference grids, including: detecting at least one facial landmark in the image data; mapping the at least one facial landmark to a corresponding node of the user grid; and selecting the reference grid from the plurality of reference grids based on the relative positions of the index nodes of the user grid and the nodes of the user grid corresponding to the at least one facial landmark.
[0055] Example 10: A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of a computing device, are configured to cause the at least one processor to: capture, via an image sensor of the computing device, image data including an initial image of the face of a user; generate, based on the image data, a user grid representing the face of the user; identify index nodes in the user grid corresponding to a set portion of the face of the user captured in the image data; identify index nodes in a reference grid, the index nodes of the reference grid corresponding to a set portion of the reference grid, the set portion of the reference grid corresponding to the set portion of the user grid; position a virtual frame of a head-mounted wearable device at a position on the reference grid corresponding to the index nodes of the reference grid; project the reference grid and the virtual frame onto the user grid; and adjust the position of the virtual frame to correspond to the index nodes of the user grid.
[0056] Example 11: The non-transitory computer-readable medium as described in Example 10, wherein the instructions cause the at least one processor to: identify the index nodes in the user grid including identifying the nasion node in the user grid, the nasion node corresponding to the position of the nasion part of the face of the user captured in the image data; and identify the index nodes in the reference grid including identifying the nasion node in the reference grid, the nasion node corresponding to the position of the nasion part of the face represented by the reference grid.
[0057] Example 12: The non-transitory computer-readable medium as described in Example 10 or Example 11, wherein the instructions cause the at least one processor to: perform a rigid transformation of the reference grid and the virtual frame onto the user grid to project the reference grid and the virtual frame onto the user grid.
[0058] Example 13: The non-transitory computer-readable medium as described in Example 12, wherein the reference grid includes a plurality of nodes, and wherein the instructions cause the at least one processor to perform the rigid transformation including a rotation operation, a translation operation, and a scaling operation on a subgroup of the plurality of nodes of the reference grid to adapt the reference grid to the user grid.
[0059] Example 14: The non-transitory computer-readable medium as described in any one of Examples 10 to 13, wherein the instructions cause the at least one processor to: detect one or more facial landmarks in the image data; and generate the user grid based on the one or more facial landmarks by a machine learning model.
[0060] Example 15: The non-transitory computer-readable medium as described in any one of Examples 10 to 14, wherein the instructions cause the at least one processor to: output an adapted image, the adapted image including a rendering of the virtual frame superimposed on the initial image of the face of the user generated based on the image data at positions corresponding to the index nodes of the user grid.
[0061] Example 16: The non-transitory computer-readable medium as described in Example 15, wherein the instructions cause the at least one processor to: compare the positions of the index nodes of the reference grid with the positions of the index nodes of the user grid, including: detecting a distance between the index nodes in the reference grid and the index nodes in the user grid; determining a corresponding pixel distance between the index nodes of the reference grid and the index nodes of the user grid; and adjusting the position of the virtual frame in the adapted image based on the pixel distance.
[0062] Example 17: The non-transitory computer-readable medium according to any one of Examples 10 to 16, wherein the instructions cause the at least one processor to capture a two-dimensional image of the user's face, and wherein the user mesh is a three-dimensional mesh corresponding to the user's face, and the reference mesh is a three-dimensional mesh generated based on previously collected data representing a plurality of subjects.
[0063] Example 18: The non-transitory computer-readable medium according to any one of Examples 10 to 17, wherein the instructions cause the at least one processor to select a reference mesh from a plurality of reference meshes, including: detecting at least one facial landmark in the image data; mapping the at least one facial landmark to a corresponding node of the user mesh; and selecting the reference mesh from the plurality of reference meshes based on the relative positions of the indexed nodes of the user mesh and the nodes of the user mesh corresponding to the at least one facial landmark.
[0064] Example 19: A system, the system comprising a computing device, the computing device including: an image sensor; at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: capture image data including an initial image of a user's face; generate a user mesh based on the image data, the user mesh representing the user's face; identify a nasion node in the user mesh, the nasion node corresponding to the nasion portion of the user's face captured in the image data; identify a nasion node in a reference mesh, the nasion node of the reference mesh corresponding to the nasion portion of the reference mesh, the nasion portion of the reference mesh corresponding to the nasion portion of the user mesh; position a virtual frame of a head-mounted wearable device at a position on the reference mesh corresponding to the nasion node of the reference mesh; project the reference mesh and the virtual frame onto the user mesh; and adjust the position of the virtual frame to correspond to the nasion node of the user mesh.
[0065] A variety of implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this specification.
[0066] In addition, the logical flow depicted in the figures does not require a particular order or sequential order as shown to achieve the desired result. Additionally, other steps may be provided, or steps may be deleted from the described flow, and other components may be added to or removed from the described system. Accordingly, other embodiments are within the scope of the appended claims.
[0067] In addition to the above description, controls can be provided to users that allow the users to make choices regarding whether and when the systems, programs, or features described herein may be able to collect user information (e.g., information about a user's social network, social actions, or activities, profession, user preferences, or user's current location) and whether content or communications may be sent to the user from a server. Additionally, certain data can be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, a user's identity can be processed so that the user's personally identifiable information cannot be determined, or the user's geographic location can be generalized (such as to a city, zip code, or state level) when location information is obtained so that the user's specific location cannot be determined. Thus, the user can control what information is collected about the user, how that information is used, and what information is provided to the user.
[0068] Although certain features of the described implementations have been shown as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. Accordingly, it should be understood that the appended claims are intended to cover all such modifications and changes that fall within the scope of the implementations. It should be understood that they are presented by way of example only and not by way of limitation, and various changes in form and detail may be made. Except for mutually exclusive combinations, any part of the apparatus and / or method described herein can be combined in any combination. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described.
Claims
1. A computer-implemented method, comprising: capturing image data of an initial image including the face of the user via an application executed on a computing device operated by the user; generating a user mesh based on the image data, the user mesh representing the face of the user; identifying index nodes in the user mesh, the index nodes corresponding to a set portion of the face of the user captured in the image data; identifying index nodes in a reference mesh, the index nodes of the reference mesh corresponding to a set portion of the reference mesh, the set portion of the reference mesh corresponding to the set portion of the user mesh; positioning a virtual frame of a head-mounted wearable device at a position on the reference mesh corresponding to the index nodes of the reference mesh; projecting the reference mesh and the virtual frame onto the user mesh; and adjusting the position of the virtual frame to correspond to the index nodes of the user mesh.
2. The computer-implemented method according to claim 1, wherein: identifying the index nodes in the user mesh includes identifying a nasion node in the user mesh, the nasion node corresponding to the position of the nasion portion of the face of the user captured in the image data; and identifying the index nodes in the reference mesh includes identifying a nasion node in the reference mesh, the nasion node corresponding to the position of the nasion portion of the face represented by the reference mesh.
3. The computer-implemented method according to claim 1 or 2, wherein, projecting the reference mesh and the virtual frame onto the user mesh includes performing a rigid transformation of the reference mesh and the virtual frame onto the user mesh.
4. The computer-implemented method according to claim 3, wherein, the reference mesh includes a plurality of nodes, and wherein performing the rigid transformation includes performing rotation operations, translation operations, and scaling operations on a subgroup of the plurality of nodes of the reference mesh to adapt the reference mesh to the user mesh.
5. The computer-implemented method according to any one of the preceding claims, wherein, generating the user mesh includes: detecting one or more facial landmarks in the image data; and generating the user mesh based on the one or more facial landmarks by a machine learning model.
6. The computer-implemented method according to any one of the preceding claims, further comprising: outputting an adapted image, the adapted image including a rendering of the virtual frame superimposed on the initial image of the face of the user generated based on the image data at a position corresponding to the index nodes of the user mesh.
7. The computer-implemented method according to claim 6, wherein, adjusting the position of the virtual frame includes: comparing the positions of the index nodes of the reference mesh with the positions of the index nodes of the user mesh, including: detecting the distance between the index nodes in the reference mesh and the index nodes in the user mesh; Determine the corresponding pixel distance between the index nodes of the reference grid and the index nodes of the user grid; and Adjust the position of the virtual frame in the adapted image based on the pixel distance.
8. The computer-implemented method according to any one of the preceding claims, wherein, capturing the image data includes capturing a two-dimensional image of the face of the user; and wherein the user grid is a three-dimensional grid corresponding to the face of the user, and the reference grid is a three-dimensional grid generated based on previously collected data representing multiple subjects.
9. The computer-implemented method according to any one of the preceding claims, further comprising selecting a reference grid from a plurality of reference grids, comprising: detecting at least one facial landmark in the image data; mapping the at least one facial landmark to a corresponding node of the user grid; selecting the reference grid from the plurality of reference grids based on the relative positions of the index nodes of the user grid and the nodes in the user grid corresponding to the at least one facial landmark.
10. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of a computing device, are configured to cause the at least one processor to: Capture image data including an initial image of the face of a user through an image sensor of the computing device; Generate a user grid based on the image data, the user grid representing the face of the user; Identify index nodes in the user grid that correspond to a set portion of the face of the user captured in the image data; Identify index nodes in a reference grid, the index nodes of the reference grid corresponding to a set portion of the reference grid, the set portion of the reference grid corresponding to the set portion of the user grid; Position a virtual frame of a head-mounted wearable device at a position on the reference grid corresponding to the index nodes of the reference grid; Project the reference grid and the virtual frame onto the user grid; and Adjust the position of the virtual frame to correspond to the index nodes of the user grid.
11. The non-transitory computer-readable medium according to claim 10, wherein, the instructions cause the at least one processor to: Identifying the index nodes in the user grid includes identifying the nasion node in the user grid, the nasion node corresponding to the position of the nasion portion of the face of the user captured in the image data; and Identifying the index nodes in the reference grid includes identifying the nasion node in the reference grid, the nasion node corresponding to the position of the nasion portion of the face represented by the reference grid.
12. The non-transitory computer-readable medium according to claim 10 or 11, wherein, the instructions cause the at least one processor to: Perform a rigid transformation of the reference grid and the virtual frame to the user grid to project the reference grid and the virtual frame onto the user grid.
13. The non-transitory computer-readable medium according to claim 12, wherein, the reference grid includes a plurality of nodes, and wherein the instructions cause the at least one processor to perform the rigid transformation including a rotation operation, a translation operation, and a scaling operation on a subgroup of the plurality of nodes of the reference grid to adapt the reference grid to the user grid.
14. The non-transitory computer-readable medium according to any one of claims 10 to 13, wherein, the instructions cause the at least one processor to: detect one or more facial landmarks in the image data; and generate the user grid based on the one or more facial landmarks by a machine learning model.
15. The non-transitory computer-readable medium according to any one of claims 10 to 14, wherein, the instructions cause the at least one processor to: output an adapted image, the adapted image including a rendering of the virtual frame superimposed on the initial image of the face of the user generated based on the image data at positions corresponding to the index nodes of the user grid.
16. The non-transitory computer-readable medium according to claim 15, wherein, the instructions cause the at least one processor to: compare the positions of the index nodes of the reference grid with the positions of the index nodes of the user grid, including: detecting a distance between the index node in the reference grid and the index node in the user grid; determining a corresponding pixel distance between the index node of the reference grid and the index node of the user grid; and adjusting the position of the virtual frame in the adapted image based on the pixel distance.
17. The non-transitory computer-readable medium according to any one of claims 10 to 16, wherein, the instructions cause the at least one processor to capture a two-dimensional image of the face of the user, and wherein the user grid is a three-dimensional grid corresponding to the face of the user, and the reference grid is a three-dimensional grid generated based on previously collected data representing a plurality of subjects.
18. The non-transitory computer-readable medium according to any one of claims 10 to 17, wherein, the instructions cause the at least one processor to select a reference grid from a plurality of reference grids, including: detecting at least one facial landmark in the image data; mapping the at least one facial landmark to a corresponding node of the user grid; and selecting the reference grid from the plurality of reference grids based on the relative positions of the index nodes of the user grid and the nodes in the user grid corresponding to the at least one facial landmark.
19. A system, comprising: a computing device, the computing device including: an image sensor; at least one processor; and a memory that stores instructions, the instructions when executed by the at least one processor cause the at least one processor to: Capture image data of an initial image including a user's face; Generate a user mesh based on the image data, the user mesh representing the user's face; Identify a nasion node in the user mesh, the nasion node corresponding to the nasion portion of the user's face captured in the image data; Identify a nasion node in a reference mesh, the nasion node of the reference mesh corresponding to the nasion portion of the reference mesh, the nasion portion of the reference mesh corresponding to the nasion portion of the user mesh; Position a virtual frame of a head-mounted wearable device at a position on the reference mesh corresponding to the nasion node of the reference mesh; Project the reference mesh and the virtual frame onto the user mesh; and Adjust the position of the virtual frame to correspond to the nasion node of the user mesh.
20. The system according to claim 19, wherein, the reference mesh includes a plurality of nodes, and the user mesh includes a plurality of nodes, and wherein the instructions cause the at least one processor to project the reference mesh and the virtual frame onto the user mesh, including: Perform a rigid transformation of the reference mesh and the virtual frame onto the user mesh, including performing rotation operations, translation operations, and scaling operations on a subgroup of the plurality of nodes of the reference mesh to adapt the reference mesh to the user mesh.