Fitting of an eyeglass frame including a live fitting
The system addresses the limitations of conventional virtual try-on by enabling real-time simulation of eyeglass frames on a user's face, providing an immersive and accurate virtual try-on experience through precise 3D modeling and alignment, overcoming processing delays and occlusion issues.
Patent Information
- Application Number
- JP2022550785
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-21
- Filing Date
- 2021-02-19
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2041-02-19
AI Technical Summary
Conventional virtual try-on technologies fail to provide an experience comparable to actual try-on due to processing delays and technical challenges, making it difficult for consumers to visualize how personal accessories like eyeglasses will look on their person.
A system and method for live fitting of eyeglass frames using client and server components that enable real-time simulation of eyeglass frames on a user's face by capturing images, determining physical characteristics, and generating a 3D model to accurately render the frames in real-time, allowing users to see how they would look from different angles.
Enables a realistic and immediate virtual try-on experience by accurately simulating the placement and movement of eyeglass frames on a user's face, providing a more immersive and accurate representation of how the frames would appear and fit, reducing occlusion errors through precise 3D modeling and alignment.
Smart Images

Figure 0007713949000002 
Figure 0007713949000003 
Figure 0007713949000004
Abstract
Description
Technical Field
[0001] 〔CROSS - REFERENCE TO RELATED APPLICATIONS〕 This application claims priority to U.S. Provisional Patent Application No. 62 / 979,968 (LIVE FITTING OF GLASSES FRAMES, filed Feb. 21, 2020), which is incorporated herein by reference for all purposes.
Background Art
[0002] When making decisions about items such as personal accessories, consumers generally want to visualize how the item will look on the consumer's person. In the real world, consumers would try on the item. For example, a person purchasing glasses has to visit an optician several times to check the fit of the glasses frame and lenses. It would be more convenient if virtual try - on were possible. However, with conventional technology, due to processing delays and other technical challenges, an experience comparable to actual try - on cannot be achieved. It is desirable to enable virtual try - on in a way that is close to the real experience.
Summary of the Invention
[0003] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0004]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Embodiments for Carrying Out the Invention
[0005] The present invention can be implemented in many ways, including a process, an apparatus, a system, a composition of matter, a computer program product embodied on a computer-readable storage medium, and / or a processor, e.g., a processor configured to execute instructions stored on and / or provided thereby a memory coupled to the processor. In this specification, these implementations, or any other form that the present invention may take, may be referred to as techniques. Generally, the order of steps of the disclosed processes may be changed within the scope of the present invention. Unless otherwise specified, components such as processors or memories described as being configured to perform tasks may be implemented as general components temporarily configured to perform the tasks at a given time, or as specific components manufactured to perform the tasks. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data such as computer program instructions.
[0006] A detailed description of one or more embodiments of the present invention is provided below along with the accompanying figures that illustrate the principles of the present invention. The present invention is described in relation to such embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention includes numerous alternative forms, modifications, and equivalents. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. These details are provided for purposes of illustration, and the present invention may be practiced according to the claims without some or all of these specific details. Technical material known in the technical field related to the present invention is not described in detail so that the present invention is not unnecessarily obscured.
[0007] As used herein, the terms "live fitting" or "live try-on" refer to simulating the placement of an object on a person's body by substantially instantaneously displaying a simulation. The terms "video fitting" or "video try-on" refer to simulating the placement of an object on a person's body by displaying the simulation after some delay. An example of live fitting is providing an experience where, at substantially the same timing as a person is looking at a camera, an eyeglass frame is placed on the person's face and the person sees themselves in a mirror while wearing the glasses. An example of video fitting is uploading one or more images of a person's face, determining the placement of glasses on the person's face, and displaying a resulting image or series of images (video) of the glasses placed on the person's face.
[0008] The technology for live fitting of the disclosed eyeglass frames provides the user with the following experience. Since a virtual "mirror" is displayed on the user's electronic device screen, the user can see themselves in the mirror augmented with the selected pair of frames. The user can then sequentially try on various different frames.
[0009] Compared to video try-on technology, "live try-on" or "virtual mirror" style eyeglass try-on renders the selected pair of frames in real time onto an image of the face, allowing the user to see their face more immediately. In various embodiments, the user can immediately see the selected glasses on their face. They can engage in a live experience to see how they would look from different angles by turning their head as desired. When the user moves their head, the glasses move in the same way, simulating how the user would look if they were actually wearing the glasses. Optionally, the user may select a different pair of frames, and the live-rendered image of the user's face will be displayed with the newly selected glasses on.
[0010] While the user is moving their head, this technology collects information regarding the size and shape of the user's face and head. Various visual cues are provided as part of the interface, and the system prompts the user to move in various ways in order to collect the amount of information needed to reach an accurate representation of the user's head and face, including the appropriate scale / size of that representation. Examples of ways to determine the accurate scale / size are further described herein. Visual cues can be provided to indicate how much information has been collected and how much information is still needed. For example, the visual cue indicates when enough information has been obtained to render a high-quality and accurate virtual try-on view.
[0011] Figure 1 is a block diagram showing one embodiment of a system for live fitting of an eyeglass frame. For simplicity, this system is referred to as for live fitting of an eyeglass frame. The data generated by the system can be used in various other applications, including using the live fitting data for video fitting of the eyeglass frame.
[0012] In this example, system 100 includes client device 104, network 106, and server 108. Client device 104 is coupled to server 108 via network 106. Network 106 can include a high-speed data network and / or a telecommunications network. User 102 may interact with the client device to "try on" products, for example, provide a user image of the user's body via the device and display a virtual fitting of the product on the user's body according to the technology further described herein.
[0013] The client device 104 is configured to provide a user interface to the user 102. For example, the client device 104 may receive an input such as an image of the user captured by a camera of the device, or may observe a user interaction between the user 102 and the client device. Based on at least some of the information collected by the client device, a simulation of placing a product on the user's body is output to the user.
[0014] In various embodiments, the client device includes input components such as a camera, a depth sensor, a lidar sensor, other sensors, or a combination of multiple sensors. The camera may be configured to observe and / or capture an image of the user from which physical characteristics can be determined. The user may be instructed to operate the camera or pose for the camera, as further described herein. The information collected by the input components may be used and / or stored for making recommendations.
[0015] Server 108 is configured to output one or more images of the product integrated with the input image, such as determining physical characteristics from the input image, determining the correlation between the physical characteristics and the product, and fitting an eyeglass frame to the user's face. Server 108 may be remote from client device 104 and may be accessible via a network 106 such as the Internet. As further described with respect to FIGS. 2 and 3, various functionality may be embodied either in the client or the server. For example, functions conventionally associated with a server may be performed not only by the server but also / alternatively by the client, and vice versa. The output can be provided to the user with very little latency (if any) after the user provides the input image, such that the user experience is a live fitting of the product, and as a result, the user experience is a live fitting of the product. Virtual fitting of products to the user's face has many applications. For example, there are applications for virtually trying on facial accessories such as eyewear, cosmetics, jewelry, etc. For simplicity, the embodiments herein mainly describe the live fitting of an eyeglass frame to the user's face / head, which is not intended to be limiting, and the present technology may be applied to the fitting of other types of accessories and may also be applied to video fitting (e.g., with some latency).
[0016] FIG. 2 is a block diagram showing an embodiment of a client device for virtual fitting of an eyeglass frame. In some embodiments, client device 104 of FIG. 1 is implemented using the example of FIG. 2.
[0017] In this example, the client device includes an image storage device (storage) 202, a glasses frame information storage device 204, a 3D model storage device 214, a coarse model generator 206, a fitting engine 216, and a rendering engine 212. The client device may be implemented with additional, different, and / or fewer components than those shown in the example. Each of the image storage device 202, the glasses frame information storage device 204, and the 3D model storage device 214 may be implemented using one or more types of storage media. Each of the model generator 206, the fitting engine 216, and the rendering engine 212 may be implemented using hardware and / or software.
[0018] The image storage device 202 is configured to store a set of images. The images include, but are not limited to, RGB images and depth images, and may be in various formats or different types. In some embodiments, each set of images is associated with a recorded video or a series of snapshots of the user's face in various orientations, as further described with respect to FIG. 4. In some embodiments, each set of images is stored together with data associated with the set as a whole or with individual images of the set. In various embodiments, at least a subset of the user images is stored locally and / or remotely and may be transmitted, for example, to the server 108 for storage.
[0019] The camera 218 is configured to capture an image of the user. The captured image is stored at 202 and can be used to determine physical characteristics. As described with respect to FIG. 1, the camera can have various sensors such as a depth sensor that helps generate a model of the user's head. An example of a camera with a depth sensor is the True Depth (registered trademark) camera available in some iPhones (registered trademark). Depending on the camera hardware, it can capture various formats and types of images and data, including but not limited to RGB images and depth images.
[0020] An image can have associated intrinsic and / or extrinsic information. The intrinsic and extrinsic information can be generated by a third party (e.g., a client device application) or can be generated as further described with respect to FIG. 3. In various embodiments, the intrinsic and extrinsic information provided by a third party can be further processed using the techniques described with respect to 308 and 310 of FIG. 3. The information can be generated locally on the device or remotely by a server.
[0021] The coarse model generator 206 is configured to determine a mathematical 3D model of the user's face associated with each set of images. The coarse model generator can be implemented using a third-party mesh model such as one native to the mobile device. The model I / O framework available in iOS (registered trademark) is one such example. In various embodiments, the model can be obtained from the remote server 108 instead of generating the model locally or to complement local model information. The model generator is referred to here as a "coarse" model generator to distinguish it from that shown in FIG. 3, but the model generator 206 can be configured to generate a model having at least the same granularity as that generated by the model generator 306, depending on the techniques used and the available processing resources.
[0022] The fitting engine 216 (also sometimes referred to as a comparison engine) is configured to determine a fit between a 3D model of a user's face (stored, for example, in a 3D model storage device) and a 3D model of an eyeglass frame. In some embodiments, the fitting engine processes a rough model of the user's face. For example, the rough 3D model provides instructions (e.g., suggestions or cues) for automatically placing objects or features such as hats, glasses, facial hair, etc. on the rough model. The placement can be improved by determining additional landmarks. For example, if the rough model lacks the junction points of the ears, the fitting engine can determine those points, as further described with respect to 306 of FIG. 3.
[0023] The eyeglass frame information storage device 204 is configured to store information related to various eyeglass frames. For example, the information related to an eyeglass frame may include measurements of various areas of the frame (e.g., bridge length, lens diameter, temple distance), renderings of the eyeglass frame corresponding to various (R,t) pairs, mathematical representations of the 3D model of the eyeglass frame that can be used to render an eyeglass image for various (R,t) parameters, price, identifier, model number, category, type, eyeglass frame material, brand, part number. In some embodiments, the 3D model of each eyeglass frame includes, for example, a set of 3D points defining various positions / parts of the eyeglass frame, including one or more of a pair of bridge points and a pair of temple bend points. In various embodiments, a 2D image of the eyeglasses is generated on a client device. In other embodiments, a 2D image of the eyeglasses is generated by a server such as 108 of FIG. 1 and transmitted to the client device.
[0024] The rendering engine 212 is configured to render a 2D image of a glasses frame that is overlaid on an image. For example, the selected glasses frame may be a glasses frame for which information is stored in the glasses frame information storage device 204. For example, the image on which the glasses frame is to be overlaid may be stored as part of a set of images stored in the image storage device 202. In some embodiments, the rendering engine 212 is configured to render a glasses frame (e.g., selected by the user) for each of at least one subset of the set of images. In various embodiments, the image on which the glasses frame is overlaid is supplied from a camera. In some embodiments, the rendering engine 212 is configured to transform a 3D model of the glasses frame after the glasses frame is placed on a 3D face (e.g., a 3D model of the user's face or a 3D model of another 3D face) using external information such as an (R,t) pair corresponding to the image. The (R,t) pair is an example of external information determined for an image of a set of images associated with the user's face, where R is a rotation matrix and t is a translation vector corresponding to the image as further described with respect to 308. In some embodiments, the rendering engine 212 is also configured to perform occlusion culling on the glasses frame transformed using the occlusion body. The occlusion glasses frame in the orientation and translation (parallel movement) associated with the (R,t) pair excludes certain portions hidden from view by the occlusion body in that orientation / translation. The glasses frame rendered on the image should indicate the glasses frame in the orientation and translation corresponding to the image and may be overlaid on that image in the playback of the set of images to the user on the client device.
[0025] FIG. 3 is a block diagram showing one embodiment of a server for virtual fitting of an eyeglass frame. In some embodiments, server 108 of FIG. 1 is implemented using the example of FIG. 3. In this example, the server includes an image storage device 302, an eyeglass frame information storage device 304, a 3D model storage device 314, a model generator 306, a fitting engine 316, an external information generator 308, an internal information generator 310, and a rendering engine 312. The server may be implemented with additional, different, and / or fewer components than those illustrated. The functions described with respect to client 200 and server 300 may be embodied in any device. For example, the rough model generated by 206 may be processed (e.g., refined) locally on the client or sent to server 300 for further processing. Each of image storage device 302, eyeglass frame information storage device 304, and 3D model storage device 314 may be implemented using one or more types of storage media. Each of model generator 306, fitting engine 316, external information generator 308, internal information generator 310, and rendering engine 312 may be implemented using hardware and / or software. Each of the components is similar to their counterparts in FIG. 2 unless otherwise described.
[0026] The model generator 306 is configured to determine a mathematical 3D model of the user's face associated with each set of images. The model generator 306 can be configured to generate a 3D model from scratch or based on a rough model generated by the model generator 206. In various embodiments, the model generator is configured to execute the process of FIG. 9 to generate a 3D model. For example, a mathematical 3D model of the user's face (i.e., a mathematical model of the user's face in 3D space) can be set at the origin. In some embodiments, the 3D model of the user's face includes a set of points in 3D space that defines a set of reference points associated with features (e.g., their positions) of the user's face from the associated set of images. Examples of reference points include the endpoints of the user's eyes, the endpoints of the user's eyebrows, the user's nose bridge, the user's ear junctions, and the tip of the user's nose. In some embodiments, the mathematical 3D model determined for the user's face is called an M matrix that is determined based on a set of reference points associated with features of the user's face from the associated set of images. In some embodiments, the model generator 306 is configured to store the M matrix determined for the set of images together with the set in the image storage device 302. In some embodiments, the model generator 306 is configured to store the 3D model of the user's face in the 3D model storage device 314. The model generator 306 can be configured to execute the process of FIG. 9.
[0027] The external information generator 308 and the internal information generator 310 are configured to generate information that can be used for live try-on or video try-on. As described with respect to FIG. 2, the information may be obtained from a third party, the information can be generated based on third party information, or may be generated as follows.
[0028] The external information generator 308 is configured to determine a set of external information for each of at least one subset of the image set. For example, the image set may be stored in the image storage device 302. In various embodiments, the set of external information corresponding to the images of the image set describes one or more of the orientation and translation of the 3D model of the user's face that is determined for the image set required to result in the correct appearance of the user's face in that particular image. In some embodiments, the set of external information determined for the images of the image set related to the user's face is referred to as an (R, t) pair, where R is a rotation matrix and t is a translation vector corresponding to that image. Thus, the (R, t) pair corresponding to the image of the image set can transform the M matrix (representing the 3D of the user's face) (R×M + 1) corresponding to the image set into the appropriate orientation and translation of the user's face shown in the image related to that (R, t) pair. In some embodiments, the external information generator 308 is configured to store the (R, t) pairs determined for each of at least one subset of the image set together with the set in the image storage device 302.
[0029] The internal information generator 310 is configured to generate an internal information set of the camera related to the recording of the image set. For example, the camera is used to record the image set stored in the image storage device 302. In various embodiments, the internal information set corresponding to the camera describes a set of parameters associated with the camera. For example, the brand or type of the camera can be transmitted by the device 200. As another example, the parameters related to the camera include the focal length. In some embodiments, the set of internal information related to the camera correlates points on the scaling reference object between different images of the user with the scaling reference object in the image, and finds the internal information set representing the internal parameters of the camera by using camera calibration techniques. In some embodiments, the set of internal information related to the camera is obtained by using an auto-calibration technique that does not require a scaling reference. In some embodiments, the internal information set associated with the camera is called the I matrix. In some embodiments, the I matrix projects the 3D model version of the user's face transformed by the (R, t) pair corresponding to a specific image onto the 2D surface of the focal plane of the camera. In other words, I×(R×M + 1) projects the 3D model in the orientation and translation determined by the M matrix and the (R, t) pair corresponding to the image onto the 2D surface. The projection onto the 2D surface is the view of the user's face as seen from the camera. In some embodiments, the internal information generator 310 is configured to store in the image storage device 302 the I matrix determined for the camera associated with the image set.
[0030] In some embodiments, the fitting engine 316 is configured to determine a set of calculated bridge points included in a "desired glasses" 3D point set associated with a particular user. In various embodiments, the "desired glasses" 3D point set associated with a particular user includes markers that can be used to determine a desired alignment or fit between a 3D model of an eyeglass frame and a 3D model of the user's face. In some embodiments, when determining the set of calculated bridge points, the fitting engine 316 is configured to determine a plane in 3D space using at least three points from a set of 3D points included in the 3D model of the user's face. For example, the plane is determined using the inner corners of the two eyes and the two ear attachment points from the 3D model of the user's face. The fitting engine 316 is configured to determine a vector parallel to the plane, which may be referred to as the "face normal". The distance between the midpoint of the two inner brow points and the midpoint of the inner corners of the two eyes along the face normal is calculated and may be referred to as the "brow z delta". The fitting engine 316 is configured to determine a "bridge shift" value by multiplying the brow z delta by a predetermined coefficient. For example, the coefficient is close to 1.0 and is calculated heuristically. The fitting engine 316 is configured to determine the set of calculated bridge points by moving each of the inner corners of the two eyes of the 3D model of the user's face towards the camera in the direction of the face normal by the bridge shift value. In some embodiments, the fitting engine 316 is also configured to determine a vertical shift that is determined as a function of the distance between the midpoint of the two inner brow points and the midpoint of the inner corners of the two eyes and a predetermined coefficient.
[0031] In some embodiments, the set of calculated bridge points is further moved along the distance between the midpoint of the two inner brow points and the midpoint of the two inner corners of the eyes, based on a vertical shift. In some embodiments, the other 3D points included in the ideal glasses 3D point set are two temple flexion points, and the fitting engine 316 is configured to be set equal to the two ear attachment points of the 3D model of the user's face. In some embodiments, the initial placement of the 3D model of the glasses frame with respect to the 3D model of the user's face can be determined using the two bridge points and / or the two temple flexion points of the ideal glasses 3D point set. In some embodiments, the fitting engine 316 is configured to determine the initial placement by aligning the line between the bridge points of the 3D model of the glasses frame with the line between the calculated bridge points of the ideal glasses 3D point set associated with the user. Next, the bridge points of the 3D model of the glasses frame are positioned by the fitting engine 316 such that the midpoints of both the bridge points of the 3D model of the glasses frame and the calculated bridge points of the ideal glasses 3D point set associated with the user are in the same position or within a predetermined distance of each other. Then, the bridge points of the 3D model of the glasses frame are fixed, and the temple flexion points of the 3D model of the glasses frame are rotated about the overlapping bridge line that serves as an axis such that the temple flexion points of the 3D model of the glasses frame coincide with or are within a predetermined distance of the ear attachment points of the 3D model of the user's face. As described above, in some embodiments, the ear attachment points of the 3D model of the user's face may be referred to as the temple flexion points of the ideal glasses 3D point set associated with the user.
[0032] In some embodiments, after or alternatively to determining an initial placement of the 3D model of the eyeglass frame relative to the 3D model of the user's face, the fitting engine 316 is configured to determine a set of nose curve points in 3D space associated with the user. The set of nose curve points associated with the user may be used to determine the placement of the 3D model of the eyeglass frame relative to the 3D model of the user's face or to modify an initial placement of the 3D model of the eyeglass frame relative to the 3D model of the user's face determined using an ideal eyeglass 3D point set. In some embodiments, the fitting engine 316 is configured to determine a set of nose curve points in 3D space by morphing a predetermined 3D face to correspond to the 3D model of the user's face. In some embodiments, the predetermined 3D face includes a 3D model of a generic face. In some embodiments, the predetermined 3D face includes a predetermined set of points along a nose curve. In some embodiments, morphing the predetermined 3D face to correspond to the 3D model of the user's face includes moving corresponding positions / vertices (and their respective neighboring vertices) of the predetermined 3D face to coincide with or come closer to corresponding positions on the 3D model of the user's face. After the predetermined 3D face has been morphed, the predetermined set of points along the nose curve is also moved as a result of the morphing. Thus, after the predetermined 3D face has been morphed, the updated positions in 3D space of the predetermined set of points along the nose curve of the predetermined 3D face are referred to as a morphed set of 3D points of the morphed nose curve associated with the user.
[0033] In some embodiments, regions / features such as the nasal curve can be determined from a 3D face model or a rough model using 3D points (also called markers or vertices) and fitting the region to a set of vertices as follows. Typically, the ordering of the vertex indices of the rough head model and the 3D head model is fixed. In other words, the fitting engine can pre-record which vertices approximately correspond to regions such as the nasal curve. These vertices may change slightly during model generation and may not line up neatly on the curve. One approach to generating the nasal curve is to generate a 3D point set by selecting pre-recorded vertices on the head mesh. A plane can then be fitted to these 3D points. In other words, the fitting engine finds the plane that best approximates the space covered by these 3D points. The fitting engine then determines the projection of these points onto that plane. This results in a clean and accurate nasal curve that can be used during fitting.
[0034] In some embodiments, the fitting engine 316 is configured to modify an initial placement of a 3D model of an eyeglass frame relative to a 3D model of a user's face by determining a segment between two adjacent points of a morphed set of nose curvature points associated with the user that is closest to a bridge point of the 3D model of the eyeglass frame and calculating a normal of this segment (which may also be referred to as a "nose curvature normal"). The fitting engine 316 is then configured to position the 3D model of the eyeglass frame along the nose curvature normal toward this segment until the bridge point of the 3D model of the eyeglass frame is within a predetermined distance of the segment. In some embodiments, the fitting engine 316 is further configured to bend a temple bend point of the 3D model of the eyeglass frame to align with an ear attachment point of the 3D model of the user's face.
[0035] FIG. 4 is a flowchart showing an embodiment of a process for trying on glasses. This process can be implemented by system 100. This process can be executed for live try-on or video try-on.
[0036] In the illustrated example, the process begins by obtaining a set of images of the user's head (402). For example, when the user turns on the device's camera, the user's face is displayed on the screen as a virtual mirror. By collecting images of the user's face from various different angles, input is provided for reconstructing a 3D model of the user's head and face.
[0037] In some embodiments, the user can select a combination of eyeglass frames to be displayed on their face. The user can move their head and face, and the position and orientation of the glasses are continuously updated to track the movement of the face and stay in an appropriate position relative to the face movement.
[0038] The process determines an initial orientation of the user's head (404). For example, the process can determine whether the user's head is tilted, facing forward, etc. The orientation of the user's head can be determined in various ways. For example, the process can use the set of images of the user's head to determine a set of face landmarks and use the landmarks to determine the orientation. As further explained with respect to FIGS. 2 and 3, the landmarks can be features of the face such as bridge points, eye corners, ear junctions, etc. As another example, the orientation can be determined by using depth images and pose information provided by a third party (e.g., ARKit) as explained with respect to FIG. 9.
[0039] The process obtains an initial model of the user's head (406). The initial / default model of the user's head may be generated in various ways. For example, the model can be obtained from a third party (e.g., the rough model described with respect to FIG. 2). As another example, the model may be obtained from a server (e.g., the 3D model described with respect to FIG. 3). As yet another example, the model may be generated based on the face of a past user. The face of a "past" user may be a statistical model generated from images stored within a predetermined period (e.g., the most recent face is within the last few hours, days, weeks, etc.).
[0040] The accuracy of the 3D model is improved by additional information collected in the form of 2D color images of the face from different angles and corresponding depth sensor information of the face. In various embodiments, the process instructs the user to turn their head from left to right as a way to obtain sufficient information to construct a satisfactory 3D model. As a non-limiting example, about 10 frames are sufficient to create a 3D model of the required accuracy. Higher accuracy in the shape of the 3D model enables a better fitting of the glasses to the user's head, including the position and angle of the glasses in 3D space. Higher accuracy also enables a better analysis of which parts of the glasses are visible when worn on the face / head and which parts of the glasses are obscured by face features such as the nose or ears.
[0041] Greater accuracy in the 3D model also contributes to more accurate measurements of the user's face and head when the "scale" in the 3D model is established by determining the scale using the process of FIG. 8B or FIG. 9. As further described herein, the process of FIG. 8B determines a measure of the distance between several pairs of points on the face / head. This first distance to be measured is typically the distance between the pupils in a front view of the face, which is known as the interpupillary distance or PD. By knowing the distance between these two points in the virtual 3D space, the distance between any other two points in that space can be calculated. Other measures of interest that can also be calculated include the width of the face, the width of the bridge of the nose, the distance between the center of the nose and each pupil (dual PD), etc.
[0042] Additional measurements can also be calculated that include the scaled 3D head and a 3D model of glasses (having a known scale) fitted to the head. The temple length (the distance from the hinge of the temple arm of the glasses to the ear junction point where the temple is placed) is one example.
[0043] The process transforms (408) an initial model of the user's head corresponding to the initial orientation. For example, the model of the head can be rotated within the 3D space to correspond to the initial orientation.
[0044] The process receives (410) a user selection of a glasses frame. The user may provide the selection via a user interface by selecting a particular glasses frame from several selections, as further described with respect to FIG. 14.
[0045] The process combines the transformed model of the user's head with the model of the eyeglass frame (412). The combination of the user's head and the eyeglass frame provides an accurate representation of how the eyeglasses would look on the user's head, including a realistic visualization of the scale and placement of the glasses on facial landmarks such as the bridge of the nose and the temples. Compared to the prior art, this combination is more realistic as it reduces / eliminates inaccurate occlusion. Occlusion is when a foreground object in 3D space is in front of a background object, causing the foreground object in the image to hide the background object. Correct occlusion more realistically represents the glasses worn on the head by appropriately hiding parts of the glasses behind parts of the face. The cause of inaccurate occlusion is, among other things, when the user's head model or the eyeglass frame model is not accurately combined, or when the head pose is not accurately determined for a particular image (when the external (R,t) is not accurate), especially when the head model of the part where the glasses cross or contact the head is inaccurate. Since the fitting of the glasses to the head depends on the accuracy of the 3D head model, an inaccurate head model will result in an inaccurate fitting. Therefore, a better head model will achieve more accurate occlusion and provide a more realistic try-on experience for the user.
[0046] The process generates an image of the eyeglass frame (414) at least partially based on the combination of the transformed model of the user's head and the model of the eyeglass frame. Since the image of the eyeglass frame is 2D, it can be presented later on a 2D image of the user's head. In various embodiments, the image of the eyeglass frame can be updated in response to certain conditions being met. For example, if features of the user's face are not covered by the initial 2D image of the eyeglass frame due to an inaccurate scale, the initial 2D image can be changed to stretch and expand the frame in one or more dimensions to reflect a more accurate scale as the model is improved.
[0047] The process provides a presentation (416) that includes overlaying an image of a pair of glasses on at least one image of a set of images of the user's head. The presentation can be output on a user interface such as those shown in FIGS. 10-15.
[0048] FIG. 5 is a flowchart showing one embodiment of a process for acquiring an image of a user's head. This process can be executed as part of another process such as 402 of FIG. 4. This process can be implemented by system 100.
[0049] In the illustrated example, the process begins by receiving a set of images of the user's head (502). In various embodiments, the user may be instructed to move their head to acquire desired images. For example, the user may be instructed via the user interface to take a forward-facing image and then rotate their head to the left, then to the right, or to slowly rotate their head from one direction to another. If the user's movement is too fast or too slow, the user may be instructed to slow down or speed up.
[0050] The process stores a set of images of the user's head and associated information (504). The associated information may include sensor data such as depth data. The images and associated information may later be used to construct a model of the user's head or for other purposes, as further described herein.
[0051] The process determines whether to stop (506). For example, if a sufficient number of images (number, image quality, etc.) have been taken, the process determines that the stop condition has been met. If the stop condition is not met, the process returns to 502 to receive additional images. Otherwise, if the stop condition is met, the process ends.
[0052] Some examples of user interfaces for acquiring an image of a user's head are shown in FIGS. 10-12.
[0053] Figure 6 is a flowchart showing one embodiment of a process for live fitting of glasses. This process can be executed as part of another process such as Figure 4. This process can be implemented by system 100. In the illustrated example, the process starts by determining an event related to updating the current model of the user's face (602).
[0054] The process updates the current model of the user's face in response to the event using a set of historical frames of the user's face (604). For example, the set of historical frames of the user's face can be those obtained at 402 in Figure 4, or images obtained prior to the current recording session.
[0055] The process obtains a newly recorded frame of the user's face (606). The process can obtain the newly recorded frame by instructing the camera on the device to capture an image of the user. Feedback can be provided to the user via a user interface as shown in Figures 10 - 12, instructing the user to move their head to capture the desired image.
[0056] The process generates a corresponding image of the eyeglass frame using the current model of the user's face (608). The exemplary process is further described in Figures 8A, 8B, and 9.
[0057] The process presents an image of the eyeglass frame over the newly recorded frame of the user's face (610). An example of presenting the image is 416 in Figure 4.
[0058] The current model of the user's face can be updated when new information is available, such as new facial landmarks or depth sensor data associated with recent historical images. In various embodiments, a certain number of poses are required to generate a model of a desired density or accuracy. However, a user may sometimes turn their head too quickly and not fully capture a pose. If a pose is not fully captured, the user is prompted to return to a position where a pose can be captured. As will be further described in connection with FIG. 12, feedback can be provided on a GUI or in another form (e.g., sound or haptic feedback) to prompt the user to face the desired direction.
[0059] 7 is a flow chart illustrating one embodiment of a process for generating a corresponding image of an eyeglass frame. This process may be performed as part of another process, such as 608 in FIG. 6. This process may be performed by system 100. For example, this process may be performed if a user provides additional images after an initial model of the user's face has been formed.
[0060] In the illustrated example, the process begins by obtaining a current orientation of the user's face (702). An example of determining the current orientation is 404. In various embodiments, the current orientation can be determined based on a newly recorded frame of the user's face that includes depth sensor data. In various embodiments, the orientation can be obtained from the device. An orientation provided by the device or a third party can be used directly or further processed to improve the orientation.
[0061] The process transforms the current model of the user's face to correspond to the current orientation (704). The 3D model of the user's face can be oriented to correspond to the current orientation. Scaling can be performed to efficiently and accurately transform the current model of the user's face, as further described herein with respect to Figures 8A, 8B, and 9.
[0062] The process combines the transformed model of the user's face with the model of the eyeglass frame (706). An example of combining the transformed head model with the model of the eyeglass frame is 412.
[0063] The process generates a current image of the eyeglass frame based at least in part on the combination (708). An example of combining the transformed head model with the model of the eyeglass frame is 414.
[0064] The process generates a current image of the eyeglass frame based at least in part on the combination of the transformed model of the head and the model of the eyeglass frame (708). In various embodiments, the current image of the eyeglass frame is a 2D image suitable for being displayed on the user's head to indicate that the user is trying on the eyeglass frame. The 2D image can be generated such that artifacts and occlusions are removed when combined with the user's face.
[0065] The following figures (Figures 8A, 8B, and 9) show some examples of determining scale using either a relatively coarse head model or a relatively detailed head model.
[0066] Figure 8A is a flowchart illustrating one embodiment of a process for scaling a head model using a relatively coarse model. This process can be executed as part of another process such as 704 of FIG. 7. This process can be implemented by the system 100 using a coarse model such as the model generated by 206.
[0067] In various embodiments, the true scale of the user's face and PD (pupillary distance) can be determined. For example, the true scale and PD can be determined on an iOS (registered trademark) device from one or more RGB camera images, one or more true depth images, and 3D geometry provided by ARKit. The same concept can also be applied to Android (registered trademark) devices or other platforms where depth images and calibration information of 3D geometry can be obtained.
[0068] The process starts by receiving a two-dimensional (2D) RGB image and a depth image (802). The 2D RGB can be included in a set of RGB images of the user's head. An example of a 3D depth image is a true depth image. In various embodiments, the process obtains the 2D RGB image and / or the depth image via an API.
[0069] When a 3D model of an object is given, the model space coordinates of the object can be mapped to 2D image space. An example of the mapping is as follows. [x,y,1] T =P * V * M * [X,Y,Z,1] T Here, x and y are 2D coordinates, P, V, and M are a projection matrix, a view matrix, and a model matrix respectively, and X, Y, and Z are 3D model space coordinates.
[0070] The model matrix moves coordinates in the (scale - less) model space to the real - world coordinate system. Then, the view matrix provides translation and rotation operations so that the object is represented in the camera coordinate system. When face tracking is turned on, ARKit provides a representation of the face in the model space where the face is represented by a low - resolution inaccurate mesh (a few vertices). Further, a P matrix, a V matrix, and an M matrix are also provided, and thus, a mapping between pixel coordinates and model mesh vertices can be obtained. Assuming the P matrix (obtained from the focal length and the optical center), any point on the image can be represented in the camera coordinate system (real world dimensions) if the depth information of that point is available. In the case of a device with a depth sensor, calibration is performed in various embodiments such that the depth image is registered to the RGB image only by a difference in resolution. In some embodiments, there is no difference in resolution.
[0071] The process finds (804) coordinates associated with 2D features in the 2D RGB image. An example of a 2D feature is the iris of the eye, and thus, the coordinates are iris coordinates. The 2D feature coordinates can be found using machine learning. The iris coordinates in 2D can be used to determine the iris points in 3D, and the distance between the iris points gives the inter - pupil distance. In various embodiments, using an example of ARKit by Apple®, the iris coordinates (x, y, z) are determined for each of the left and right eyes by using ARKit. This can be determined from the device sensor information.
[0072] The process uses the resolution mapping between the 2D RGB image and the depth image and the determined 2D feature coordinates to determine (806) 3D feature coordinates in the depth image. When the iris points are determined on the RGB image, depth information can be obtained from the depth image (and in some cases, some additional processing is done in the vicinity of the iris points), and this depth value is combined with the focal length and optical center information from the projection matrix to represent the iris points in 3D coordinates in real - world dimensions.
[0073] The projection matrix has the following format (which can be obtained from ARKit).
[0074] [Number]
[0075] Given the projection matrix, the iris coordinates in the depth image, the depth value, and the depth image resolution, the iris points can be represented in a three-dimensional coordinate system with real-world dimensions using the following formula. a = depth_projection_matrix[0,0] c = depth_projection_matrix[0,2] f = depth_projection_matrix[1,1] g = depth_projection_matrix[1,2] m = depth_projection_matrix[2,2] n = depth_projection_matrix[2,3] z = depth_image[int(marker_depth_coordinates[0]), int(marker_depth_coordinates[1])] H, W = height of the depth image, the following depth image Y_clip = 1 - (int(marker_depth_coordinates[0]) / (H / 2.0)) X_clip = (int(marker_depth_coordinates[1]) / (W / 2.0)) - 1 Z = -z X = (X_clip * (-Z) - c * Z) / a Y = (y_clip * (-Z) - g * Z) / f
[0076] The process determines the feature pair distance of actual dimensions in the 2D space using the 3D feature coordinates (808). For example, the true PD can be determined using the 3D feature coordinates. The actual dimensions are useful for accurately indicating the arrangement of the glasses frame on the user's head.
[0077] FIG. 8B is a flowchart illustrating an embodiment of a process for scaling a head model using a relatively fine model. This process can be executed as part of another process such as 704 of FIG. 7. This process can be implemented by system 100 using a coarse model such as that generated by 306(206). Compared to the process of FIG. 8A, a more accurate scale can be determined, but it may require more computational power.
[0078] Given a head rotation sequence (RGB images) and a single scale image (true depth and RGB), a scale for 3D reconstruction of the head (the accurate high-resolution mesh is also referred to as the "Ditto mesh") can be obtained. One approach is to project iris points onto the 3D head and scale the head using the 3D iris distances. However, this uses only two points on the mesh and may not be accurate if there are errors in unprojection or iris detection. Another approach is to use multiple feature points on the face and calculate the pairwise distances obtained through the 3D representation based on the pairwise distances on the unprojection (positions on the 3D Ditto mesh) and ARKit information. The scale ratio of the two distances corresponding to the same pair is expected to be constant across all pairs in various embodiments. This scale ratio can be estimated by using multiple pairs as follows, or alternatively, by using the process described in FIG. 9.
[0079] The process begins by unprojecting 3D feature coordinates in a depth image onto a 3D head model in order to obtain 3D feature coordinates using external information corresponding to an RGB image (812). Examples of 3D feature coordinates include iris coordinates. Examples of 3D head models include relatively detailed models generated by model generator 306. Although not shown here, the process can receive inputs such as a 3D mesh (a model of the user's head) such as that generated by model generator 306, true depth (e.g., ARKit) information, and / or camera internal information prior to receiving the inputs as in the process of FIG. 8A. For example, using camera extrinsics, the process obtains the left eye within the Ditto mesh (3D head model) using an unprojection defined in 3D coordinates within the Ditto space. Similarly, the right eye can also be obtained.
[0080] The process determines a first feature pair distance using the 3D feature coordinates (814). The first feature pair distance is based on the Ditto mesh. The first feature pair distance is a pairwise distance in the position in the Ditto mesh model of the user's head.
[0081] The process determines a second feature pair distance using the true depth image (816). The second feature pair distance is based on the ARKit information. The second feature pair distance is a pairwise distance obtained from the true depth information.
[0082] The process determines a scale factor as the ratio of the first feature pair distance to the second feature pair distance (818). For example, the first feature pair distance is compared (e.g., divided) by the second feature pair distance to obtain a scale factor (also called a scale ratio). The scale factors are expected to be constant, but if they are not exactly the same, an average can be taken. The scale factor can be used to determine PD and the true scale.
[0083] Described herein is the use of ARKit and depth information to add scale to 3D reconstruction by using single-scale images (RGB + true depth). These concepts can be extended to provide a live try-on experience on true depth devices. Given a depth image, a high-precision mesh of the face can be determined / obtained from the RGB image, projection matrix, and view matrix (camera extrinsics) (e.g., by extending current methods to known methods such as the Ditto 3D reconstruction algorithm). Then, given each new image (RGB and / or depth image + extrinsics), the initial mesh and given extrinsics can be refined to provide the user with an accurate live try-on or video try-on experience.
[0084] FIG. 9 is a flowchart illustrating one embodiment of a process for scaling and generating a head model. This process can be executed as part of another process such as 704 of FIG. 7. This process can be implemented by system 100 using camera information such as the image set and associated information of 504. This process is an alternative to the processes described in FIGS. 8A and 8B. In various embodiments, in addition to the pose information provided by a framework such as ARKit, the captured depth and RGB images are used to generate a life-size head mesh. One advantage is that existing information (pose, rough head model) can be leveraged and built later by incorporating video try-on (offline processing) / improved live try-on. This reduces processing time by eliminating the need to determine camera information and pose information.
[0085] The process starts (902) by receiving one or more RGB images, one or more depth sensor images, pose information, and camera intrinsics. This information can be generated by a device with a depth sensor and via a library provided by a native framework such as ARKit. For example, ARKit provides a core head model and pose information which are extrinsics for the images. A camera with a depth sensor can generate a depth image corresponding to a standard RGB image. Camera intrinsics refers to information such as the focal length.
[0086] The process uses the camera intrinsics to generate 3D points at the actual scale for each point within each depth sensor image (904). The camera intrinsics provides information about the characteristics of the camera, which can be used to map points from the depth sensor image to 3D points at the actual scale. The process of FIG. 8A (or a part thereof) can be applied to generate 3D points by processing all points / pixels in the image (not necessarily just iris points).
[0087] The process uses the pose information to merge the 3D points from the images into a point cloud with the actual scale (906). The point cloud represents a rough region or structure of the 3D head model.
[0088] The process uses past head scans from a storage device to generate a model of the user's face with the actual scale that matches the shape of the point cloud (908). The generated model is clean and accurate for the user's head. To obtain a clean and accurate model of the user's head, scan histories are registered to the 3D point cloud. The head scan history can be a statistical model aggregated by using a set of scan histories.
[0089] Scaling (e.g., as a result of the processes of FIGS. 8A, 8B, or 9) can be used to generate or modify the size of a 3D head model (Ditto mesh) in 3D space. The scaled head can be used to generate a 2D image of a selected eyeglass frame (e.g., as used by 608).
[0090] The following figures show several graphical user interfaces (GUIs). The GUI can be rendered on the display of the client device 104 of FIG. 1 corresponding to various steps of the live fitting process.
[0091] FIG. 10 shows an example of a frame fitting graphical user interface obtained in some embodiments. This GUI conveys how well the eyeglass frame fits the user's face. When the glasses are first extended onto the face (before sufficient image data regarding the user's face is collected), there may not be enough face data collected to accurately determine how a particular pair of glasses fits in various areas (face width, optical center, bridge of the nose, and temples).
[0092] As the user rotates left and right, face data is collected and processed, 3D understanding of the head (construction of a 3D model of the user's face) is performed, and the fit across each part can be accurately evaluated. - As shown in the illustration, the GUI conveys one or more of the following. - Emphasize face elements (bridge, temples, etc.) for fitting. - Indicate whether fitting for a particular face region is being processed and / or indicate the degree of processing completed. - When fitting is processed, a score (e.g., red, yellow, green, here represented by a grayscale gradient) is displayed to indicate the degree to which the glasses fit a particular face element (are suitable).
[0093] FIG. 11 shows an example of a frame scale graphical user interface obtained in some embodiments. This GUI conveys the scale of the eyeglass frame relative to the user's face. When the glasses are first extended over the face, there may not be enough face data collected to accurately determine the scale (the relative size of the frame to the user's face). In various embodiments, the glasses are initially displayed as an "ideal" size such that they appear to fit the user's face (1100), even if the frame is too small or too large. The true scale of the glasses can be determined after additional user face images are acquired (1102). An example of a method for determining the scale / true size of the user's face is described with respect to FIGS. 8A, 8B, and 9. Here, while the user is moving their head from left to right (dashed line), the glasses follow the user's face and the user experience is like looking in a mirror. As the user rotates and more face data is collected, the glasses expand and contract to the correct size and seat / fit more accurately on the face (1104). Here, it has been found that the frame is larger than the initial "ideal" size.
[0094] FIG. 12 shows an example of a desired and captured face angle graphical user interface obtained in some embodiments. This GUI conveys the angle of the face that has been successfully captured and / or the angle of the face that is desired to be captured (e.g., the desired angle has not yet been fully processed). When the glasses are first extended over the face, there may not be enough facial data collected to accurately determine how a particular pair of glasses fits into key areas (the width of the face, the optical center, the bridge of the nose, and the temples). As the user rotates left and right, face data is collected and processed to obtain a 3D understanding of the head (constructing a 3D model of the user's face) so that the fit across the area can be accurately evaluated. A side turn captures a clip (e.g., a video frame), which in turn enables the user to see themselves at key angles when the frame is extended.
[0095] The GUI conveys one or more of the following. - In a first portion, an image of the user's face with the eyeglass frame the user is trying on is displayed. In a second portion (the bottom strip in this example), the captured desired face angle is displayed. - An indicator indicating the desired angle has been captured but is still processing (1200). First, since a forward-facing image is desired, the indicator (circular arrow) indicates that this is the image being captured. When the image is captured, the indicator is replaced with the captured image (1202). - An initial (forward-facing) image of the user without glasses within the strip (1202). - The guidance within the strip prompts the user to rotate left and right, or in one direction or another (e.g., left 1204 - 1208 or right 1210). - An image of the user without glasses when the desired angle is being processed is shown in the strip below 1202 - 1210.
[0096] Figure 13 shows an example of a split-screen graphical user interface obtained in some embodiments. With this GUI, the user can view both live try-on in one part of the screen and video try-on in another part of the screen. In various embodiments, the default display is live try-on (the frame extended onto the face in real time). When the user rotates to capture the required angle, the image is processed for video-based try-on. A split screen is displayed during processing. When the processing is complete, the video try-on becomes visible. The user can drag a slider to switch between live try-on and video try-on.
[0097] Figure 14 shows an example of a graphical user interface for displaying various eyeglass frames obtained in some embodiments. This GUI can be used for live try-on and / or video try-on.
[0098] For example, in the case of live try-on, the initial screen is a live try-on (a frame extended to the face in real time). On the strip, other selected or recommended frames are also displayed as live try-ons. In various embodiments, the main try-on and the strip try-on are the same live feeds but feature different frames. The user can swipe the strip up and down to view different frames.
[0099] For example, in the case of video try-on, the initial screen is a live try-on (a frame extended to the face in real time). On the strip, other selected frames or recommended frames are displayed as video try-ons. When the try-on is processed, the strip appears. Each video try-on can be interacted with independently. The user can swipe the strip up and down to view different frames.
[0100] FIG. 15 is a diagram showing an example of a graphical user interface having an insertion obtained in some embodiments. When glasses are first extended over the face, there may not be enough facial data collected to accurately confirm how a particular pair of glasses fits into the key areas (face width, visual center, bridge of the nose, and temples). As the user rotates left and right, facial data is collected and processed to obtain a 3D understanding of the head (construct a 3D model of the user's face), and the fit across the area can be accurately evaluated.
[0101] In addition, a side turn captures a clip (video frame), which in turn enables the user to see themselves at the key angle when the frame is extended.
[0102] The GUI conveys one or more of the following. - The initial screen is a live try-on (a frame extended to the face in real time) (1500). - The inserted image shows a processed video try-on representing the range in which the required frames have been received and processed. - As the video fitting is processed, the inserted image becomes clearer (progress from 1500 to 1506).
[0103] The technology disclosed herein has many advantages over conventional live fitting products, including the ability to store various images and image sequences from live virtual mirror sessions with different glasses-wearing heads and different poses of the face. The technology disclosed herein provides the ability to create image sequences representing the natural movements of a user wearing different frames. In various embodiments, fitting information, (a series of) images, and the like are saved from the session and used to show the user additional different types of frames even after the live session has ended.
[0104] The technology disclosed herein can be integrated with other types of video fitting for glasses processes / systems. This video fitting approach has proven to be a very useful way for people interested in purchasing new glasses to see how they would look with different pairs of glasses. In this use case, the user records an image, uploads it for analysis, and then the recorded image is saved and used to create a 3D reconstruction of the user's face. These images are saved for later use, and the 3D model of the face is saved for subsequent rendering requests that utilize various different glasses frames according to the user's requests.
[0105] The foregoing embodiments have been described in some detail for purposes of clarity of understanding, but the present invention is not limited to the details provided. There are many alternative ways to implement the present invention. The disclosed embodiments are exemplary and not limiting.
Claims
1. A system, wherein a processor determines an event related to updating a current model of a user's face, in response to the event, updates the current model of the user's face using a set of historical recording frames of the user's face, acquires a newly recorded frame of the user's face, determines a face normal vector using at least three points from the current model of the user's face, determines a bridge shift value based on two eyebrow points and two inner corners of the eyes from the current model of the user's face, determines a set of calculated bridge points based on the face normal vector and the bridge shift value, generates a corresponding image of an eyeglass frame using the current model of the user's face based on the set of calculated bridge points, a processor configured to present an image of the eyeglass frame on the newly recorded frame of the user's face; and a memory coupled to the processor and configured to provide instructions to the processor. A system comprising:
2. The system according to claim 1, wherein the event is at least partially based on an elapsed time, a number of the newly recorded frames satisfying a threshold, or a detected orientation of the user's face.
3. The system according to claim 1, wherein the set of historical recording frames of the user's face includes recording frames within a time threshold.
4. The system according to claim 1, wherein the current model of the user's face is obtained from depth sensor data collected by a device.
5. The system according to claim 1, wherein the current model of the user's face is generated based at least in part on a past user's face.
6. A system further comprising updating the current model of the user's face by the following steps: acquiring a current orientation of the user's face; converting the current model of the user's face corresponding to the current orientation; combining the converted model of the user's face with a model of the eyeglass frame; and generating a current image of the eyeglass frame based at least in part on the combination of the converted model of the user's face and the model of the eyeglass frame. The system according to claim 1, further comprising updating by.
7. The converting the current model of the user's face corresponding to the current orientation includes scaling the user's face by the following steps, a system comprising: Receiving a two-dimensional (2D) RGB image and a depth image; Determining coordinates related to 2D features in the 2D RGB image; Determining 3D feature coordinates in the depth image using a resolution mapping between the 2D RGB image and the coordinates related to the 2D features found in the depth image; Determining an actual-size feature pair distance in 2D space using the 3D feature coordinates; The system according to claim 6, comprising:
8. The converting the current model of the user's face corresponding to the current orientation includes scaling the user's face by the following steps, a system comprising: Receiving a two-dimensional (2D) RGB image and a depth image; Unprojecting 3D feature coordinates in the depth image onto a 3D head model and obtaining 3D feature coordinates using external information corresponding to the RGB image; Determining a first feature pair distance using the 3D feature coordinates; Determining a second feature pair distance using the depth image; Determining a scale factor as a ratio of the first feature pair distance and the second feature pair distance; The system according to claim 6, comprising scaling the user's face thereby.
9. The converting the current model of the user's face corresponding to the current orientation includes scaling the user's face by the following steps, a system comprising: Receiving one or more RGB images, one or more depth sensor images, pose information, and camera intrinsics; For each point in each depth sensor image, generating a 3D point at an actual scale using camera intrinsics; Merging 3D points from the images into a point cloud at an actual scale using pose information; Generating a model of the user's face at an actual scale that matches the shape of the point cloud using a historical head scan from a storage device; generating a corresponding image of the eyeglass frame using the model of the face of the user that has been generated, including scaling the face of the user, the system according to claim 6.
10. A system including a camera configured to record a frame of the face of the user, the camera further including the camera including at least one depth sensor, the system according to claim 1.
11. Presenting the image of the eyeglass frame on the newly recorded frame of the face of the user includes rendering the image of the eyeglass frame on the newly recorded frame of the face of the user within a graphical user interface, the system according to claim 1.
12. Obtaining the newly recorded frame of the face of the user includes presenting feedback to the user regarding at least one of the angle of the captured face and the angle of the face that has not yet been fully processed, the system according to claim 1.
13. The processor is further configured to present information regarding the degree of fit for at least one area of the face of the user, the system according to claim 1.
14. Obtaining the newly recorded frame of the face of the user and presenting the image of the eyeglass frame on the newly recorded frame of the face of the user are performed substantially simultaneously, the system according to claim 1.
15. The processor is further configured to present an image of the eyeglass frame on a previously recorded frame of the face of the user, the system according to claim 1.
16. The image of the eyeglass frame on the previously recorded frame of the face of the user and the image of the eyeglass frame on the newly recorded frame of the face of the user are presented side by side, the system according to claim 15.
17. The processor is receiving a user selection of a specific eyeglass frame from among the selection of eyeglass frames, The system according to claim 1, further configured to present the image of the specific eyeglass frame on the newly recorded frame of the face of the user.
18. The system according to claim 1, wherein the processor is further configured to output a progress status of acquiring the newly recorded frame of the user's face.
19. The processor acquires an image set of a user's head, determines an initial orientation of the user's head, acquires an initial model of the user's head, transforms the initial model of the user's head corresponding to the initial orientation, receives a user selection of an eyeglass frame, determines a face normal vector using at least three points from the transformed model of the user's head, determines a bridge shift value based on two brow points and two inner corners of the eyes from the transformed model of the user's head, determines a set of calculated bridge points based on the face normal vector and the bridge shift value, combines the transformed model of the user's head with the model of the eyeglass frame based on the set of calculated bridge points, generates an image of the eyeglass frame based at least in part on a combination of the transformed model of the user's head and the model of the eyeglass frame, a processor configured to provide a presentation including overlaying an image of the eyeglass frame on at least one image of the image set of the user's head; a system comprising the processor and a memory coupled to the processor and configured to provide instructions to the processor.
20. The processor receives one or more RGB images, one or more depth sensor images, pose information, and camera intrinsics, generates 3D points at actual scale for each point in each depth sensor image using camera intrinsics, merges 3D points from the images into a point cloud at actual scale using pose information, generates a 3D head model at actual scale that matches the shape of the point cloud using a historical head scan from a storage device, determines a face normal vector using at least three points from the 3D head model determines a bridge shift value based on two brow points and two inner corners of the eyes from the 3D head model determines a set of calculated bridge points based on the face normal vector and the bridge shift value A processor configured to generate a corresponding image of an eyeglass frame using the generated 3D head model and the set of calculated bridge points; A system comprising the memory coupled to the processor and configured to provide instructions to the processor. **Claim 21** Determining an event related to updating a current model of a user's face; Updating the current model of the user's face using the set of historical frames of the user's face in response to the event; Obtaining a newly recorded frame of the user's face; Generating a corresponding image of an eyeglass frame using the current model of the user's face; Determining a face normal vector using at least three points from the current model of the user's face; Determining a bridge shift value based on two eyebrow points and two inner corners of the eyes from the current model of the user's face; Determining a set of calculated bridge points based on the face normal vector and the bridge shift value; Presenting an image of the eyeglass frame on the newly recorded frame of the user's face based on the set of calculated bridge points. A method including the above steps. **Claim 22** A computer program product embodied in a non-transitory computer-readable medium, which determines an event related to updating a current model of a user's face, updates the current model of the user's face using the set of historical frames of the user's face in response to the event, obtains a newly recorded frame of the user's face, generates a corresponding image of an eyeglass frame using the current model of the user's face, determines a face normal vector using at least three points from the current model of the user's face, determines a bridge shift value based on two eyebrow points and two inner corners of the eyes from the current model of the user's face, determines a set of calculated bridge points based on the face normal vector and the bridge shift value. A computer program product embodied in a non-transitory computer-readable medium including computer instructions for presenting an image of the eyeglass frame on the newly recorded frame of the user's face based on the set of the calculated bridge points.
Citation Information
Patent Citations
Method and system for creating custom products
JP2016537716A
3D face authentication and recognition based on bilateral symmetry analysis
US20060078172A1
2d image-based 3D glasses virtual try-on system
US20160035133A1
Rendering glasses shadows
US20160217609A1