Head-mounted device and method for adjusting rendering of virtual content items
By adjusting rendering of virtual content items in head-mounted devices using gaze view sensors and processors in head-mounted devices, the problem of changing position of virtual content items in traditional augmented reality head-mounted devices is solved, improving the fidelity of the augmented reality experience.
Patent Information
- Application Number
- CN202080083109.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-13
- Filing Date
- 2020-11-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-11-11
AI Technical Summary
When a traditional augmented reality head-mounted device moves the user's face, the apparent distance and angular position of the virtual content items are prone to frequent changes, resulting in a decrease in fidelity of the enhanced scene and a decrease in user experience.
By equipping the internally oriented gaze view sensor in the head-mounted device, the distance and angle changes between the image rendering device and the user's eyes are determined, and the rendering of the virtual content item is adjusted by the processor to maintain the stable apparent position of the virtual content item relative to the real-world object.
Effectively stabilize the apparent location of virtual content items relative to real-world objects, reduce changes in virtual content items due to head-mounted devices, and improve the fidelity of the augmented reality experience.
Smart Images

Figure CN114761909B_ABST
Abstract
Description
[0001] Priority
[0002] This patent application claims priority to a non - provisional application filed on December 13, 2019, with Serial No. 16 / 713,462 and titled "Content Stabilization for Head - Mounted Displays", which is assigned to the assignee of this application and is hereby incorporated by reference in its entirety. Background Art
[0003] In recent years, augmented reality software applications that combine real - world images from a user's physical environment with computer - generated imagery or virtual objects (VOs) have grown in popularity and use. Augmented reality software applications can add graphics, sound, and / or tactile feedback to the natural world around the user of the application. Images, video streams, and information about people and / or objects can be superimposed on the visual world and presented to the user as an augmented scene on a wearable electronic display or a head - mounted device (e.g., smart glasses, augmented reality glasses, etc.). Summary of the Invention
[0004] Aspects include a head - mounted device for use in an augmented reality system, the head - mounted device being configured to compensate for movement of the device on a user's face. In aspects, a head - mounted device can include: a memory; sensors; and a processor coupled to the memory and the sensors, where the processor can be configured to: receive information from the sensors, where the information can indicate the position of the head - mounted device relative to a reference point on the user's face; and adjust the rendering of virtual content items based on the position.
[0005] In some aspects, the information received from the sensors is related to the position of the head - mounted device relative to the user's eyes, where the processor can be configured to adjust the rendering of virtual content items based on the position of the head - mounted device relative to the user's eyes. In some aspects, the sensors can include infrared (IR) sensors and an IR light source configured to emit IR light towards the user's face. In some aspects, the sensors can include ultrasonic sensors configured to emit ultrasonic pulses towards the user's face and determine the position of the head - mounted device relative to a reference point on the user's face. In some aspects, the sensors can include a first camera, and in some aspects, the sensors can further include a second camera. In some aspects, the head - mounted device can further include: an image rendering device coupled to the processor and configured to render virtual content items.
[0006] In some aspects, the processor may also be configured to: determine an angle from an image rendering device on the head-mounted device to the user's eyes; and adjust the rendering of the virtual content item based on the determined angle to the user's eyes and the determined distance between the head-mounted device and a reference point on the user's face.
[0007] Some aspects may include a method of adjusting the rendering of a virtual content item in an augmented reality system to compensate for movement of the head-mounted device on the user, the method may include: determining the position of the head-mounted device relative to a reference point on the user's face; and adjusting the rendering of the virtual content item based on the determined position of the head-mounted device relative to the reference point on the user's face. In some aspects, the reference point on the user's face may include the user's eyes.
[0008] Some aspects may further include receiving information related to the position of the head-mounted device relative to the user's eyes from a sensor, wherein adjusting the rendering of the virtual content item based on the determined position of the head-mounted device relative to a reference point determined on the user's face may include: adjusting the rendering of the virtual content item based on the distance and angle to the user's eyes determined according to the position of the head-mounted device relative to the user's eyes.
[0009] In some aspects, determining the position of the head-mounted device relative to a reference point on the user's face may include: determining the position of the head-mounted device relative to the reference point on the user's face based on information received from an infrared (IR) sensor and an IR light source configured to emit IR light towards the user's face.
[0010] In some aspects, determining the position of the head-mounted device relative to a reference point on the user's face may include: determining the position of the head-mounted device relative to the reference point on the user's face based on information received from an ultrasonic sensor configured to emit ultrasonic pulses towards the user's face.
[0011] In some aspects, determining the position of the head-mounted device relative to a reference point on the user's face may include: determining a change in the position of the head-mounted device relative to the reference point on the user's face; and adjusting the rendering of the virtual content item based on the determined position of the head-mounted device relative to the reference point on the user's face may include: adjusting the rendering of the virtual content item based on the change in the position of the head-mounted device relative to the reference point on the user's face.
[0012] In some aspects, determining the position of a head-mounted device relative to a reference point on a user's face can include performing time-of-flight measurements by a processor based on signals emitted by sensors on the head-mounted device. In some aspects, determining the position of a head-mounted device relative to a reference point on a user's face can include performing triangulation operations by a processor based on images captured by an imaging sensor on the head-mounted device.
[0013] Further aspects include a non-transitory processor-readable storage medium storing processor-executable instructions configured to cause a processor in a head-mounted device or an associated computing device to perform the operations of any of the methods outlined above. Further aspects include a head-mounted device or an associated computing device having functionality for performing any of the methods outlined above. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings incorporated herein and constituting a part of this specification illustrate example embodiments of various embodiments and, together with the general description given above and the detailed description given below, serve to explain the features of the claims.
[0015] Figure 1A is a diagram of a head-mounted device (e.g., augmented reality glasses) according to various embodiments that can be configured to perform vision-based registration operations that account for changes in distance / angle between cameras of the head-mounted device.
[0016] Figure 1B is a system block diagram showing a computer architecture and sensors that can be included in a head-mounted device according to various embodiments, the head-mounted device being configured to perform vision-based registration operations that account for changes in distance / angle between cameras of the head-mounted device and the user's eyes.
[0017] Figures 2A - 2E is a diagram of an imaging system adapted to display electronically generated images or virtual content items on a head-up display system.
[0018] Figures 3 - 5 is a processor flow diagram showing additional methods of performing vision-based registration operations according to various embodiments that account for changes in distance / angle between cameras of the head-mounted device and the user's eyes.
[0019] Figure 6 is a component block diagram of a mobile device suitable for implementing some embodiments.
[0020] Figure 7 is a component diagram of an example computing device suitable for use with various embodiments. DETAILED DESCRIPTION
[0021] Each embodiment will be described in detail with reference to the accompanying drawings. Wherever possible, the same reference numerals will be used throughout the drawings to refer to the same or like parts. References to specific examples and implementations are for illustrative purposes only and are not intended to limit the scope of the claims.
[0022] An augmented reality system works by displaying elements of virtual content such that they appear on, near, or associated with volumetric objects in the real world. The term "augmented reality system" refers to any system that renders virtual content items within a scene that includes real-world objects, including systems that render virtual content items such that they appear to be floating within the real world, mixed reality systems, and video see-through systems, which render images of real-world objects (e.g., obtained via an outward-facing camera) combined with virtual content items. In a common form of augmented reality system, a user wears a head-mounted device that includes: an outward-facing camera that captures images of the real world; a processor (which may be a separate computing device or a processor within the head-mounted device) that generates virtual content items (e.g., images, text, icons, etc.) and uses the images from the outward-facing camera to determine how or where the virtual content items should be rendered to appear on, near, or associated with a selected real-world object; and an image rendering device (e.g., a display or a projector) that renders the images such that the virtual content items appear to the user to be at the determined location relative to the real-world object.
[0023] For ease of describing the various embodiments, the positioning of virtual content items relative to a selected real-world object is referred to herein as "registration", and an item is "registered" with a real-world object when the virtual content item appears to the user to be on, near, or associated with the selected real-world object. As used herein, an item is "associated" with a real-world object when the augmented reality system attempts to render the virtual content item such that the item appears to be registered with the real-world object (i.e., appears to be at a fixed relative position, near, or remaining at a fixed relative position).
[0024] When registering a virtual content item associated with a real-world object in the distance, the head-mounted device renders the virtual content such that it appears to the user to be at the same distance as the real-world object, even though the virtual content is being rendered by an image rendering device (e.g., a projector or a display) that is only a few centimeters from the user's eyes. This can be achieved by using lenses in the image rendering device that refocus the light projected or displayed from the virtual content such that when the user is viewing a real-world object in the distance, the light is focused on the retina by the lens of the user's eye. Thus, even though the image of the virtual content is generated within a few millimeters of the user's eyes, the virtual content appears to be in focus as if it were at the same distance from the user as the real-world object associated with it through the augmented reality system.
[0025] Conventional head-mounted devices for augmented reality applications typically include thick or bulky nose bridges and frames, or are designed to be secured to the user's head via a headband. As augmented reality software applications continue to grow in popularity and use, there is expected to be an increased consumer demand for new types of head-mounted devices that have thinner or lighter nose bridges and frames and can be worn without a headband, similar to reading glasses or goggles. Due to these and other new features, it is more likely that the nose bridge will slip down from the user's nose, the frame will move or shift on the user's face, and the user will frequently adjust the location, position, and orientation of the head-mounted device on the user's nose and face (similar to how people currently adjust their reading glasses or goggles). Such movement and adjustment of the device can change the position, orientation, distance, and angle between the camera of the head-mounted device, the user's eyes, and the electronic display of the head-mounted device.
[0026] Projecting or lensing a virtual content item from an image rendering device (e.g., a projector or a display) near the user's eyes such that the virtual content appears registered and in focus with a real-world object when the user views the real-world object in the distance in a head-mounted device involves the use of waveguides, laser projection, lenses, or projectors. Such rendering techniques make the apparent position of the rendered content sensitive to changes in the distance and angle between the user's eyes and the image rendering device (e.g., a projector or a display). If such distance and angle are kept fixed, then the virtual content item can remain in focus and appear to remain registered with the real-world object at a fixed relative position determined by the augmented reality system. However, if the distance and angle between the user's eyes and the image rendering device (e.g., a projector or a display) change (e.g., if the head-mounted device slides down from the user's nose or the user repositions the head-mounted device onto the user's face), this will change the apparent depth and / or location of the virtual content item, while the distance to the real-world object and the location of the real-world object do not appear to change (i.e., the virtual content item appears to move relative to the real-world object). Due to the short distance from the image rendering device (e.g., a projector or a display) to the user's eyes compared to the distance to the real-world object, even small changes in the distance and angle of the head-mounted device will appear to cause the virtual content item to move through a large angle and distance compared to the distant object.
[0027] Even when not directly interacting with the virtual object, the user may make subtle movements (such as head, neck, or facial movements) and sudden movements (e.g., running, jumping, bending, etc.), which may affect the apparent position and orientation of the virtual content item relative to the real-world object. Such user movements may also cause the head-mounted device to move or shift on the user's nose or face, which changes the distance and angle between the virtual object image rendering device (e.g., a projector or a display) and the user's eyes. This may result in significant changes in the position, apparent distance, and / or orientation of the virtual content item relative to the real-world object. Similarly, when the user manually adjusts the position of the head-mounted device on the user's face, any movement changes the orientation, distance, and / or position of the display optics relative to the user's eyes, depending on the amount and direction / rotation of the movement. These movements of the head-mounted device relative to the user's eyes may cause the virtual object to appear at a different distance from the real-world object (e.g., out of focus when the user views the real-world object) and at a different angular location compared to the real-world object. This sensitivity of the apparent distance and angular position of the virtual object to the movement of the head-mounted device on the user's face may affect the fidelity of the augmented scene and degrade the user experience.
[0028] Some traditional solutions attempt to improve the accuracy of registration by collecting information from external sensing devices (such as magnetic or ultrasonic sensors communicatively coupled to a head-mounted device) to determine the position and orientation relative to the user's eyes, and using that information to adjust the location where a virtual object will be rendered during the positioning phase of vision-based registration. For example, traditional head-mounted devices may include gyroscopes and accelerometers, which can sense the rotation of the device through three rotational axes and the movement through three dimensions (i.e., 6 degrees of freedom). While these traditional solutions (especially in combination with the images provided by an outward-facing camera) provide information to an augmented reality system to achieve the realignment of virtual reality items with associated real-world objects (e.g., update the registration of virtual reality items), such sensors do not account for the movement of the head-mounted device relative to the user's eyes. Instead, most traditional vision-based registration techniques / technologies assume fixed positions and orientations of the outward-facing camera and the inward-facing image rendering device (e.g., waveguide, projector, or display) relative to the user's eyes. As a result, traditional augmented reality head-mounted devices may exhibit frequent changes in the apparent distance and angular position of virtual content items relative to distant objects (due to the movement of the device on the user's face).
[0029] Generally speaking, each embodiment includes a head-mounted device equipped with an outward-facing world view image sensor / camera and an inward-facing gaze view sensor / camera. The inward-facing gaze view sensor / camera can be configured to determine or measure changes in the distance and angle (referred to herein as "position" as defined below) between an image rendering device (e.g., a projector or a display) and the user's eyes. In a typical head-mounted device, the image rendering device (e.g., a projector or a display) will be at a fixed distance from the outward-facing camera (or more precisely, the image plane of the outward-facing camera). Thus, while an augmented reality system determines the appropriate rendering of virtual content items to appear associated with one or more real-world objects (i.e., to appear on or near one or more real-world objects), for this purpose, the process assumes a fixed distance and angular relationship between the outward-facing camera image plane and the image rendering device (e.g., a projector or a display). To correct for changes in the position of the image rendering device relative to the user's eyes due to movement of the head-mounted device on the user's face, a processor within the head-mounted device or communicating with the head-mounted device (e.g., via a wireless or wired link) can be configured to use distance and angle measurements from sensors configured to determine changes in the distance and angle to the user's face or eyes, and adjust the rendering of virtual content items (e.g., augmented imagery) to account for changes in the distance / angle between the head-mounted device and the user's eyes. Such adjustment can serve to stabilize the virtual content item relative to the apparent location relative to the real-world object, such that when the head-mounted device is displaced on the user's head, the virtual content remains at the same apparent location relative to the observed real world as the location determined by the augmented reality system.
[0030] In some embodiments, the head-mounted device can be configured to determine the distance and angle (or changes in distance and / or angle) between a registration point on the head-mounted device (e.g., a distance / angle sensor) and a registration point on the user's face (e.g., the user's eyes) with respect to six axes or degrees of freedom (i.e., the X, Y, Z, roll, pitch, and yaw axes and dimensions). For ease of reference, the terms "position" and "change in position" are used herein as a general reference to the distance and angular orientation between the head-mounted device and the user's eyes, and are intended to include any dimensional or angular measurement with respect to the six axes or degrees of freedom. For example, the head-mounted device moving down from the user's nose will result in changes in distance along the X and Z axes (e.g.) and rotation about the pitch axis, all combinations of which can be referred to herein as a change in the position of the head-mounted device relative to the user's eyes.
[0031] Assuming that the head-mounted device is rigid, there will be constant distance and angular relationships between an outward-facing image sensor, a projector or display that renders an image of a virtual content item, and an inward-facing sensor configured to determine the position of the head-mounted device relative to a reference point on the user's face. For example, the sensor can measure the distance (or change in distance) and angle (or change in angle) to a point on the user's face to determine the position (or change in position) along six axes or degrees of freedom. Additionally, the measurement of the position of the sensor relative to a registration point on the user's face can be related to both the outward-facing image sensor and the projector or display by a fixed geometric transformation. Thus, the inward-facing distance and angle measurement sensors can be placed anywhere on the head-mounted device, and the processor can use the distance and angle measurements to determine the position of the head-mounted device relative to the user's eyes and adjust the rendering of the virtual content item such that it appears to the user to remain registered with the associated real-world object.
[0032] In various embodiments, various types of sensors can be used to determine the relative position or change in position of the head-mounted device relative to the user's eyes. In an example embodiment, the sensor can be an inward-facing infrared sensor that generates small flashes of infrared light, detects the reflection of the small flashes of infrared light from the user's eyes, and determines the distance and angle (or change in distance and / or angle) between the outward-facing image sensor and the user's eyes by performing a time-of-flight measurement on the detected reflection. In another example, a single visible light camera can be configured to determine the change in distance and / or angle between the outward-facing image sensor and the user's eyes based on the change in position of features observed between two or more images. In another example, two spaced-apart imaging sensors (i.e., a binocular image sensor or a stereo camera) can be used to determine the distance by image processing performed by a processor to determine the angle to a common reference point on the user's face (e.g., the pupil of one eye) in each sensor and use triangulation to calculate the distance. In another example embodiment, the sensor can be one or more capacitive contact sensing circuits that can be embedded inside the head-mounted device to contact the user's face (e.g., the bridge of the nose, the eyebrow region, or the temple region) and are configured to output capacitance data that a processor can analyze to determine whether the device has moved or shifted on the user's face. In another example embodiment, the sensor can be an ultrasonic transducer that generates ultrasonic pulses, detects the echo of the ultrasonic pulses from the user's face, and determines the distance (or change in distance) between the outward-facing image sensor and the user's face by performing a time-of-flight measurement on the detected echo. The distance or change in distance can be determined based on the time between the generation of the IR flash or ultrasound and the speed of light or sound. The sensor can be configured to determine the angle to the user's eyes and thus also measure the change in the angle orientation of the head-mounted device (and thus the outward-facing camera) and the user's eyes. The processor can then use such measurements to determine the position (or change in position) of the head-mounted device relative to the user's eyes and determine the adjustments to be made to the rendering of the virtual content item such that the item appears to remain registered with the associated real-world object.
[0033] The head-mounted device processor can render an image of the adjusted or updated virtual content at the updated display location (i.e., distance and angle) to generate an enhanced scene such that the virtual content item remains registered with the real-world object as determined by the augmented reality system.
[0034] The term "mobile device" is used herein to refer to any one or all of the following: cellular phones, smart phones, Internet of Things (IoT) devices, personal or mobile multimedia players, laptop computers, tablet computers, ultrabooks, palmtop computers, wireless email receivers, cellular phones enabled with multimedia Internet, wireless game controllers, head-mounted devices, and similar electronic devices that include a programmable processor, memory, and circuitry for sending and / or receiving wireless communication signals to / from a wireless communication network. While the various embodiments are particularly useful in mobile devices such as smart phones and tablet devices, these embodiments are generally useful in any electronic device that includes communication circuitry for accessing a cellular or wireless communication network.
[0035] The phrase "head-mounted device" and the acronym (HMD) are used herein to refer to any electronic display system that presents a combination of computer-generated imagery and real-world images from a user's physical environment (i.e., the images the user would see without the use of glasses) and / or enables the user to view the generated images in the context of a real-world scene. Non-limiting examples of head-mounted devices include helmets, glasses, virtual reality glasses, augmented reality glasses, electronic goggles, and other similar technologies / devices or may be included in helmets, glasses, virtual reality glasses, augmented reality glasses, electronic goggles, and other similar technologies / devices. As described herein, a head-mounted device may include a processor, memory, a display, one or more cameras (e.g., world view camera, gaze view camera, etc.), one or more six-degree-of-freedom triangulation scanners, and a wireless interface for connecting to the Internet, a network, or another computing device. In some embodiments, the head-mounted device processor may be configured to execute or perform an augmented reality software application.
[0036] In some embodiments, the head-mounted device may be an accessory for a mobile device (e.g., a desktop computer, laptop computer, smart phone, tablet computer, etc.) and / or receive information from the mobile device, where all or part of the processing occurs in the mobile device (e.g., in Figure 6 and Figure 7It is executed on a processor of a computing device (such as the one shown in ) etc. Thus, in various embodiments, the head-mounted device can be configured to perform all processing locally on a processor in the head-mounted device, offload all major processing to a processor in another computing device (such as a laptop computer etc. located in the same room as the head-mounted device), or split major processing operations between a processor in the head-mounted device and a processor in the other computing device. In some embodiments, the processor in the other computing device can be a server in the "cloud", and the processor in the head-mounted device or in an associated mobile device communicates with the server via a network connection (such as a cellular network connection to the Internet).
[0037] The phrase "six degrees of freedom (6-DOF)" is used herein to refer to the degrees of freedom of movement of a head-mounted device or its components (relative to the user's eyes / head, computer-generated images or virtual objects, real-world objects, etc.) in three-dimensional space or relative to three perpendicular axes with respect to the user's face. The position of the head-mounted device on the user's head can change, such as moving in the forward / backward direction or along the X-axis (surge), in the left / right direction or along the Y-axis (sway), and in the up / down direction or along the Z-axis (heave). The orientation of the head-mounted device on the user's head can change, such as rotating around three perpendicular axes. The term "roll" can refer to rotation along the longitudinal axis or tilting from side to side on the X-axis. The term "pitch" can refer to rotation along the transverse axis or tilting forward and backward on the Y-axis. The term "yaw" can refer to rotation along the normal axis or turning left and right on the Z-axis.
[0038] A variety of different methods, technologies, solutions, and / or techniques (collectively referred to herein as "solutions") can be used to determine the location, position, or orientation of points on a user's face (e.g., points on the facial structure around the user's eyes, eyes, eye sockets, eye corners, corneas, pupils, etc.), any or all of which can be implemented by, included in, and / or used by various embodiments. As described above, various types of sensors (including IR, image sensors, binocular image sensors, capacitive contact sensing circuits, and ultrasonic sensors) can be used to measure the distance and angle from the sensors on the head-mounted device to points on the user's face. The processor can apply trilateration or multilateration to the measurements made by the sensors as well as accelerometer and gyroscope sensor data to determine the six-degree-of-freedom (DOF) position (i.e., distance and angular orientation) of the head-mounted device relative to the user's face. For example, the head-mounted device can be configured to send a sound (e.g., ultrasound), light, or radio signal to a target point, measure how long it takes for the reflection of the sound, light, or radio signal to be detected by the sensors on the head-mounted device, and use any or all of the above techniques (e.g., time of arrival, angle of arrival, etc.) to estimate the distance and angle between the lens or camera of the head-mounted device and the target point. In some embodiments, a processor (such as the processor of the head-mounted device) can use a three-dimensional (3D) model (e.g., 3D reconstruction) of the user's face when processing an image captured by an inward-facing image sensor (e.g., a digital camera) to determine the position of the head-mounted device relative to the user's eyes.
[0039] As discussed above, an augmented reality system can use vision-based registration techniques / technologies to align the rendered image of a virtual content item (e.g., computer-generated imagery, etc.) such that the item appears to be registered with a real-world object to the user. Thus, based on the position and orientation of the head-mounted device on the user's face or head, the device projects or otherwise renders the virtual content item such that it appears to the user as desired relative to the associated real-world object. Examples of vision-based registration techniques include optical-based registration, video-based registration, registration with artificial markers, registration with natural markers, multi-camera model-based registration, hybrid registration, and registration by blur estimation.
[0040] For example, an augmented reality system may perform operations including the following: a head-mounted device captures an image of a real-world scene (i.e., the physical environment around the user); processes the captured image to identify four known points from the image of the real-world scene; sets four separate tracking windows around these points; determines a camera calibration matrix (M) based on these points; determines the "world", "camera", and "screen" coordinates of the four known points based on the camera calibration matrix (M); determines a projection matrix based on these coordinates; and uses the projection matrix to sequentially superimpose each of the generated virtual objects onto the image of the real-world scene.
[0041] Each of the vision-based registration techniques mentioned above may include a localization phase, a rendering phase, and a merging phase. During the localization phase, the head-mounted device may capture an image of the real-world scene (i.e., the physical environment around the user) and determine where the virtual objects are to be displayed in the real-world scene. During the rendering phase, the head-mounted device may generate two-dimensional images of the virtual content items. During the merging phase, the head-mounted device may render the virtual content items via an image rendering device (e.g., a projector or a display) so that these items appear to be superimposed on or superimposed with the real-world scene for the user. Specifically, the head-mounted device may present the superimposed image to the user such that the user can view and / or interact with the virtual content items in a manner that appears natural to the user, regardless of the position and orientation of the head-mounted device. In a mixed reality device employing video see-through technology, the merging phase may involve rendering both the real-world image and the virtual content items together in the display.
[0042] After or as part of the merging phase, as the position and / or orientation of the head-mounted device on the user's head changes (such as when the user manually repositions the head-mounted device on the user's face), the processor may obtain distance and / or angle measurements to the registration points on the user's face from the inward-facing distance and angle sensors and adjust the rendering of the virtual content items so that the virtual content items appear to remain in the same relative position (i.e., distance and angle) to the real-world scene.
[0043] Figure 1A FIG. shows a head-mounted device 100 that may be configured according to various embodiments. In Figure 1AIn the example shown, the head-mounted device 100 includes a frame 102, two optical lenses 104, and a processor 106. The processor 106 is communicatively coupled to an outward-facing world view image sensor / camera 108, an inward-facing gaze view sensor / camera 110, a sensor array 112, a memory 114, and a communication circuit 116. In some embodiments, the head-mounted device 100 may include capacitive contact sensing circuitry along the arms 120 of the frame or in the bridge 122 of the head-mounted device 100. In some embodiments, the head-mounted device 100 may also include sensors for monitoring physical conditions (e.g., location, movement, acceleration, orientation, altitude, etc.). The sensors may include any one or all of the following: gyroscopes, accelerometers, magnetometers, magnetic compasses, altimeters, odometers, and pressure sensors. The sensors may also include various biosensors for collecting information related to the environment and / or user conditions (e.g., heart rate monitors, body temperature sensors, carbon sensors, oxygen sensors, etc.). The sensors may also be located external to the head-mounted device 100 and paired or grouped with the head-mounted device 100 via a wired or wireless connection (e.g., etc.).
[0044] In some embodiments, the processor 106 may also be communicatively coupled to an image rendering device 118 (e.g., an image projector). The image rendering device 118 may be embedded in the arm portion 120 of the frame 102 and configured to project an image onto the optical lens 104. In some embodiments, the image rendering device 118 may include a light-emitting diode (LED) module, a light tunnel, a homogenizing lens, an optical display, a folding mirror, or other components (known projectors or head-mounted displays). In some embodiments (e.g., embodiments where the image rendering device 118 is not included or not used), the optical lens 104 may be, or may include, a transparent or partially transparent electronic display. In some embodiments, the optical lens 104 includes an image generating element, such as a transparent organic light-emitting diode (OLED) display element or a liquid crystal on silicon (LCOS) display element. In some embodiments, the optical lens 104 may include separate left and right eye display elements. In some embodiments, the optical lens 104 may include a light guide for transmitting light from the display element to the wearer's eyes or operate as a light guide.
[0045] The outward-facing or world view image sensor / camera 108 may be configured to capture real-world images from the user's physical environment and send the corresponding image data to the processor 106. The processor 106 may combine the real-world images with computer-generated imagery or virtual objects (VOs) to generate an enhanced scene and render the enhanced scene on the electronic display or the optical lens 104 of the head-mounted device 100.
[0046] The inward-facing or gaze view sensor / camera 110 can be configured to obtain image data from the user's eyes or facial structures around the user's eyes. For example, the gaze view sensor / camera 110 can be configured to produce a small flash (such as infrared light), capture their reflections from the user's eyes (such as eye sockets, eye corners, corneas, pupils, etc.), and send the corresponding image data to the processor 106. The processor 106 can use the image data received from the gaze view sensor / camera 110 to determine the optical axis of each eye in the user's eyes, the gaze direction of each eye, the user's head orientation, various eye gaze speed or acceleration values, the angular change of the eye gaze direction, or other similar gaze-related information. In addition, the processor 106 can use the image data received from the gaze view sensor / camera 110 and / or other sensor components in the head-mounted device 100 (such as capacitive contact sensing circuits) to determine the distance and angle between each eye in the user's eyes and the world view image sensor / camera 108 or the optical lens 104.
[0047] In some embodiments, the head-mounted device 100 (or the inward-facing or gaze view sensor / camera 110) can include a scanner and / or a tracker, which are configured to determine the distance and angle between each eye in the user's eyes and the world view image sensor / camera 108 or the optical lens 104. The scanner / tracker can be configured to measure the distance and angle to registration points on the user's face (such as points on the user's eyes or facial structures around the user's eyes). Various types of scanners or trackers can be used. In one example embodiment, the scanner / tracker can be an inward-facing binocular image sensor, which images the user's face or eyes and determines the distance and angle to the registration points by triangulation. In another example embodiment, the scanner / tracker can be an IR sensor, which is configured to project a laser beam towards the registration points and determine the coordinates of the registration points by measuring the distance between the scanner / tracker and the points (such as by measuring the time of flight of the reflected IR flash) and two angles (such as via an absolute rangefinder, interferometer, angle encoder, etc.). In some embodiments, the scanner / tracker can include ultrasonic transducers (such as piezoelectric transducers), which are configured to perform time-of-flight measurements by emitting ultrasound towards the surface of the user's eyes or facial features around the user's eyes, measuring the time between the sound pulse and the detected echo, and determining the distance based on the measured time of flight.
[0048] In some embodiments, the scanner / tracker can be or include a triangulation scanner configured to determine the coordinates of an object (e.g., a user's eye or facial structure around the user's eye) based on any one of various image triangulation techniques known in the art. Generally, the triangulation scanner projects light (e.g., light from a laser line probe) or a two-dimensional (2D) light pattern (e.g., structured light) over an area onto the object surface. The triangulation scanner can include a camera coupled to a light source (e.g., an IR light-emitting diode laser) that has a fixed mechanical relationship with the outward-facing camera on the head-mounted device. The projected light ray or light pattern emitted from the light source can be reflected from the surface on the user's face (e.g., one or both of the user's eyes) and imaged by the camera. Since the camera and the light source are arranged in a fixed relationship with respect to each other, the distance and angle to the object surface can be determined based on trigonometric principles from the projected light ray or pattern, the captured camera image, and the baseline distance separating the light source and the camera.
[0049] In some embodiments, the scanner / tracker can be configured to acquire a series of images and register these images relative to each other such that the position and orientation of each image relative to the other images are known, use features (e.g., fiducial points) located in the images to match overlapping regions of adjacent image frames, and determine the distance and angle between each of the user's eyes and the world view image sensor / camera 108 or the optical lens 104 based on the overlapping regions.
[0050] In some embodiments, the processor 106 of the head-mounted device 100 may be configured to use localization and mapping techniques (such as Simultaneous Localization and Mapping (SLAM), Visual Simultaneous Localization and Mapping (VSLAM), and / or other techniques known in the art) to construct and update a map of the visual environment and / or determine the distances and angles between each of the user's eyes, each of the optical lenses 104 in the optical lenses 104, and / or each of the world view image sensors / cameras 108 in the world view image sensors / cameras 108. For example, the world view image sensor / camera 108 may include a monocular image sensor that captures images or frames from the environment. The head-mounted device 100 may identify prominent objects or features within the captured images, estimate the size and scale of the features in the images, compare the identified features with each other and / or with features of known size and scale in a test image, and identify correspondences based on the comparison. Each correspondence may be a set of values or an information structure that identifies a feature (or feature point) in one image as having a high probability of being the same feature in another image (e.g., a subsequently captured image). In other words, a correspondence may be a set of image points in correspondence (e.g., a first point in a first image and a second point in a second image, etc.). The head-mounted device 100 may generate a homography matrix information structure based on the identified correspondences and use the homography matrix to determine its pose (e.g., position, orientation, etc.) within the environment.
[0051] Figure 1B FIG. shows a computer architecture and various sensors that may be included in the head-mounted device 100, which is configured to account for changes in the distance / angle between the camera of the head-mounted device and the user's eyes. In Figure 1B the example shown, the head-mounted device 100 includes a main board assembly 152, an image sensor 154, a microcontroller unit (MCU) 156, an infrared (IR) sensor 158, an inertial measurement unit (IMU) 160, a laser distance and angle sensor (LDS) 162, and an optical flow sensor 164.
[0052] Sensors 154, 158 - 164 can collect information useful for implementing SLAM technology in the head - mounted device 100. For example, the optical flow sensor 164 can measure optical flow or visual motion and output a measurement based on the optical flow / visual motion. Optical flow can identify or define the pattern of apparent motion of objects, surfaces, and edges in a visual scene caused by the relative motion between an observer (e.g., the head - mounted device 100, the user, etc.) and the scene. The MCU 156 can use the optical flow information to determine the visual or relative motion between the head - mounted device 100 and real - world objects near the head - mounted device 100. Based on the visual or relative motion of the real - world objects, the processor can use SLAM technology to determine the distance and angle to the real - world objects. By determining the distance and angle to the real - world objects, the augmented reality system can determine the virtual distance for rendering virtual content items such that they appear to be at the same distance (i.e., in focus) as the real - world objects with which the virtual content is to be registered. In some embodiments, the optical flow sensor 164 can be an image sensor coupled to the MCU 156 (or the processor 106 shown in Figure 1A ), and the MCU 156 is programmed to run an optical flow algorithm. In some embodiments, the optical flow sensor 164 can be a vision chip that includes an image sensor and a processor on the same chip or die.
[0053] The head - mounted device 100 can be equipped with various additional sensors, including gyroscopes, accelerometers, magnetometers, magnetic compasses, altimeters, cameras, optical readers, orientation sensors, monocular image sensors, and / or similar sensors for monitoring physical conditions (e.g., location, motion, acceleration, orientation, altitude, etc.) or collecting information useful for implementing SLAM technology.
[0054] Figure 2A An imaging system, shown as part of an augmented reality system, is adapted to render electronically - generated images or virtual content items on a head - mounted device. Figure 2AThe example illustrations in [the figure] include the user's retina 252, the real-world object retinal image 254, the eye focus 256, the virtual content item (virtual ball) retinal image 258, the virtual content item focus 260, the user's eye 262 (e.g., the eye lens, pupil, cornea, etc.), the user's iris 264, the light projection 266 for the virtual content item, the surface or facial features 268 around the user's eye, the sensor 272 for measuring the distance (d) 270 and angle between the sensor 272 and the facial features 268, the image sensor lens 274, the image rendering device 276, the light 278 from the real-world object, the field of view (FOV) cone 280 of the image sensor lens 274, and the real-world object 282. In some embodiments, the image rendering device 276 can be a projector or a display (e.g., a waveguide, a laser projector, etc.) configured to render an image of the virtual content item to appear at a distance from and near the real-world object.
[0055] Figure 2B It shows that when the user adjusts the position of the head-mounted device on the user's face or when the head-mounted device is displaced on the user's nose, the distance (d’) 270 between the sensor 272 and the facial features 268 changes. Although the movement of the head-mounted device on the user's face changes the position (e.g., distance and / or angle) of the image rendering device 276 relative to the user's eye 262 and face 268 (which changes the perceived location of the virtual content item 258), the distance D 284 to the real-world object 282 does not change. Thus, the perceived location of the real-world object 282 (which is related to the location of the object image 254 on the retina 252) remains unchanged, but as the virtual content retinal image 258 changes relative to the real-world object retinal image 254, the change in the position of the image rendering device 276 relative to the user's eye 262 and face 268 changes the location where the user perceives the virtual content item (virtual ball).
[0056] Figure 2C It shows that a system configured to perform vision-based adjustments according to various embodiments can determine the position (or change in position) of the image sensor 274 relative to the eye 262 and adjust the rendering of the virtual content item 258 based on the position or change in position such that the virtual content appears stable relative to the real-world object and natural to the user, regardless of the position of the head-mounted device on the user's head.
[0057] In some embodiments, the system may be configured to determine a new distance (d') 270 (or a change in distance) between the sensor 272 and the user's eye 262 (or another reference point on the user's face) through various techniques such as measuring the time-of-flight of an IR flash or an ultrasonic pulse, triangulating an image obtained from a binocular image sensor. The system may adjust the rendering of the virtual content item 258 based on the determined distance or change in distance between the image sensor and the user's eye such that the virtual content item (the ball in the figure) appears again (i.e., the retinal image) to be registered with the real-world object (the racket in the figure) because the retinal image 258 of the ball is at the same focal length as and next to the retinal image 254 of the real-world object.
[0058] Figure 2D and Figure 2E illustrates that in some embodiments, determining the position of the image sensor 274 relative to the user's eye 262 may include determining a change in the angular orientation of the image rendering device 276 due to movement of the head-mounted device on the user's face. For example, as Figure 2D shown, as the head-mounted device slides down from the user's nose, the separation distance d' will change, and the vertical orientation of the head-mounted device will drop a height h relative to the centerline 290 of the user's eye 262, resulting in a change about the pitch axis. As a result of the change in the vertical distance h, the angle of the projection 277 from the image rendering device 276 to the iris 264 will change by an amount a1, which will further change the apparent location of the virtual content item unless the rendering of the virtual content is adjusted. As a further example, as Figure 2E shown, if the head-mounted device rotates on or around the user's face, the angle between the projections 277 from the image rendering device 276 to the iris 264 will change by an amount a2, which will also change the apparent location of the virtual content item unless the rendering of the virtual content is adjusted.
[0059] In some embodiments, such as embodiments where the inward-facing sensor 272 is a binocular image sensor, the sensor may also be configured to measure vertical translation and rotational changes based on the angular change in the position of the user's eye 262 (or another registration point associated with the eye on the user's face) relative to the sensor 272. For example, the measurement of vertical translation and rotational changes may be determined through image processing that identifies features or fiducials (e.g., iris, eye corner, mole, etc.) discerned on the user's face in multiple images, and tracks the change in the apparent position of the feature across multiple images and from two image sensors, and uses trigonometry to determine the distance from the sensor to the eye and the rotational movement of the head-mounted device relative to the user's eye (i.e., changes in the projection angle and distance). In some embodiments, the rotational change measurement may be obtained through different sensors on the head-mounted device. By determining such changes in the position and orientation of the image rendering device 276 relative to the user's eye, the processor may apply one or more trigonometric transforms to appropriately adjust the rendering of the virtual content item.
[0060] Figure 3 Method 300 for adjusting the rendering of a virtual content item in a head-mounted device to account for changes in the position (e.g., distance / angle) of the head-mounted device relative to the user's eye is shown in accordance with some embodiments. All or part of method 300 may be performed by components or a processor (e.g., processor 106) in the head-mounted device or by a processor in another computing device (e.g., the device shown in Figure 6 and Figure 7 .
[0061] Prior to initiating method 300, an augmented reality system implemented in a head-mounted device, for example, may capture an image of the real-world scene using an outward-facing image sensor from the physical environment surrounding the head-mounted device. The captured image may include surfaces, people, and objects physically present in the user's field of view. The augmented reality system may process the captured image to identify the location of real-world objects or surfaces in the image of the captured real-world scene. For example, the augmented reality system may perform image processing operations to identify prominent objects or features within the captured image, estimate the size and scale of the features in the image, compare the identified features with each other and / or with features of known size and scale in a test image, and identify correspondences based on the comparison. As part of such operations, the augmented reality system may generate a homography matrix information structure based on the identified correspondences and use the homography matrix to identify the location of real-world objects or surfaces in the captured image. The augmented reality system may generate an image of a virtual object and determine the display location of the virtual object based on the location of the real-world object or surface.
[0062] In block 302, the head-mounted device processor can process the information received from the inward-facing sensors to determine the position (or change in position) of the head-mounted device relative to a reference point on the face of the user of the head-mounted device. In some embodiments, the reference point on the user's face can be a facial feature having a fixed location relative to the user's eyes. In some embodiments, the reference point can be one or both of the user's eyes. As described, if the head-mounted device is rigid, the position of the sensors relative to the virtual content image rendering device and the image plane of the outward-facing imaging sensors will be fixed, and thus, changes in the distances and angles measured by the sensors at any location on the head-mounted device will translate into equivalent changes in the distances and angles between the virtual content image rendering device and the image plane of the outward-facing imaging sensors. Similarly, because the features on the user's face are typically in fixed positions relative to the user's eyes, the sensors can provide the processor with information indicating the position of the head-mounted device relative to any of a variety of registration points on the user's face, such as one or more of the following: the user's eyes (or portions thereof), cheeks, eyebrows, nose, etc., because any changes will translate into equivalent changes in the distances and angles to the user's eyes. For example, the information received by the processor from the sensors can relate to the position of the head-mounted device relative to one or both of the user's eyes.
[0063] As described herein, any one of a variety of sensors can be used to make the measurements used by the processor in block 302. For example, in some embodiments, a head-mounted device processor can use image processing (such as tracking the location or relative movement of features on a user's face from one image to the next) such that an inward-facing image sensor provides information regarding the relative position (e.g., in terms of distance and angle) of the head-mounted device or image rendering device relative to the user's eyes or surrounding facial structure. In some embodiments, the inward-facing sensor can be combined or coupled to an IR emitter (e.g., an IR LED laser), the IR emitter can be configured to emit a flash of IR light, and the sensor can be configured to detect the scattering of the IR light and provide information regarding the relative position (e.g., in terms of distance and angle) of the head-mounted device or image rendering device relative to the user's face or eyes based on a time-of-flight measurement between the time the IR flash is emitted and the time the scattered light is detected. The IR imaging sensor can also provide information regarding the relative position (e.g., angle) of the head-mounted device or image rendering device relative to the user's eyes based on the image location of the reflected (as opposed to scattered) light from the eyes. In some embodiments, the processor can cause a piezoelectric transducer to emit an acoustic pulse (such as an ultrasonic pulse) and record the time of the echo from the user's eyes or surrounding facial structure to determine the time of flight, from which the relative distance between the sensor and the user's face can be calculated. In block 302, the head-mounted device processor can determine the position of the outward-facing image sensor relative to the user's eyes based on the time-of-flight measurement information provided by the sensor. In some embodiments, the sensor can be a binocular image sensor, and the processor can use triangulation operations and / or 6-DOF calculations to determine the position of the head-mounted device relative to a reference point (e.g., one or both of the user's eyes) on the user's face based on the information obtained by the inward-facing sensor. In some embodiments, the sensor can use a 3D rendering of the user's face when processing the images obtained by the image sensor to determine the position of the head-mounted device relative to the user's eyes. In some embodiments, the sensor can be one or more capacitive contact sensing circuits that can be embedded inside the head-mounted device to contact the user's face (e.g., the bridge of the nose, the eyebrow area, or the temple area) and are configured to output capacitance data that the processor can analyze to determine whether the device has moved or shifted on the user's face.
[0064] In block 304, the head-mounted device processor may determine and apply an adjustment to the rendering of the display position of a virtual content item or virtual object by an image rendering device (e.g., in terms of focus and / or location) based on the position (or change in position) of the head-mounted device relative to a reference point on the user's face (e.g., the user's eyes) determined in block 302. In some embodiments, the processor may be configured to determine and apply an adjustment to the rendering of the display location of a virtual content item or virtual object by an image rendering device (e.g., in terms of focus and / or location) based on the position (or change in position) of the image rendering device relative to one or both of the user's eyes. The processor may adjust the virtual content rendering to account for changes in position, orientation, distance, and / or angle between the outward-facing camera of the head-mounted device, the user's eyes, and the image rendering device (e.g., projector or display) of the head-mounted device. For example, as the bridge of the head-mounted device slides down the user's nose, the adjustment to the virtual content rendering may be adjusted for changes in the distance and height (relative to the eye centerline) of the image rendering device. As another example, the adjustment to the virtual content rendering may be adjusted for rotational changes in the projection angle due to movement of the head-mounted device with a thin or lightweight frame on the user's face.
[0065] The processor may repeat the operations in blocks 302 and 304 continuously, periodically, or sporadically triggered by an event to frequently adjust the rendering of the virtual content item so that the image appears to remain stable relative to the real world despite movement of the head-mounted device on the user.
[0066] Figure 4 Method 400 for performing a vision-based registration operation that accounts for changes in the position (i.e., distance / angle) of a head-mounted device relative to a user's eyes is shown. All or part of method 400 may be performed by components or a processor (e.g., processor 106) in the head-mounted device or a processor in another computing device (e.g., the device shown in Figure 6 and Figure 7 ).
[0067] In block 402, an outward-facing image sensor of the head-mounted device may capture an image of the real-world scene from the physical environment surrounding the head-mounted device. The captured image may include surfaces, people, and objects physically present in the field of view of the user / wearer. As part of the operation in block 402, the captured image may be provided to the head-mounted device processor.
[0068] In block 404, the head-mounted device processor can process the captured image to identify the location of real-world objects or surfaces in the image of the captured real-world scene (to which one or more virtual content items can be related in an augmented reality scene). For example, the head-mounted device 100 can perform image processing operations to identify prominent objects or features within the captured image, estimate the size and scale of features in the image, compare the identified features with each other and / or with features of known size and scale in a test image, and identify correspondences based on the comparison. As part of identifying the location of real-world objects or surfaces in block 404, the processor can use known methods (such as VSLAM or binocular triangulation if the sensors are binocular) to determine the distance to such objects or surfaces. Also as part of the operations in block 404, the head-mounted device 100 can generate a homography matrix information structure based on the identified correspondences and use the homography matrix to identify the location of real-world objects or surfaces in the captured image.
[0069] In block 406, as part of generating an augmented reality scene, the head-mounted device processor can generate an image of a virtual object and determine the display location of the virtual object based on the location of the real-world object or surface. The head-mounted device processor can perform the operations in block 406 as part of the localization phase of vision-based registration in the augmented reality process.
[0070] In block 302, the head-mounted device processor can process information received from the inward-facing sensors to determine the position (or change in position) of the head-mounted device relative to a reference point on the user's face (such as one or both eyes of the user of the head-mounted device), as described in the similarly numbered blocks for method 300.
[0071] In block 410, the head-mounted device processor can adjust the image of the virtual content item based on the determined position of the outward-facing image sensor or image rendering device on the head-mounted device relative to one or both eyes of the user. The processor can adjust the image (e.g., scale down, scale up, rotate, translate, etc.) to account for minor movements and adjustments that change the relative position and orientation (e.g., in terms of distance and angle) between the camera of the head-mounted device, the user's eyes, and the electronic display of the head-mounted device. For example, this adjustment can compensate when the bridge of the head-mounted device slides down the user's nose, the thin or lightweight frame of the head-mounted device moves or shifts on the user's face, or the user manually adjusts the location, position, and / or orientation of the head-mounted device on the user's nose or face.
[0072] In block 412, the head-mounted device processor can adjust the rendering of the adjusted image of the virtual content item such that the content appears registered with the real-world scene (e.g., adjacent to or superimposed on the real-world scene) at the same depth of focus as the real-world object. The head-mounted device processor can perform the operations in block 412 as part of a merging phase of a vision-based registration process.
[0073] In block 414, the head-mounted device processor can use the image rendering device of the head-mounted device to render the augmented scene. For example, the image rendering device can be a projector (e.g., a laser projector or a waveguide) that projects an image of the virtual content item onto the lens of the head-mounted device or into the user's eye such that the virtual content is in focus with and appears near the associated distant real-world object. The head-mounted device processor can perform the operations in block 414 as part of a rendering phase of vision-based registration for an augmented reality scene. In block 414, the head-mounted device can present an overlay image to the user such that the user can view and / or interact with the virtual object.
[0074] The processor can repeat the operations in blocks 302 and 410 - 414 continuously, periodically, or sporadically triggered by an event to frequently adjust the rendering of the virtual content item such that the image appears stable relative to the real world despite movement of the head-mounted device on the user. Additionally, the processor can repeat method 400 periodically or sporadically triggered by an event to update the augmented reality scene, such as adding, removing, or changing virtual content items, and adjusting the apparent position of the virtual content item relative to real-world objects as the user moves through the environment.
[0075] Figure 5 Method 500 for performing vision-based registration operations that account for changes in distance / angle between an outward-facing camera of a head-mounted device and the user / wearer's eyes is shown. All or part of method 500 can be performed by components or a processor (e.g., processor 106) in the head-mounted device or by a processor in another computing device (e.g., the device shown in Figure 6 and Figure 7 ). In blocks 402 - 406, the head-mounted device processor can perform operations that are the same as or similar to the operations in similarly numbered blocks of method 400 described in reference Figure 4
[0076] In block 508, the head-mounted device processor may communicate with an inward-facing image sensor of the head-mounted device to cause the inward-facing image sensor to generate a brief flash of infrared light toward the user's eyes. The inward-facing gaze view sensor / camera may be configured to capture an image of the user's eyes and perform trigonometric calculations to determine the distance to each eye (or cornea, pupil, etc.) and the angle between each eye (or cornea, pupil, etc.) and an outward-facing camera of the head-mounted device (e.g., a world view camera, a gaze view camera) and / or an image rendering device (e.g., a projector or a display).
[0077] In block 510, the head-mounted device processor may receive information received from the inward-facing image sensor when the inward-facing image sensor captures the reflection or scattering of the brief flash of infrared light from the user's eyes or surrounding facial features. The reflection may be captured by the inward-facing gaze view sensor / camera or another sensor in the head-mounted device.
[0078] In block 512, the head-mounted device processor may use the information received from the sensor to determine the position of the outward-facing camera relative to one or both eyes of the user by performing time-of-flight measurements and / or trigonometric calculations based on the captured reflected light and / or scattered light. For example, the head-mounted device processor may process the information received from the sensor related to the position of the head-mounted device relative to the user's eyes to determine the translation, tilt, roll, pitch, and yaw (or their corresponding triangles) movement of the head-mounted device or the image sensor relative to the user's eyes.
[0079] In block 514, the head-mounted device processor may adjust the rendering of the virtual object based on the determined position of the outward-facing image sensor relative to one or both eyes of the user. The head-mounted device processor may perform the operation in block 514 as part of a rendering phase of vision-based registration for an augmented reality scene. In block 514, the head-mounted device may present an overlay image to the user such that the user can view and / or interact with the virtual object.
[0080] The processor may repeat the operations in blocks 508-514 continuously, periodically, or sporadically triggered by an event to frequently adjust the rendering of the virtual content item so that the image appears to remain stable relative to the real world despite the movement of the head-mounted device on the user. Additionally, the processor may periodically or sporadically triggered by an event repeat method 500 to update the augmented reality scene, such as adding, removing, or changing virtual content items, and adjusting the apparent position of the virtual content item relative to real-world objects as the user moves through the environment.
[0081] Various embodiments may be implemented on various mobile devices, inFigure 6 An example of a mobile device in the form of a smart phone is shown. For example, an image captured by the imaging sensor of the head-mounted device 100 can be wirelessly transmitted (e.g., via a Bluetooth or WiFi wireless communication link 610) to the smart phone 600, where the processor 601 can perform a part of the processing in any of the methods 200, 300, 400, and 500, and then send the result back to the head-mounted device 100. The smart phone 600 may include a processor 601 coupled to an internal memory 602, a display 603, and a speaker 604. Additionally, the smart phone 600 may include an antenna 605 for transmitting and receiving wireless signals 610, and the wireless signals 610 may be connected to a wireless data link and / or a cellular phone transceiver 606 coupled to the processor 601. The smart phone 600 generally also includes menu selection buttons or rocker switches 607 for receiving user input.
[0082] A typical smart phone 600 also includes a sound codec (CODEC) circuit 608 that digitizes the sound received from the microphone into data packets suitable for wireless transmission, and decodes the received sound data packets to generate an analog signal provided to the speaker to generate sound. In addition, one or more of the processor 601, the wireless transceiver 606, and the CODEC 608 may include digital signal processor (DSP) circuitry (not shown separately).
[0083] The methods of various embodiments can be implemented in various personal computing devices (such as Figure 7 the laptop computer 700 shown). For example, an image captured by the imaging sensor of the head-mounted device 100 can be wirelessly transmitted (e.g., via a Bluetooth or WiFi wireless communication link 708) to the laptop computer 700, where the processor 701 can perform a part of the processing in any of the methods 200, 300, 400, and 500, and then send the result back to the head-mounted device 100. The laptop computer 700 will generally include a processor 701 coupled to a volatile memory 702 and a large-capacity non-volatile memory (such as a disk drive 704 of flash memory). The laptop computer 700 may also include a floppy disk drive 705 coupled to the processor 706. The computer receiver device 700 may also include several connector ports or other network interfaces coupled to the processor 701 for establishing a data connection or receiving an external memory receiver device, such as a universal serial bus (USB) or A connector socket, or other network connection circuitry for coupling a processor 701 to a network (e.g., a communication network). In a notebook configuration, the computer housing includes a touchpad 710, a keyboard 712, and a display 714 all coupled to the processor 701. Other configurations of the computing device may include a computer mouse or trackball known to be coupled to a processor (e.g., via a USB input), which may also be used in conjunction with the various embodiments.
[0084] The processor can be any programmable microprocessor, microcomputer, or one or more multiprocessor chips that can be configured by software instructions (applications) to perform various functions, including the functions of the various embodiments described in this application. In some mobile devices, multiple processors may be provided, such as one processor dedicated to wireless communication functions and one processor dedicated to running other applications. Generally, software applications may be stored in the internal memory 606 before they are accessed and loaded into the processor. The processor may include internal memory sufficient to store the software application instructions.
[0085] The various embodiments shown and described are provided only as examples to illustrate the various features of the claims. However, the features shown and described with respect to any given embodiment are not necessarily limited to the associated embodiment and may be used or combined with the other embodiments shown and described. Additionally, the claims are not intended to be limited by any one exemplary embodiment. For example, one or more of the operations in these methods may be replaced or combined with one or more of the operations of these methods.
[0086] The foregoing method descriptions and process flow diagrams are provided only as illustrative examples and are not intended to require or imply that the operations of the various embodiments be performed in the order given. As will be understood by those skilled in the art, the operations in the foregoing embodiments may be performed in any order. Words such as "thereafter," "subsequently," "then," etc. are not intended to limit the order of the operations; these words are used to guide the reader through the description of the method. Additionally, any reference to a claim element in the singular form (e.g., using the articles "a," "an," or "the") should not be construed as limiting the element to the singular.
[0087] The various illustrative logical blocks, functional components, functional units, circuits, and algorithmic operations described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, functional units, circuits, and operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the claims.
[0088] Hardware for implementing or performing the various illustrative logics, logic blocks, functional components, and circuits described in connection with the embodiments disclosed herein can be implemented using a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.
[0089] In one or more embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a non-transitory computer-readable medium or a non-transitory processor-readable medium. Operations of a method or algorithm disclosed herein may be embodied in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium that can be accessed by a computer or a processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, operations of a method or algorithm may reside as one or any combination of code and / or instructions, or a set of code and / or instructions, on a non-transitory processor-readable medium and / or a computer-readable medium, which may be incorporated into a computer program product.
[0090] The foregoing description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the claims. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and implementations without departing from the scope of the claims. Thus, the present disclosure is not intended to be limited to the embodiments and implementations described herein, but is to be accorded the widest scope consistent with the claims and the principles and novel features disclosed herein.
Claims
1. A head-mounted device for use in an augmented reality system, comprising: A memory; A contact sensor; And A processor coupled to the memory and the contact sensor, wherein the processor is configured to: Receive information from the contact sensor, wherein the information indicates a position change of the head-mounted device relative to a reference point on the face of a user wearing the head-mounted device; and Adjust the rendering of a virtual content item based on the position change.
2. The head-mounted device according to claim 1, wherein The information received from the contact sensor is related to a position change of the head-mounted device relative to the eyes of the user, and wherein the processor is configured to adjust the rendering of the virtual content item based on the position change of the head-mounted device relative to the eyes of the user.
3. The head-mounted device according to claim 2, further comprising: An infrared (IR) sensor and an IR light source configured to emit IR light towards the face of the user.
4. The head-mounted device according to claim 2, further comprising: An ultrasonic sensor configured to emit ultrasonic pulses towards the face of the user.
5. The head-mounted device according to claim 2, further comprising a first camera configured to image a first eye of the user.
6. The head-mounted device according to claim 5, further comprising a second camera configured to image a second eye of the user.
7. The head-mounted device according to claim 1, further comprising: An image rendering device coupled to the processor and configured to render the virtual content item.
8. The head-mounted device according to claim 1, wherein, The processor is further configured to: Determine an angle from the image rendering device on the head-mounted device to the eyes of the user; And Adjust the rendering of the virtual content item based on the determined angle to the eyes of the user and the determined distance between the head-mounted device and the reference point on the face of the user.
9. The head-mounted device according to claim 1, wherein, The contact sensor includes a capacitive contact sensing circuit.
10. The head-mounted device according to claim 1, wherein, The contact sensor is embedded on a surface of the head-mounted device for contacting the face of the user.
11. The head-mounted device according to claim 10, wherein, The contact sensor is configured to contact the nose of the user.
12. The head-mounted device according to claim 1, further comprising a biosensor.
13. The head-mounted device according to claim 1, further comprising one or more of a gyroscope sensor and an accelerometer.
14. A method for adjusting the rendering of a virtual content item in an augmented reality system to compensate for movement of a head-mounted device on a user, comprising: Receiving information from a contact sensor, wherein the information indicates a position change of the head-mounted device relative to a reference point on the face of a user wearing the head-mounted device; and Adjusting the rendering of the virtual content item based on the position change of the head-mounted device relative to the reference point on the face of the user.
15. The method according to claim 14, wherein The reference point on the face of the user includes the eyes of the user.
16. The method according to claim 14, further comprising: Determining the position of the head-mounted device relative to the reference point on the face of the user based on information received from an infrared (IR) sensor and an IR light source configured to emit IR light towards the face of the user.
17. The method according to claim 14, further comprising: Determining a position of the head-mounted device relative to the reference point on the user's face based on information received from an ultrasonic sensor configured to emit ultrasonic pulses towards the user's face.
18. The method according to claim 14, further comprising: Determining a position of the head-mounted device relative to the reference point on the user's face using time-of-flight measurements based on signals emitted by a sensor on the head-mounted device.
19. The method according to claim 14, further comprising: Determining a position of the head-mounted device relative to the reference point on the user's face based on one or more images captured by an imaging sensor on the head-mounted device.
20. The method according to claim 14, wherein, The contact sensor includes a capacitive contact sensing circuit.
21. The method according to claim 14, wherein The contact sensor is embedded in a surface of the head-mounted device for engaging the user's face.
22. The method according to claim 21, wherein, The contact sensor is configured to engage the user's nose.
23. A non-transitory processor-readable medium having processor-executable instructions stored thereon, the processor-executable instructions being configured to cause a processor of a head-mounted device to perform operations including: Receiving information from a contact sensor indicating a change in position of the head-mounted device relative to a reference point on the face of a user wearing the head-mounted device; and Adjusting rendering of a virtual content item based on the change in position of the head-mounted device relative to the reference point on the user's face.
24. The non-volatile processor-readable medium according to claim 23, wherein, The stored processor-executable instructions are configured to cause the processor of the head-mounted device to perform operations further including: Receiving information related to a position of the head-mounted device relative to the user's eyes, and wherein the stored processor-executable instructions are configured to cause the processor of the head-mounted device to perform operations such that adjusting the rendering of the virtual content item includes adjusting the rendering of the virtual content item based on the position of the head-mounted device relative to the user's eyes.
25. The non-volatile processor-readable medium according to claim 23, wherein, The stored processor-executable instructions are configured to cause the processor of the head-mounted device to perform operations further including: Determining a position of the head-mounted device relative to the reference point on the user's face based on information received from an infrared (IR) sensor and an IR light source configured to emit IR light towards the user's face.
26. The non-volatile processor-readable medium according to claim 23, wherein, The stored processor-executable instructions are configured to cause the processor of the head-mounted device to perform operations further including: Determining a position of the head-mounted device relative to the reference point on the user's face based on information received from an ultrasonic sensor configured to emit ultrasonic pulses towards the user's face.
27. The non-volatile processor-readable medium according to claim 23, wherein, The stored processor-executable instructions are configured to cause the processor of the head-mounted device to perform operations further including: Determining a position of the head-mounted device relative to the reference point on the user's face using time-of-flight measurements based on signals emitted by a sensor on the head-mounted device.
28. The non-volatile processor-readable medium according to claim 23, wherein, The stored processor-executable instructions are configured to cause the processor of the head-mounted device to perform operations further including: Determine the position of the head-mounted device relative to the reference point on the face of the user based on one or more images captured by an imaging sensor on the head-mounted device.
29. A head-mounted device, comprising: A unit configured to receive, from a contact sensor, information indicating a change in the position of the head-mounted device relative to a reference point on the face of a user wearing the head-mounted device; And A unit configured to adjust the rendering of a virtual content item based on the change in the position of the head-mounted device relative to the reference point on the face of the user.
30. The head-mounted device according to claim 29, wherein, The reference point on the face of the user includes the eyes of the user.
31. The head-mounted device according to claim 29, further comprising: A unit configured to determine the position of the head-mounted device relative to the reference point on the face of the user based on an infrared (IR) sensor and an IR light source configured to emit IR light towards the face of the user.
32. The head-mounted device according to claim 29, further comprising: A unit configured to determine the position of the head-mounted device relative to the reference point on the face of the user based on ultrasonic pulses emitted towards the face of the user.
33. The head-mounted device according to claim 29, further comprising: A unit configured to determine the position of the head-mounted device relative to the reference point on the face of the user using time-of-flight measurements based on signals emitted by a sensor on the head-mounted device.
34. The head-mounted device according to claim 29, further comprising: A unit configured to determine the position of the head-mounted device relative to the reference point on the face of the user based on one or more images captured by an imaging sensor on the head-mounted device.
Citation Information
Patent Citations
Opacity filter for see-through head mounted display
CN102540463A
Periocular test for mixed reality calibration
US20180096503A1
Head-mountable display device and method
US20190377191A1