Techniques for camera focusing in mixed reality environments with hand gesture interaction

The HMD device's adjustable-focus PV camera reduces autofocus hunting by using a depth sensor to track hand movements, enhancing image quality and power efficiency during mixed reality interactions.

JP7764535B2Active Publication Date: 2025-11-05MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024081899
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-31
Filing Date
2024-05-20
Publication Date
2025-11-05
Estimated Expiration
2040-04-24

AI Technical Summary

Technical Problem

Mixed reality head-mounted display (HMD) devices experience autofocus hunting due to user hand movements, degrading image quality and consuming excessive power.

Method used

An adjustable-focus PV camera in the HMD device uses a depth sensor to track hand location and motion, controlling autofocus based on hand characteristics to reduce hunting.

Benefits of technology

Reduces autofocus hunting, improving image quality and power efficiency by intelligently managing autofocus operations during hand interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764535000001
    Figure 0007764535000001
  • Figure 0007764535000002
    Figure 0007764535000002
  • Figure 0007764535000003
    Figure 0007764535000003
Patent Text Reader

Abstract

To reduce auto-focus hunting occurrences during camera operations.SOLUTION: An adjustable-focus PV camera in a mixed-reality head-mounted display (HMD) device operates with an auto-focus subsystem that is configured to be triggered based on location and motion of a user's hands. The HMD device is equipped with a depth sensor that is configured to capture depth data from the surrounding physical environment to detect and track the user's hand location, movements, and gestures in three-dimensions. The hand tracking data from the depth sensor may be assessed to determine whether hand characteristics are detected, its size, motion, speed, etc. within a particular region of interest (ROI) in the field of view of the PV camera. The auto-focus subsystem uses the assessed hand characteristics as an input to control auto-focus of the PV camera to reduce auto-focus hunting occurrences.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background Mixed reality head-mounted display (HMD) devices can utilize photo and video (PV) cameras that capture still and / or video images of the surrounding physical environment to facilitate various user experiences, including recording and sharing of the mixed reality experience. The PV camera can include autofocus, autoexposure, and autobalance capabilities. In some scenarios, hand movements of an HMD device user can result in the autofocus subsystem hunting while attempting to resolve a sharp image of the physical environment. For example, hand movements of the user as they interact with a hologram rendered by the HMD device can result in the camera refocusing each time the hand is detected in the scene by the camera. This autofocus hunting phenomenon can reduce the quality of the user experience for both the local HMD device user and remote users who may be viewing the mixed reality user experience captured by the local HMD device. Summary of the Invention

[0002] overview

[0002] An adjustable-focus PV camera in a mixed reality head-mounted display (HMD) device operates with an autofocus subsystem configured to be triggered based on the location and motion of a user's hand to reduce the occurrence of autofocus hunting during PV camera operation. The HMD device includes a depth sensor configured to capture depth data from the surrounding physical environment to detect and track the location, movement, and gestures of a user's hand in three dimensions. Hand tracking data from the depth sensor can be evaluated to determine hand characteristics—such as which hand or which part of the user's hand is detected, its size, motion, and speed—within a specific region of interest (ROI) within the PV camera's field of view (FOV). The autofocus subsystem uses the evaluated hand characteristics as input to control the PV camera's autofocus to reduce the occurrence of autofocus hunting. For example, if hand tracking indicates that a user is utilizing hand motion while interacting with a hologram, the autofocus subsystem can inhibit autofocus triggering to reduce the hunting phenomenon.

[0003]

[0003] Reducing autofocus hunting can be beneficial because it can be an undesirable noise for HMD device users (who may detect frequent PV camera lens motion) and can also result in degradation of the quality of images and videos captured by the PV camera. Reducing autofocus hunting can also improve the operation of the HMD device by reducing the power consumed by the autofocus motor or other mechanisms.

[0004] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Moreover, the claimed subject matter is not limited to implementations that solve any or all of the disadvantages noted in any part of this disclosure. It should be understood that the above-described subject matter may be implemented as a computer-controlled device, a computer process, a computing system, or an article of manufacture, such as one or more computer-readable storage media. These and various other features will become apparent from a review of the following Detailed Description and associated drawings. [Brief explanation of the drawings]

[0005] DESCRIPTION OF THE DRAWINGS [Figure 1]

[0005] An exemplary mixed reality environment is shown in which holograms are rendered on a see-through mixed reality display system of a head-mounted display (HMD) device while a user views the surrounding physical environment. [Figure 2]

[0006] 1 illustrates an exemplary environment in which a local HMD device, a remote HMD device, and a remote service may communicate over a network. [Figure 3]

[0007] 1 illustrates an exemplary architecture of an HMD device. [Figure 4]

[0008] 1 illustrates a local user in a physical environment interacting with an exemplary virtual object. [Figure 5]

[0008] Figure 1 illustrates a local user in a physical environment interacting with an exemplary virtual object. [Figure 6]

[0009] 1 illustrates an exemplary FOV of a local HMD device from the perspective of a local user, including a field of view of a physical environment on which virtual objects are rendered using a mixed reality display system. [Figure 7]

[0010] 1 illustrates an exemplary configuration in which content is shared from a local HMD device user to a remote user. [Figure 8]

[0011] 1 shows a remote user operating a remote tablet computer that displays a composite image including real-world elements and virtual objects transmitted from a local user's HMD device. [Figure 9]

[0012] 1 illustrates exemplary hand motions and gestures from the perspective of a local user in the FOV of a local HMD device. [Figure 10]

[0012] Figure 3 shows exemplary hand motions and gestures from the perspective of a local user in the FOV of a local HMD device. [Figure 11]

[0012] Figure 3 shows exemplary hand motions and gestures from the perspective of a local user in the FOV of a local HMD device. [Figure 12]

[0013] 1 illustrates an exemplary spherical coordinate system representing the horizontal FOV. [Figure 13] 1 shows an exemplary spherical coordinate system representing the vertical FOV. [Figure 14]

[0014] 1 illustrates an exemplary region of interest (ROI) within the FOV of an HMD device from the perspective of a local user using a spherical coordinate system. [Figure 15]

[0015] FIG. 10 illustrates various data provided as inputs into the autofocus subsystem of a local HMD device for illustrative purposes. [Figure 16]

[0016] 1 illustrates a classification of exemplary items in a physical environment that may be detected by a depth sensor of a local HMD device. [Figure 17]

[0017] 10 illustrates an exemplary process performed by an autofocus subsystem of a local HMD device when processing frames of content. [Figure 18]

[0018] 1 illustrates an exemplary classification of characteristics used by the autofocus subsystem in determining whether to trigger or inhibit autofocus. [Figure 19]

[0019] 1 is a flowchart of an exemplary method performed by an HMD device or other suitable electronic device utilizing an autofocus subsystem. [Figure 20] 1 is a flowchart of an exemplary method performed by an HMD device or other suitable electronic device utilizing an autofocus subsystem. [Figure 21] 1 is a flowchart of an exemplary method performed by an HMD device or other suitable electronic device utilizing an autofocus subsystem. [Figure 22]

[0020] FIG. 1 is a schematic block diagram of an exemplary remote service or computer system that may be used in part to implement the present techniques for setting camera focus in a mixed reality environment with hand gesture interaction. [Figure 23]

[0021] FIG. 1 is a block diagram of an example data center that may be used at least in part to implement the present techniques for setting camera focus in a mixed reality environment with hand gesture interaction. [Figure 24]

[0022] FIG. 1 is a schematic block diagram of an example architecture of a computing device, such as a smartphone or tablet computer, that may be used to implement the present techniques for setting camera focus in a mixed reality environment with hand gesture interaction. [Figure 25]

[0023] FIG. 1 is a pictorial diagram of an illustrative example of a mixed reality HMD device. [Figure 26]

[0024] FIG. 1 is a block diagram of an illustrative example of a mixed reality HMD device. DETAILED DESCRIPTION OF THE INVENTION

[0006]

[0025] Like reference numerals refer to like elements throughout the drawings, and elements are not drawn to scale unless otherwise indicated.

[0007] Detailed Description

[0026] FIG. 1 illustrates an exemplary mixed reality environment 100 supported on an HMD device 110 that combines real-world elements and computer-generated virtual objects to enable a variety of user experiences. A user 105 can utilize the HMD device 110 to experience the mixed reality environment 100, which is virtually rendered on a see-through mixed reality display system and, in some implementations, may include audio and / or haptic / tactile sensations. In this particular, non-limiting example, an HMD device user physically walks through a real-world urban area, including city streets with various buildings, stores, etc. The field of view (FOV) from the user's perspective (represented by a dashed line in FIG. 1 ) of the real-world urban space provided by the HMD device changes as the user moves through the environment, and the device can render holographic virtual objects onto the real-world field of view. Here, the holograms include various virtual objects, including a tag 115 identifying a company, directions 120 to a point of interest in the environment, and a gift box 125. Virtual objects within the FOV coexist with real objects in a three-dimensional (3D) physical environment to create a mixed reality experience. Virtual objects can be positioned in relation to the real-world physical environment, such as a gift box on a sidewalk, or in relation to the user, such as the direction in which they move with the user.

[0008]

[0027] FIG. 2 illustrates an exemplary environment in which local and remote HMD devices can communicate with each other and with a remote service 215 over a network 220. The network may be comprised of various network-connected devices to enable communication between computing devices and may include any one or more of a local area network, a wide area network, the Internet, the World Wide Web, etc. In some embodiments, an ad-hoc (e.g., peer-to-peer) network between devices can be created using, for example, Wi-Fi, Bluetooth, or near-field communication (NFC), as representatively shown by dashed arrow 225. A local user 105 can operate a local HMD device 110, which can communicate with remote HMD devices 210 operated by individual remote users 205. The HMD device can perform various tasks like a typical computer (e.g., a personal computer, a smartphone, a tablet computer, etc.), and can also perform additional tasks based on the configuration of the HMD device. Tasks may include sending email or other messages, searching the web, sending photos or videos, interacting with holograms, and using a camera to send a live stream of the surrounding physical environment, among other tasks.

[0009]

[0028] HMD devices 110 and 210 can communicate with remote computing devices and services, such as remote service 215. The remote service may be, for example, a cloud computing platform set up in a data center that can enable the HMD devices to leverage various solutions provided by the remote service, such as artificial intelligence (AI) processing, data storage, data analysis, etc. While FIG. 2 also shows an HMD device and a server, the HMD device may also communicate with other types of computing devices, such as smartphones, tablet computers, laptops, personal computers, and the like (not shown). For example, a user experience implemented on a local HMD device can be shared with a remote user, as described below. Images and video of a mixed reality scene that a local user observes on their HMD device, along with sound and other experience elements, can be received and rendered on a laptop computer at a remote location.

[0010]

[0029] FIG. 3 illustrates an exemplary system architecture of an HMD device, such as local HMD device 110. While various components are depicted in FIG. 3, the listed components are not exhaustive, and other components not shown that facilitate the functionality of the HMD device are possible, such as a global positioning system (GPS), other input / output devices (keyboard and mouse), etc. The HMD device may have one or more processors 305, such as a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence (AI) processing unit. The HMD device may have memory 310 that may store data and instructions executable by processor 305. The memory may include short-term memory devices, such as random access memory (RAM), and may further include long-term memory devices, such as flash storage and solid-state drives (SSDs).

[0011]

[0030] The HMD device 110 may include an I / O (input / output) system 370 comprised of various components to allow a user to interact with the HMD device. Exemplary, but non-exhaustive, components include a speaker 380, a gesture subsystem 385, and a microphone 390. As representatively shown by arrow 382, ​​the gesture subsystem may interoperate with a depth sensor 320 that may collect depth data about the user's hands and thereby enable the HMD device to perform hand tracking.

[0012]

[0031] The depth sensor 320 can communicate collected data about the hand to a gesture subsystem 385 that handles actions associated with the user's hand movements and gestures. The user can have interactions with holograms on the HMD device's display, such as moving holograms, selecting holograms, shrinking or enlarging holograms (e.g., with a pinch motion), among other interactions. Exemplary holograms that the user can control include buttons, menus, images, results from a web-based search, and character figures, among other holograms.

[0013]

[0032] The see-through mixed reality display system 350 may include a microdisplay or imager 355 and a mixed reality display 365, such as a waveguide-based display that uses a surface relief grating to render virtual objects on the HMD device 110. The processor 305 (e.g., an image processor) may be operatively connected to the imager 355 to provide image data, such as video data, so that images may be displayed using the light engine and the waveguide display 365. In some implementations, the mixed reality display may be configured as a near-eye display that includes an exit pupil expander (EPE) (not shown).

[0014]

[0033] The HMD device 110 may include many types of sensors 315 to provide a user with an integrated and immersive experience within a mixed reality environment. A depth sensor 320 and a photo / video (PV) camera 325 are exemplary sensors shown, but other sensors not shown, such as infrared sensors, pressure sensors, and motion sensors, are possible. The depth sensor can operate using various types of depth sensing techniques, such as structured light, passive stereo, active stereo, time-of-flight, pulsed time-of-flight, phased time-of-flight, or light detection and ranging (LIDAR). Typically, depth sensors operate with an IR (infrared) light source, although some sensors may operate with an RGB (red, green, blue) light source. Generally, a depth sensor detects the distance to a target and uses a point cloud representation to construct an image representing the external surface properties of the target or physical environment. The points or structures of the point cloud data can be stored in memory locally, at a remote service, or a combination thereof.

[0015]

[0034] The PV camera 325 can be configured with an adjustable focus to capture images, record video of the physical environment surrounding the user, or transmit content from the HMD device 110 to a remote computing device, such as the remote HMD device 210 or another computing device (e.g., a tablet or personal computer). The PV camera can be implemented as an RGB camera for capturing a scene within the three-dimensional (3D) physical space in which the HMD device operates.

[0016]

[0035] The camera subsystem 330 associated with the HMD device may be used at least in part for the PV camera and may include an auto-exposure subsystem 335, an auto-balance subsystem 340, and an auto-focus subsystem 345. The auto-exposure subsystem may perform automatic adjustment of image brightness according to the amount of light reaching the camera sensor. The auto-balance subsystem may automatically correct color differences based on the lighting so that white is displayed properly. The auto-focus subsystem may focus the PV camera lens to ensure that captured and rendered images are sharp, which is typically implemented by mechanical movement of the lens relative to the image sensor.

[0017]

[0036] The composite generator 395 generates composite content that combines the physical world scene captured by the PV camera 325 with images of virtual objects generated by the HMD device. The composite content can be recorded or transmitted to a remote computing device, such as an HMD device, a personal computer, a laptop, a tablet, a smartphone, or the like. In a typical implementation, the images are non-holographic 2D representations of the virtual objects. However, in alternative implementations, data can be transmitted from a local HMD device to a remote HMD device to enable remote rendering of holographic content.

[0018]

[0037] The communications module 375 can be used to send and receive information to external devices, such as the remote HMD device 210, the remote service 215, or other computing devices. The communications module can include, for example, a network interface controller (NIC) for wireless communication with a router or similar network-connected device, or a radio supporting one or more of Wi-Fi, Bluetooth™, or near field communication (NFC) transmissions.

[0019]

[0038] 4 and 5 illustrate an exemplary physical environment in which a user 105 interacts with holographic virtual objects observable by the user through a see-through mixed reality display system on an HMD device (note that the holographic virtual objects in this illustrative example are only visible through the HMD device and cannot be projected into free space to allow viewing by the naked eye, for example). In FIG. 4, the virtual objects include a vertically oriented panel 405 and a cylindrical object 410. In FIG. 5, the virtual objects include a horizontally oriented virtual building model 505. The virtual objects are positioned at various locations relative to the 3D space of the physical environment, including a plant 415 and a photo 420. Although not labeled, floors, walls, or doors are also part of the real physical environment.

[0020]

[0039] FIG. 6 shows the field of view (FOV) 605 of an exemplary mixed reality scene viewed using a see-through display from the perspective of a user of an HMD device. The user 105 can observe holograms of portions of the physical world as well as virtual objects 405 and 410 generated by the local HMD device 110. The holograms can be positioned anywhere within the physical environment but are typically placed between 50 centimeters and 5 meters from the user, for example, to minimize user discomfort from mismatched adaptive eye vergence. The user typically interacts with the holograms using a mixture of up-down, left-right, and in-out hand motions, as shown in FIGS. 9-11 and described in the accompanying text. In some implementations, interaction can occur at some spatial distance away from the location of the rendered hologram. For example, a virtual button exposed on a virtual object can be pressed by the user by performing a tapping gesture at a predetermined distance from the object. The specific user-hologram interactions utilized in a given implementation may vary.

[0021]

[0040] FIG. 7 illustrates an exemplary environment in which a remote user 205 operates a remote tablet device 705 that renders content 710 from a local user's HMD device 110. In this example, the rendering includes composite content including a scene of the local user's physical environment captured by a PV camera on the local HMD device and a 2D non-holographic rendering of virtual objects 405 and 410. As shown in FIG. 8, the composite rendering 805 is substantially similar to what the local user sees through a see-through mixed reality display on the local HMD device. Thus, the remote user can observe the local user's hands interacting with virtual object 405 as well as portions of the surrounding physical environment, including plants, walls, photos, and doors. The received content at the remote device 705 can include still images and / or video streamed in real time or can include recorded content. In some implementations, the received content can include data enabling the remote rendering of 3D holographic content.

[0022]

[0041] 9-11 illustrate exemplary hand motions and gestures that may be performed by the local user 105 while operating the local HMD device 110. FIG. 9 illustrates the user's vertical (e.g., upward and downward) hand movements as the user manipulates the virtual object 405, as representatively indicated by reference numeral 905. FIG. 10 illustrates the user's horizontal (e.g., left to right) hand movements to manipulate the virtual object 405, as representatively indicated by reference numeral 1005. FIG. 11 illustrates the user's in / out movement within the mixed reality space, for example, by performing a "bloom" gesture, as representatively indicated by reference numeral 1105. Other directional movements not illustrated in FIGS. 9-11 are also possible while the user is operating the local HMD device, such as circular movements, regular movements, and various hand gestures involving manipulating the user's fingers.

[0023]

[0042] 12 and 13 show exemplary spherical coordinate systems representing horizontal and vertical fields of view (FOV). In a typical implementation, the spherical coordinate system may use the radial distance from the user to a point in 3D space, the azimuth angle from the user to the point in 3D space, and the polar angle (or altitude) between the user and the point in 3D space to coordinate a point in the physical environment. FIG. 12 shows a user's overhead view depicting the horizontal FOV in relation to various sensors, displays, and components within the local HMD device 110. The horizontal FOV has an axis extending parallel to the ground, with its origin located at the HMD distance, for example, between the user's eyes. Different components may be associated with different angles of the horizontal FOV α. h , which are typically relatively narrow in relation to the user's human binocular FOV. FIG. 13 shows the vertical FOV α associated with the various sensors, displays, and components within the local HMD device. v 1 shows a side view of a user depicting a vertical FOV axis extending perpendicular to the ground and having its origin at the HMD device. The vertical FOV may also vary with the angle of components within the HMD device.

[0024]

[0043] 14 illustrates an example HMD device FOV 605 in which an example region of interest (ROI) 1405 is shown. The ROI is a statically or dynamically defined region within the HMD device FOV 605 (FIG. 6) that may be utilized by the autofocus subsystem of the HMD device 110 in determining whether to focus on a hand movement or gesture. The ROI displayed in FIG. 14 is for illustrative purposes only; in a typical implementation, a user will not perceive the ROI while viewing content on a see-through mixed reality display on the local HMD device.

[0025]

[0044] The ROI 1405 can be implemented as a 3D region of space that can be expressed using spherical or Cartesian coordinates. By using a spherical coordinate system, the ROI can, in some implementations, be dynamic according to the measured distance from the user and the effect that distance has on the azimuth and polar angles. Typically, the ROI can be located in a central region of the display system FOV because that is likely where the user will gaze, but the ROI can be anywhere within the display system FOV, such as an off-center location. The ROI can have a static location, size, and shape relative to the FOV, or in some embodiments, can be dynamically positioned, sized, and shaped. Thus, the ROI can be any static or dynamic 2D shape or 3D volume, depending on the implementation. During interaction with a hologram of a virtual object, the hand of the user 105 in FIG. 14 can be positioned within the ROI defined by a set of spherical coordinates.

[0026]

[0045] 15 shows an example diagram of data provided into the autofocus subsystem 345 of the camera subsystem 330. The autofocus subsystem collectively utilizes this data to control autofocus operations in a manner that reduces the autofocus hunting phenomenon produced by hand movements while interacting with a rendered holographic virtual object.

[0027]

[0046] Data provided into the autofocus subsystem includes data representing the physical environment from the PV camera 325 and the depth sensor 320 or other forward-facing sensor 1525. Data from the forward-facing sensor may include depth data as captured by the depth sensor 320, although other sensors may also be used to capture the physical environment around the user. Thus, the term forward-facing sensor 1525 is used herein to reflect the use of one or more depth sensors, cameras, or other sensors to capture the physical environment, as well as the user's hand movements and gestures, as described in further detail below.

[0028]

[0047] FIG. 16 illustrates exemplary categories of items that may be picked up and collected by a forward-facing sensor 1525, such as depth data, from a physical environment, as representatively indicated by reference numeral 1605. Items that may be picked up by the forward-facing sensor 1525 may include, among other objects, a user's hand 1610, physical real-world objects (e.g., chairs, beds, couches, tables) 1615, people 1620, and structures (e.g., walls, floors) 1625. The forward-facing sensor may or may not be able to recognize objects from the collected data, but may perform spatial mapping of the environment based on the collected data. However, the HMD device may be configured to detect and recognize hands to enable gesture input and further affect the autofocus operation described herein. The captured data shown in FIG. 15 that is transmitted to the autofocus subsystem includes hand data associated with the user of the HMD device, as described in further detail below.

[0029]

[0048] 17 shows an example diagram of an autofocus subsystem 345 receiving recorded frames of content 1705 (e.g., for streaming content, recording video, or capturing images) and using captured hand data to autofocus on the frames of content. The autofocus subsystem can be configured to have one or more criteria that, if met or not met, determine whether the HMD device triggers or inhibits autofocus operation.

[0030]

[0049] The autofocus operation can include an autofocus subsystem that automatically focuses on content within a ROI of the display FOV (FIG. 14). Satisfying a criterion may indicate, for example, that a user is using their hand in a manner that allows the user to clearly see their hand, making their hand the user's focal point within the ROI. For example, if a user interacts with a hologram within the ROI, the autofocus subsystem may not want to focus on the user's hand because their hand is used to pass through to control the hologram, but the hologram is still the user's primary point of interest. In other embodiments, the user's hand may be transient in the ROI and thus not a focal point for focusing. Conversely, if a user uses their hand in a manner separate from the hologram, such as to generate a new hologram or open a menu, the autofocus subsystem may choose to focus on the user's hand. The set criterion provides assistance to the autofocus subsystem to intelligently focus or not focus on the user's hand, thereby reducing the hunting phenomenon and improving the quality of recorded content when livestreaming for remote users or playing back for local users. In essence, the implementation of the criterion helps determine whether the user's hands are a point of interest for the user within the FOV.

[0031]

[0050] In step 1710, the autofocus subsystem determines whether one or more hands are present within the ROI. The autofocus subsystem may obtain data about the hands from the depth sensor 320 or another forward-facing sensor 1525. The collected data may be adjusted to a corresponding location on the display FOV to estimate the location of the user's physical hands in relation to the ROI. This may be done frame-by-frame or by using groups of frames.

[0032]

[0051] In step 1715, if one or more of the user's hands are not present in the ROI, the autofocus subsystem continues to operate normally by autofocusing on the environment detected within the ROI. In step 1720, if one or more hands are detected within the ROI, the autofocus subsystem determines whether characteristics of the hands indicate that the user is interacting with a hologram or that the user's hands are not a point of interest. In step 1730, if the user's hands are determined to be a point of interest, the autofocus subsystem triggers autofocus operation of the camera on content within the ROI. In step 1725, if the user's hands are determined to be not a point of interest, the autofocus subsystem inhibits autofocus operation of the camera.

[0033]

[0052] FIG. 18 shows an exemplary classification of hand characteristics used by the autofocus subsystem to determine whether to trigger or disable autofocus operation, as representatively indicated by reference numeral 1805. The characteristics can include the pace of hand movement within and around the ROI 1810, which the autofocus subsystem uses to trigger or inhibit focus operation when captured hand data indicates that the hand pace meets or exceeds, or does not meet or exceed, a preset speed limit (e.g., in meters per second). Thus, for example, when hand data indicates that the hand pace meets or does not meet a preset speed limit, autofocus operation can be inhibited even when a hand is present within the ROI. This prevents the lens from experiencing hunting when a user sporadically moves their hand in front of a forward-facing sensor or when the hand is transient.

[0034]

[0053] Another characteristic that may affect whether an autofocus operation is triggered or disabled includes the duration 1815 that one or more hands are positioned within the ROI and statically positioned (e.g., within a particular region of the ROI). The autofocus operation may be disabled if one or more hands are not statically positioned for a duration that meets a preset threshold limit (e.g., 3 seconds). Conversely, the autofocus operation may be performed when one or more hands are statically positioned within the ROI or within a region of the ROI for a duration that meets a preset threshold time limit.

[0035]

[0054] The detected hand size 1820 can also be used to determine whether to trigger or disable an autofocus operation. Using the user's hand size as a criterion in the autofocus subsystem's decision can, for example, help prevent another user's hand from affecting the autofocus operation. The user's hand pose (e.g., whether the hand pose indicates device input) 1825 can be used by the autofocus subsystem to determine whether to trigger or suppress an autofocus operation. For example, certain poses can be extraneous or sporadic hand movements, while some hand poses can be used for input or the user can be recognized as pointing at something. A hand pose identified as being productive can be the reason the autofocus subsystem focuses on the user's hand.

[0036]

[0055] The direction of motion (e.g., front-to-back, left-to-right, inward-to-outward, diagonal, etc.) 1830 can be utilized to determine whether to trigger or inhibit an autofocus operation. Which of the user's hands (e.g., left or right) is detected within the forward-facing sensor FOV 1835 and which part of the hand is detected 1840 can also be used in determining whether to trigger or inhibit an autofocus operation. For example, one hand may have relatively greater determinative power regarding the user's point of interest within the ROI. For example, a user may typically use one hand to interact with holograms, and therefore, that hand is not necessarily the point of interest. In contrast, the opposite hand may be used to open menus, be used as a pointer within mixed reality space, or otherwise be a point of interest to the user. Other characteristics 1845 not shown can also be used as criteria in determining whether to trigger or inhibit an autofocus operation.

[0037]

[0056] 19-21 are flowcharts of example methods 1900, 2000, and 2100 that may be performed using the local HMD device 110 or other suitable computing device. Unless specifically stated, the methods or steps shown in the flowcharts and described in the accompanying text are not constrained to any particular order or sequence. In addition, methods or some of their steps may occur or be performed simultaneously, and depending on the requirements of such implementation, it is not necessary for all methods or steps to be performed in a given implementation, and some methods or steps may be optionally utilized.

[0038]

[0057] 19, in step 1905, the local HMD device enables autofocus operation of a camera configured to capture a scene in a local physical environment surrounding the local HMD device over a field of view (FOV). In step 1910, the local HMD device collects data regarding the user's hand using a set of sensors. In step 1915, the local HMD device inhibits autofocus operation of the camera based on the collected data if the user's hand does not meet one or more criteria of the autofocus subsystem.

[0039]

[0058] In step 2005 of Figure 20, the computing device captures data to track one or more hands of a user while using the computing device in a physical environment. In step 2010, the computing device selects a region of interest (ROI) within a field of view (FOV) of the computing device. In step 2015, the computing device determines from the captured hand tracking data whether a portion of the user's hand or hands is located within the ROI. In step 2020, the computing device triggers or disables autofocus operation of a camera disposed on the computing device that is configured to capture the scene.

[0040]

[0059] In step 2105 of FIG. 21 , a computing device renders at least one hologram on a display, the hologram including a virtual object positioned in a physical environment at a known location. In step 2110, the computing device captures location tracking data of one or more hands of a user in the physical environment by using a depth sensor. In step 2115, the computing device determines from the location tracking data whether one or more hands of the user are interacting with a hologram at the known location. In step 2120, the computing device triggers operation of an autofocus subsystem in response to determining that one or more hands of the user are not interacting with a hologram. In step 2125, the computing device inhibits operation of the autofocus subsystem in response to determining that one or more hands of the user are interacting with a hologram.

[0041]

[0060] 22 is a schematic block diagram of an exemplary computer system 2200, such as a PC (personal computer) or server, in which the present technique for setting camera focus in a mixed reality environment with hand gesture interaction may be implemented. For example, the HMD device 110 may be in communication with the computer system 2200. The computer system 2200 includes a processor 2205, a system memory 2211, and a system bus 2214 that couples various system components including the system memory 2211 to the processor 2205. The system bus 2214 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, or a local bus using any of a variety of bus architectures. The system memory 2211 includes a read-only memory (ROM) 2217 and a random access memory (RAM) 2221. A basic input / output system (BIOS) 2225, containing the basic routines that help transfer information between elements within the computer system 2200, such as during start-up, is stored in the ROM 2217. Computer system 2200 may further include an internally disposed hard disk drive 2228 that reads from and writes to a hard disk (not shown), a magnetic disk drive 2230 that reads from and writes to a removable magnetic disk 2233 (e.g., a floppy disk), and an optical disk drive 2238 that reads from or writes to a removable optical disk 2243, such as a CD (compact disc), DVD (digital versatile disc), or other optical media. The hard disk drive 2228, magnetic disk drive 2230, and optical disk drive 2238 are connected to system bus 2214 by a hard disk drive interface 2246, a magnetic disk drive interface 2249, and an optical drive interface 2252, respectively. The drives and their associated computer-readable storage media provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for computer system 2200.Although illustrative examples include a hard disk, a removable magnetic disk 2233, and a removable optical disk 2243, other types of computer-readable storage media that can store data that is accessible by a computer, such as a magnetic cassette, a flash memory card, a digital video disk, a data cartridge, a random access memory (RAM), a read-only memory (ROM), and the like, can also be used in some applications of the present techniques to set camera focus within a mixed reality environment with hand gesture interaction. Additionally, the term "computer-readable storage medium" as used herein includes one or more example media types (e.g., one or more magnetic disks, one or more CDs, etc.). For purposes of this specification and claims, the phrase "computer-readable storage medium" and variations thereof are intended to encompass non-transitory embodiments and do not include waves, signals, and / or other transitory and / or intangible communication media.

[0042]

[0061] A number of program modules, including an operating system 2255, one or more application programs 2257, other program modules 2260, and program data 2263, may be stored on the hard disk, magnetic disk 2233, optical disk 2243, ROM 2217, or RAM 2221. A user may enter commands and information into the computer system 2200 through input devices such as a keyboard 2266 and a pointing device 2268, such as a mouse. Other input devices (not shown) may include a microphone, joystick, game pad, satellite dish, scanner, trackball, touch pad, touch screen, touch-sensing device, voice command module or device, user motion or user gesture capture device, or the like. These and other input devices are often connected to the processor 2205 through a serial port interface 2271 coupled to the system bus 2214, but may also be connected by other interfaces, such as a parallel port, game port, or universal serial bus (USB). A monitor 2273 or other type of display device is also connected to the system bus 2214 via an interface, such as a video adapter 2275. In addition to the monitor 2273, personal computers typically include other peripheral output devices (not shown), such as speakers and printers. The illustrative example shown in Figure 22 also includes a host adapter 2278, a small computer system interface (SCSI) bus 2283, and an external storage device 2276 connected to the SCSI bus 2283.

[0043]

[0062] Computer system 2200 can operate in a networked environment using logical connections to one or more remote computers, such as remote computer 2288. The remote computer 2288 may be selected as another personal computer, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the elements described above in relation to computer system 2200, although only a single representative remote memory / storage device 2290 is shown in Figure 22. The logical connections depicted in Figure 22 include a local area network (LAN) 2293 and a wide area network (WAN) 2295. Such networked environments are often deployed in, for example, offices, enterprise-wide computer networks, intranets and the Internet.

[0044]

[0063] When used in a LAN networking environment, the computer system 2200 is connected to the local area network 2293 through a network interface or adapter 2296. When used in a WAN networking environment, the computer system 2200 typically includes a broadband modem 2298, a network gateway, or other means for establishing communications over a wide area network 2295, such as the Internet. The broadband modem 2298, which may be internal or external, is connected to the system bus 2214 via a serial port interface 2271. In a networked environment, program modules relating to the computer system 2200, or portions thereof, may be stored in a remote memory storage device 2290. It should be noted that the network connections shown in FIG. 22 are for illustrative purposes only, and other means of establishing a communications link between computers may be used, depending on the particular requirements of the application of the present technique for setting camera focus in a mixed reality environment with hand gesture interaction.

[0045]

[0064] FIG. 23 is a high-level block diagram of an exemplary data center 2300 that provides cloud or distributed computing services that can be used to implement the present techniques for setting camera focus in a mixed reality environment with hand gesture interaction. For example, an HMD device 105 can utilize the solutions provided by the data center 2300, such as receiving streaming content. Multiple servers 2301 are managed by a data center management controller 2302. A load balancer 2303 distributes request and computational workloads across the servers 2301 to avoid situations where a single server may become overwhelmed. The load balancer 2303 maximizes the available capacity and performance of resources within the data center 2300. A router / switch 2304 supports data traffic between the servers 2301 and between the data center 2300 and external resources and users (not shown) via an external network 2305, which may be, for example, a local area network (LAN) or the Internet.

[0046]

[0065] Servers 2301 may be standalone computing devices and / or may be configured as individual blades within a rack of one or more server devices. Servers 2301 have input / output (I / O) connectors 2306 that manage communication with other database entities. One or more host processors 2307 on each server 2301 run a host operating system (O / S) 2308 that supports multiple virtual machines (VMs) 2309. Each VM 2309 can run its own O / S, such that each VM O / S 2310 on a server is different, the same, or a mixture of both. VM O / S 2310 may, for example, be different versions of the same O / S (e.g., different VMs running different current and past versions of the Windows® operating system). Additionally or alternatively, VM O / S 2310 may be provided by different manufacturers (e.g., some VMs run the Windows® operating system, while other VMs run the Linux® operating system). Each VM 2309 may also run one or more applications (Apps) 2311. Each server 2301 also includes storage 2312 (e.g., a hard disk drive (HDD)) and memory 2313 (e.g., RAM) that can be accessed and used by the host processor 2307 and VMs 2309 to store software code, data, etc. In one embodiment, the VMs 2309 may utilize the data plane APIs disclosed herein.

[0047]

[0066] Datacenter 2300 provides pooled resources that allow customers to dynamically provision or scale applications as needed without having to add servers or additional network connections. This allows customers to obtain the computing resources they need without having to acquire, provision, and manage infrastructure for each application on an ad-hoc basis. Cloud computing datacenter 2300 allows customers to dynamically scale resources up or down to meet the current needs of their business. In addition, datacenter operators can offer usage-based services to customers so that they pay for the resources they use only as they need them. For example, a customer may initially use one VM 2309 on server 23011 to run their application 2311. As demand for application 2311 increases, datacenter 2300 can scale resources on the same server 23011 and / or on new servers 2301 as needed. N Further VMs 2309 can be started on the application. These further VMs 2309 can be disabled later if the demand for the application subsides.

[0048]

[0067] Datacenter 2300 can provide guaranteed availability, disaster recovery, and backup services. For example, a datacenter can designate one VM 2309 on server 23011 as the primary location for a customer's application and can launch a second VM 2309 on the same or a different server as a standby or backup if the first VM or server 23011 fails. Datacenter management controller 2302 automatically shifts incoming user requests from the primary VM to the backup VM without requiring customer intervention. While datacenter 2300 is shown as a single location, it should be understood that server 2301 can be distributed across multiple locations across the globe to provide additional redundancy and disaster recovery capabilities. Additionally, datacenter 2300 can be an on-premise private system serving a single enterprise user, a publicly accessible distributed system serving multiple unrelated customers, or a combination of both.

[0049]

[0068] Domain Name System (DNS) server 2314 translates domains and hostnames into IP (Internet Protocol) addresses for all roles, applications, and services within data center 2300. DNS log 2315 maintains a record of which domain names have been translated by roles. It should be understood that DNS is used herein as an example, and that other name translation services and domain name logging services may be used to identify dependencies.

[0050]

[0069] Data center health monitor 2316 monitors the health of the physical systems, software, and environment within data center 2300. Health monitor 2316 provides feedback to data center managers when problems are detected with servers, blades, processors, or applications within data center 2300, or when network bandwidth or communication issues arise.

[0051]

[0070] 24 illustrates an example architecture 2400 of a computing device, such as a smartphone, tablet computer, laptop computer, or personal computer, for the present technique for setting camera focus in a mixed reality environment with hand gesture interaction. The computing device of FIG. 24 may also be an alternative to the HMD device 110, which may benefit from reduced hunting in the autofocus subsystem. While several components are depicted in FIG. 24 , other components disclosed herein, but not shown, are possible with the computing device.

[0052]

[0071] The architecture 2400 shown in FIG. 24 includes one or more processors 2402 (e.g., a central processing unit, a dedicated artificial intelligence chip, a graphics processing unit, etc.), a system memory 2404 including RAM (random access memory) 2406 and ROM (read-only memory) 2408, and a system bus 2410 that operatively and functionally couples the components in the architecture 2400. A basic input / output system (BIS), containing the basic routines that help transfer information between elements within the architecture 2400, such as during start-up, is typically stored in ROM 2408. The architecture 2400 further includes a mass storage device 2412 for storing software code or other computer-executable code utilized to implement applications, a file system, and an operating system. The mass storage device 2412 is connected to the processor 2402 through a mass storage controller (not shown) connected to the bus 2410. The mass storage device 2412 and its associated computer-readable storage media provide non-volatile storage for the architecture 2400. Although the descriptions of computer-readable storage media included in this specification refer to mass storage devices such as hard disks or CD-ROM drives, those skilled in the art will appreciate that computer-readable storage media may be any available storage media that can be accessed by architecture 2400.

[0053]

[0072] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. For example, computer-readable media includes, without limitation, RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), Flash memory or other semiconductor memory technology, CD-ROM, DVD, HD-DVD (High Definition DVD), Blu-Ray or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by architecture 2400.

[0054]

[0073] According to various embodiments, architecture 2400 can operate in a networked environment by using logical connections to remote computers through a network. Architecture 2400 can connect to the network through a network interface unit 2416 connected to bus 2410. It should be understood that network interface unit 2416 can also be utilized to connect to other types of networks and remote computer systems. Architecture 2400 can also include an input / output controller 2418 for receiving and processing input from a number of other devices, including controls such as a keyboard, mouse, touchpad, touchscreen, buttons and switches, or an electronic stylus (not shown in FIG. 24 ). Similarly, input / output controller 2418 can provide output to a display screen, user interface, printer, or other type of output device (also not shown in FIG. 24 ).

[0055]

[0074] It can be appreciated that the software components described herein, when loaded and executed within processor 2402, can transform processor 2402 and architecture 2400 as a whole from a general-purpose computing system into a special-purpose computing system customized to facilitate the functions presented herein. Processor 2402 can be constructed from any number of transistors or other discrete circuit elements, which can individually or collectively assume any number of states. More particularly, processor 2402 can operate as a finite state machine in response to executable instructions contained within the software modules disclosed herein. These computer-executable instructions define the manner in which processor 2402 transitions between states, thereby transforming processor 2402 by transforming the transistors or other discrete hardware elements that make up processor 2402.

[0056]

[0075] Encoding the software modules presented herein may also transform the physical structure of the computer-readable storage media presented herein. The specific transformation of the physical structure may depend on various factors in different implementations of this description. Examples of such factors may include, without limitation, the technique used to implement the computer-readable storage medium, whether the computer-readable storage medium is characterized as primary or secondary storage, and the like. For example, if the computer-readable storage medium is implemented as a semiconductor-based memory, the software disclosed herein may be encoded on the computer-readable storage medium by transforming the physical state of the semiconductor memory. For example, the software may transform the state of transistors, capacitors, or other individual circuit elements that make up the semiconductor memory. The software may also transform the physical state of such components in order to store data thereon.

[0057]

[0076] As another example, the computer-readable storage media disclosed herein may be implemented using magnetic or optical technology. In such implementations, the software presented herein may transform the physical state of the magnetic or optical media when the software is encoded therein. These transformations may include altering the magnetic properties of specific locations within a given magnetic media. These transformations may also include altering the physical features or characteristics of specific locations within a given optical media to change the optical properties of those locations. Other transformations of physical media are possible without departing from the scope and spirit of this description, and the above examples are provided solely to facilitate this description.

[0058]

[0077] In view of the foregoing, it can be appreciated that many types of physical transformations occur within architecture 2400 to store and execute the software components presented herein. It can also be appreciated that architecture 2400 can include other types of computing devices, including wearable devices, handheld computers, embedded computer systems, smartphones, PDAs, and other types of computing devices known to those skilled in the art. It is also contemplated that architecture 2400 may not include all of the components shown in FIG. 24, may include other components not explicitly shown in FIG. 24, or may utilize an entirely different architecture than that shown in FIG. 24.

[0059]

[0078] 25 shows one particular illustrative example of a see-through mixed reality display system 2500, and FIG. 26 shows a functional block diagram of the system 2500. The illustrative display system 2500 provides a complementary description to the HMD device 110 depicted throughout the figures. The display system 2500 includes one or more lenses 2502 forming part of a see-through display subsystem 2504, such that images may be displayed using the lenses 2502 (e.g., using projection onto the lenses 2502, using one or more waveguide systems incorporated in the lenses 2502, and / or by any other suitable manner). The display system 2500 further includes one or more outward-facing image sensors 2506 configured to acquire images of the background scene and / or physical environment viewed by the user, and may include one or more microphones 2508 configured to detect sound, such as voice commands, from the user. The outward-facing image sensor 2506 may include one or more depth sensors and / or one or more two-dimensional image sensors. In an alternative embodiment, as described above, the mixed reality or virtual reality display system may display mixed reality or virtual reality images through a viewfinder mode of the outward-facing image sensor instead of incorporating a see-through display subsystem.

[0060]

[0079] Display system 2500 may further include a gaze detection subsystem 2510 configured to detect the direction of gaze or the direction or location of focus of each of the user's eyes, as described above. Gaze detection subsystem 2510 may be configured to determine the gaze direction of each of the user's eyes in any suitable manner. For example, in the illustrated illustrative example, gaze detection subsystem 2510 includes one or more light flash sources 2512, such as infrared light sources, configured to cause a flash of light to reflect off each of the user's eyes, and one or more image sensors 2514, such as inward-facing sensors, configured to capture images of each of the user's eyes. Light flashes from the user's eyes and / or changes in the location of the user's pupils, determined from image data collected using image sensor 2514, may be used to determine the direction of gaze.

[0061]

[0080] Additionally, the location where the line of sight projected from the user's eye intersects with the external display can be used to determine the object the user is gazing at (e.g., the displayed virtual object and / or the real background object). The gaze detection subsystem 2510 can have any suitable number and configuration of light sources and image sensors. In some implementations, the gaze detection subsystem 2510 can be omitted.

[0062]

[0081] Display system 2500 may also include additional sensors. For example, display system 2500 may include a global positioning (GPS) subsystem 2516 to allow the location of display system 2500 to be determined. This may assist in identifying real-world objects, such as buildings, that may be located within the user's immediate physical environment.

[0063]

[0082] The display system 2500 may further include one or more motion sensors 2518 (e.g., inertial, multi-axis gyroscope, or accelerometer) for detecting the movement and position / orientation / pose of the user's head when the user wears the system as part of a mixed reality or virtual reality HMD device. The motion data, potentially along with eye tracking phosphene data and outward facing image data, can be used for image stabilization as well as gaze detection to help correct for blur in images from the outward facing image sensor 2506. The use of motion data can allow tracking of changes in gaze location even when the image data from the outward facing image sensor 2506 cannot be resolved.

[0064]

[0083] Additionally, the microphone 2508 and gaze detection subsystem 2510, as well as the motion sensor 2518, can be utilized as user input devices such that a user can interact with the display system 2500 not only via eye, neck, and / or head gestures, but also, in some cases, via verbal commands. It can be understood that the sensors shown in Figures 25 and 26 and described in the accompanying text are included for illustrative purposes and are in no way intended to be limiting, as any other suitable sensor and / or combination of sensors can be utilized to meet the needs of a particular implementation. Some implementations can utilize, for example, biometric sensors (e.g., for detecting heart and respiratory rate, blood pressure, brain activity, body temperature, etc.) or environmental sensors (e.g., for detecting temperature, humidity, altitude, UV (ultraviolet) light levels, etc.).

[0065]

[0084] The display system 2500 may further include a controller 2520 having a logic subsystem 2522 and a data storage subsystem 2524 in communication with the sensors, gaze detection subsystem 2510, display subsystem 2504, and / or other components through a communication subsystem 2526. The communication subsystem 2526 may also facilitate the display system operating in conjunction with remotely located resources such as processing, storage, power, data, and services. That is, in some implementations, the HMD device may operate as part of a system in which resources and capabilities may be distributed among different components and subsystems.

[0066]

[0085] The storage subsystem 2524 may contain instructions stored thereon that are executable by the logic subsystem 2522, for example, to receive and interpret input from sensors, identify the user's location and movement, identify real-world objects using surface reconstruction and other techniques, and dim / fade the display based on the distance to an object to allow the object to be observed by the user, among other tasks.

[0067]

[0086] The display system 2500 can be configured with one or more audio transducers 2528 (e.g., speakers, earphones, etc.) so that audio can be utilized as part of the mixed reality or virtual reality experience. The power management subsystem 2530 can include one or more batteries 2531 and / or protection circuit modules (PCMs) and associated charger interfaces 2534 and / or remote power interfaces to provide power to components within the display system 2500.

[0068]

[0087] It can be understood that display system 2500 is described for purposes of illustration and is therefore not intended to be limiting. It should be further understood that the display device can include additional and / or alternative sensors, cameras, microphones, input devices, output devices, etc., other than those shown, without departing from the scope of the present arrangements. Additionally, the physical arrangement of the display device and its various sensors and subcomponents can have a variety of different forms without departing from the scope of the present arrangements.

[0069]

[0088] Various exemplary embodiments of the present technique for setting camera focus in a mixed reality environment with hand gesture interaction are presented here for purposes of illustration and not as an all-inclusive list of all embodiments. One example includes a method performed by a head-mounted display (HMD) to optimize autofocus implementation, the method including: enabling autofocus operation of a camera in the HMD device configured to capture a scene in a local physical environment surrounding the HMD device over a field of view (FOV), the camera being a member of a set of one or more sensors operatively coupled to the HMD device; acquiring data related to a user's hand using the set of sensors; and inhibiting the autofocus operation of the camera based on the collected data related to the user's hand failing to satisfy one or more criteria of the autofocus subsystem.

[0070]

[0089] In another example, the set of sensors collects data describing a local physical environment in which the HMD device operates. In another example, the HMD device includes a see-through mixed reality display through which a local user observes the local physical environment, and on which the HMD device renders one or more virtual objects. In another example, a scene captured by a camera across an FOV and the rendered virtual objects are transmitted as content to a remote computing device. In another example, the scene captured by a camera across the FOV and the rendered virtual objects are mixed by the HMD device into a composite signal that is recorded. In another example, the method further includes specifying a region of interest (ROI) within the FOV, and criteria for the autofocus subsystem include collected data indicating that one or more of the user's hands are local within the ROI. In another example, the ROI includes a three-dimensional space within the local physical environment. In another example, the ROI is dynamically variable in at least one of size, shape, or location. In another example, the method further includes evaluating one or more hand characteristics to determine whether the collected data about the user's hands meets one or more criteria of the autofocus subsystem. In another example, the hand characteristics include a pace of hand movement. In another example, the hand characteristics include which part of the hand. In another example, the hand characteristics include a duration for which the one or more hands are positioned within the ROI. In another example, the hand characteristics include one or more hand sizes. In another example, the hand characteristics include one or more hand poses. In another example, the hand characteristics include one or more hand motion directions. In another example, the camera includes a PV (photo / video) camera, and the set of sensors includes a depth sensor configured to collect depth data within the local physical environment and thereby track one or more of the HMD device user's hands.

[0071]

[0090] A further example is one or more hardware-based non-transitory computer-readable memory devices storing computer-readable instructions that, when executed by one or more processors in the computing device, cause the computing device to capture data to track one or more hands of a user while the user is using the computing device in a physical environment; select a region of interest (ROI) within a field of view (FOV) of the computing device, where the computing device renders one or more virtual objects on a see-through display coupled to the computing device to enable the user to simultaneously view the physical environment and the one or more virtual objects as a mixed reality user experience; and select from the captured hand tracking data portions of the user's one or more hands in the FOV. The computing device includes one or more hardware-based non-transitory computer-readable memory devices configured to determine whether a user's hand is located within the OI and, in response to the determining, trigger or disable autofocus operation of a camera disposed on the computing device, the camera configured to capture a scene including at least a portion of the physical environment within the FOV, the autofocus operation being triggered in response to one or more hand characteristics derived from the captured hand tracking data indicating that the user's one or more hands are a focus position of the user within the ROI, and the autofocus operation being disabled in response to one or more hand characteristics derived from the captured hand tracking data indicating that the user's one or more hands are a focus position of the user within the ROI.

[0072]

[0091] In another example, the captured hand tracking data is from a forward-facing depth sensor operably coupled to the computing device. In another example, the computing device includes a head-mounted display (HMD) device, a smartphone, a tablet computer, or a portable computer. In another example, the ROI is located in a central region of the FOV. In another example, the executed instructions cause the computing device to adjust the hand tracking data captured in the physical environment within the ROI of the FOV.

[0073]

[0092] A further example includes a computing device configured to be worn on a user's head, the computing device configured to reduce undesirable hunting phenomena in an autofocus subsystem associated with the computing device, the computing device including a display configured to render holograms, an adjustable-focus PV (photo / video) camera operatively coupled to the autofocus subsystem and configured to capture adjustably focused images of a physical environment in which the user is located, a depth sensor configured to capture depth data related to the physical environment in three dimensions, one or more processors, and one or more hardware-based memory devices storing computer-readable instructions, the instructions being executed by the one or more processors. and causing a computing device to render at least one hologram on a display, the hologram including a virtual object positioned in the physical environment at a known location; capturing location tracking data of one or more hands of a user in the physical environment using a depth sensor; determining from the location tracking data whether one or more hands of the user are interacting with a hologram at the known location; triggering operation of an autofocus subsystem in response to determining that one or more hands of the user are not interacting with the hologram; and inhibiting operation of the autofocus subsystem in response to determining that one or more hands of the user are interacting with the hologram.

[0074]

[0093] In another example, operation of the autofocus subsystem is inhibited in response to a determination from the location tracking data that one or more of the user's hands interact with the hologram using an inward or outward motion.

[0075]

[0094] Although the subject matter has been described in language specific to structural features and / or method acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. 1. A method performed by a head-mounted display (HMD) device operable by a user to optimize an autofocus implementation, comprising: Enabling autofocus operation of a camera disposed within the HMD device, the camera being configured to capture a scene in a local physical environment surrounding the HMD device over a field of view (FOV), the camera being a member of a set of one or more sensors operably coupled to the HMD device; providing a gaze detection subsystem including one or more sensors within the HMD device configured to detect a position of a focal point within the FOV of one or more eyes of the user; Specifying a region of interest (ROI) within the FOV for controlling the autofocus operation; determining whether the user is interacting with one or more virtual objects within the ROI; using the detected position of the user's focus to inhibit the autofocus operation of the camera in a portion of a scene in the local physical environment in which the one or more virtual objects are rendered during interaction with the one or more virtual objects; A method comprising:

2. The method of claim 1 , wherein the set of sensors collects data describing the local physical environment.

3. 3. The method of claim 2, wherein the HMD device includes a see-through mixed reality display through which a local user observes the local physical environment, and on which the HMD device renders one or more virtual objects.

4. The method of claim 3 , wherein the scene captured by the camera across the FOV and the rendered virtual objects are transmitted as content to a remote computing device.

5. The method of claim 3 , wherein the scene and the rendered virtual objects captured by the camera across the FOV are mixed into a composite signal that is recorded by the HMD device.

6. The method of claim 2 , wherein the autofocus operation is triggered based at least in part on the local physical environment within the FOV.

7. The method of claim 1 , wherein the ROI comprises a three-dimensional space within the local physical environment.

8. The method of claim 1 , wherein the ROI is dynamically variable in at least one of size, shape, or position.

9. 2. The method of claim 1, wherein the set of sensors includes a depth sensor, and further comprising capturing one or more characteristics of an HMD device user's hands using the depth sensor to track one or more of the user's hands to determine whether to trigger or inhibit the autofocus operation of the camera.

10. The method of claim 9 , wherein the characteristics of the hand include a pace of hand movement.

11. The method of claim 9 , wherein the characteristics of the hand include which part of the hand it is.

12. The method of claim 9 , wherein the characteristics of the hands include a duration for which the one or more hands are positioned within the ROI.

13. The method of claim 1 , wherein the camera comprises a PV (photo / video) camera.

14. One or more hardware-based non-transitory computer-readable memory devices storing computer-readable instructions that, when executed by one or more processors within a computing device available to a user, cause the computing device to: using a gaze detection subsystem to capture data about the position of focus of the user's eyes while the user utilizes the computing device within a physical environment; identifying a region of interest (ROI) within a field of view (FOV) of the computing device for controlling autofocus operation of a forward-facing camera disposed on the computing device, the ROI being configured to capture a scene including at least a portion of the physical environment within the FOV, the computing device rendering one or more virtual objects on a see-through display coupled to the computing device to enable the user to simultaneously view the physical environment and the one or more virtual objects as a mixed reality user experience; determining whether the user is interacting with a virtual object within the ROI; determining the extent to which the user interacts with the virtual object from the captured positional data of the user's focus; responsive to said determining, triggering or disabling said autofocus operation; Let them do this, the autofocus operation is triggered in response to a determination that a user's interaction with the virtual object is the user's focal point within the ROI; and one or more hardware-based non-transitory computer-readable memory devices, wherein the autofocus operation is disabled in response to determining that a user's interaction with the virtual object is transient within the ROI.

15. 15. The one or more hardware-based non-transitory computer-readable memory devices of claim 14, wherein the instructions further cause the computing device to capture hand tracking data from a depth sensor operatively coupled to the computing device and to trigger or disable autofocus operation of the camera using the captured hand tracking data.

16. 15. The one or more hardware-based non-transitory computer-readable memory devices of claim 14, wherein the computing device comprises a head-mounted display (HMD) device, a smartphone, a tablet computer, or a handheld computer.

17. The one or more hardware-based non-transitory computer-readable memory devices of claim 14 , wherein the ROI is located in a central region of the FOV.

18. 20. The one or more hardware-based non-transitory computer-readable memory devices of claim 17, wherein the gaze detection subsystem includes an inward-facing sensor configured to capture images of one or more user's eyes.

19. 1. A computing device configurable to be worn on a user's head, the computing device configured to reduce undesirable hunting phenomena in an autofocus subsystem associated with the computing device, the computing device comprising: a display configured to render holograms; and an adjustable-focus PV (photo / video) camera operatively coupled to the autofocus subsystem and configured to capture adjustably focused images of a physical environment in which the user is located; a gaze detection subsystem configured to detect the user's gaze in the physical environment, including one or more of a direction of gaze, a direction of focus, or a location of focus; one or more processors; and one or more hardware-based memory devices storing computer-readable instructions that, when executed by the one or more processors, cause the computing device to: Rendering at least one hologram on the display, the hologram including a virtual object positioned in the physical environment at a known position; capturing gaze tracking data about the user's gaze in the physical environment using the gaze detection subsystem; determining from the gaze tracking data whether one or more hands of the user are interacting with a hologram at the known location; triggering operation of the autofocus subsystem in response to determining that the user is not interacting with the hologram; inhibiting operation of the autofocus subsystem in response to determining that the user is interacting with the hologram.

20. 20. The computing device of claim 19, wherein operation of the autofocus subsystem is inhibited in response to a determination from the gaze tracking data that the user's interaction with the hologram includes inward and outward movements of one or more of the user's hands.

21. 1. A method performed by a head-mounted display (HMD) device operable by a user to optimize an autofocus implementation, comprising: Enabling autofocus operation of a camera in the HMD device, the camera being configured to capture a scene in a local physical environment surrounding the HMD device over a field of view (FOV); Detecting the position of the user's eye focus within the FOV using a gaze detection subsystem including a sensor within the HMD device; determining whether the user is interacting with a virtual object within the FOV; During interaction with the virtual object, using the user's focal point position within the FOV to inhibit the autofocus operation of the camera in a portion of a scene within the local physical environment in which the virtual object is rendered; A method comprising:

22. Enabling autofocus operation of a camera in a user's head-mounted display (HMD) device, the camera being configured to capture a scene in a physical environment surrounding the HMD device over a field of view (FOV); determining whether the user is interacting with a virtual object within the FOV; using a position of the user's focal point within the FOV during interaction with the virtual object to inhibit the autofocus operation of the camera in a portion of a scene within the physical environment in which the virtual object is rendered; Rendering content from the user's HMD device on a remote user's computing device, the content including the interaction with the virtual object; and A method comprising:

23. 23. The method of claim 21 or 22, wherein the virtual objects include a tag identifying a company, directions to a point of interest, and a gift box.

24. The method of claim 21 , wherein the virtual object is positioned relative to the local physical environment.

25. The method of claim 21 or 22, wherein the interaction with the virtual object comprises moving, selecting, shrinking or enlarging the virtual object.

26. The method of claim 21 or 22, wherein the interaction occurs at a spatial distance away from the location of the virtual object.

27. 22. The method of claim 21, further comprising rendering content from the HMD device to a computing device, the content including the user's interactions.

28. 23. The method of claim 21 or 22, wherein the FOV includes a statically or dynamically defined region of interest (ROI) for controlling the autofocus operation, the ROI being unaware of the user.

29. 30. The method of claim 28, wherein the autofocus operation of the camera includes automatically focusing on content within the ROI.

30. The method of claim 22 , wherein the virtual object is positioned relative to the user.

31. A head-mounted display (HMD) device configurable to be worn on a user's head, comprising: a camera configured to capture a scene in a physical environment surrounding the HMD device over a field of view (FOV); a processor; and a memory storing instructions that, when executed by the processor, cause the HMD device to: enabling autofocus operation of the camera; Detecting a position of the user's focal point within the FOV; determining whether the user is interacting with a virtual object within the FOV; and during interaction with the virtual object, using the detected position of the user's focus to inhibit the autofocus operation of the camera in a portion of a scene in the physical environment in which the virtual object is rendered.

32. The HMD device of claim 31 , wherein the interaction occurs at a spatial distance away from the location of the virtual object.

33. The HMD device of claim 31 , wherein the instructions, when executed by the processor, further cause the HMD device to render content from the HMD device of the user to a computing device of a remote user, the content including rendering the interaction of the user on the computing device of the remote user.

34. 32. The HMD device of claim 31, wherein the FOV includes a statically or dynamically defined region of interest (ROI) for controlling the autofocus operation, the ROI being unaware of the user.

35. 35. The HMD device of claim 34, wherein the autofocus operation of the camera includes automatically focusing on content within the ROI.

Citation Information

Patent Citations

  • Small DC motor

    JP1988056157A

  • Information processing apparatus and information processing method, and computer program

    JP2016194744A

  • Object placement based on gaze in a virtual reality environment

    JP2017530438A

  • Autonomous computing and telecommunications head-up displays glasses

    US20140266988A1

  • Camera auto-focus based on eye gaze

    US20150003819A1