Responding to user requests related to images

JP2026145020APending Publication Date: 2026-09-09APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026028430
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2026-01-26
Filing Date
2026-02-25
Publication Date
2026-09-09

Smart Images

  • Figure 2026145020000001_ABST
    Figure 2026145020000001_ABST
Patent Text Reader

Abstract

Regarding responding to user requests related to images. [Solution] This specification discloses an example process for capturing images in a multi-device system. The example method determines whether an input image meets quality standards, and if the input image does not meet the quality standards, a prompt is displayed on a device other than the device that acquired the input image.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] (Cross-Reference to Related Application) This application relates to the following co-pending provisional application, namely U.S. Provisional Patent Application No. 63 / 765,276 entitled “MULTIDEVICE CAMERA SELECTION” filed on February 28, 2025, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure generally relates to responding to user requests associated with images. [[Background Art]]

[0003] The development of computer systems for interacting with and / or providing three-dimensional scenes has expanded significantly in recent years. Examples of three-dimensional scenes (e.g., environments) include physical scenes and augmented reality scenes. [[Summary of the Invention]]

[0004] Exemplary methods are disclosed herein. An example of a method comprises, in a first computer system in communication with one or more image sensors: acquiring a first image using the one or more image sensors; receiving a user request associated with the first image; and in response to acquiring the first image and receiving the user request, causing a second computer system to provide a prompt to capture a second image using the second computer system in accordance with a determination that a quality of the first image does not satisfy a quality criterion, and generating a response to the user request based on the first image and providing an output including the response to the user request based on the first image in accordance with a determination that the quality of the first image satisfies the quality criterion.

[0005] Examples of non-temporary computer-readable storage media are disclosed herein. An example of a non-temporary computer-readable storage medium stores one or more programs. The one or more programs are configured to be executed by one or more processors of a first computer system communicating with one or more image sensors. The one or more programs include instructions for acquiring a first image using one or more image sensors, receiving a user request related to the first image, and, in response to acquiring the first image and receiving the user request, prompting a second computer system to capture a second image using the second computer system, according to a determination that the quality of the first image does not meet a quality standard, and, according to a determination that the quality of the first image meets a quality standard, generating a response to the user request based on the first image, and providing an output including the response to the user request based on the first image.

[0006] Examples of computer systems are disclosed herein. An example of a first computer system is configured to communicate with one or more image sensors. The first computer system comprises one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for acquiring a first image using one or more image sensors, receiving a user request related to the first image, and, in response to acquiring the first image and receiving the user request, prompting a second computer system to capture a second image using the second computer system in accordance with a determination that the quality of the first image does not meet a quality standard, and, in accordance with a determination that the quality of the first image meets a quality standard, generating a response to the user request based on the first image, and providing an output including the response to the user request based on the first image.

[0007] An example of a first computer system is configured to communicate with one or more image sensors. The first computer system includes means for acquiring a first image using one or more image sensors, means for receiving user requests related to the first image, and means for prompting a second computer system to capture a second image using the second computer system in response to acquiring a first image and receiving a user request, according to a determination that the quality of the first image does not meet quality standards, and for generating a response to the user request based on the first image, according to a determination that the quality of the first image meets quality standards, and providing an output including the response to the user request based on the first image.

[0008] An example of a computer program product includes one or more programs configured to run by one or more processors of a first computer system communicating with one or more image sensors. The one or more programs include instructions for acquiring a first image using one or more image sensors, receiving a user request related to the first image, and, in response to acquiring the first image and receiving the user request, prompting a second computer system to capture a second image using the second computer system, according to a determination that the quality of the first image does not meet quality standards, and, according to a determination that the quality of the first image meets quality standards, generating a response to the user request based on the first image, and providing output including the response to the user request based on the first image.

[0009] By prompting a second computer system to capture a second image when the first image does not meet quality standards, smoother interaction between the user and multiple devices of the system becomes possible. Specifically, during the process of taking a photograph, the second computer system can indicate to the user that the second computer system can automatically take another photograph that meets the quality standards. This provides the user with more information and reduces the number of inputs the user needs to provide to capture an image in order to complete the requested task. In this way, the efficiency and accuracy of user-device interaction are improved (for example, by reducing the number of inputs required to capture a suitable image and by providing the user with additional information about device capabilities), resulting in reduced power consumption and improved device battery life as the user can use the device more quickly and efficiently.

[0010] An example of a method is disclosed herein. An example of a method is a computer system communicating with one or more visual imaging sensors and one or more audio output devices, which includes receiving natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing region of a three-dimensional (3D) scene; capturing image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors; providing a first audio output corresponding to the first object via one or more audio output devices in response to receiving the natural language input corresponding to the first object in the 3D scene, according to a determination that a visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene satisfies a condition; and providing a second audio output corresponding to the first object via one or more audio output devices according to a determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy a condition.

[0011] Examples of non-temporary computer-readable storage media are disclosed herein. An example of a non-temporary computer-readable storage medium stores one or more programs. The one or more programs are configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices. The one or more programs include instructions for receiving natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing region of the 3D scene; capturing image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors; and, in response to receiving natural language input corresponding to the first object in the 3D scene, providing a first audio output corresponding to the first object via one or more audio output devices according to a determination that a visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene satisfies a condition; and providing a second audio output corresponding to the first object via one or more audio output devices according to a determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy a condition.

[0012] An example of a computer system is disclosed herein. The example of a computer system is configured to communicate with one or more visual imaging sensors and one or more audio output devices. The computer system comprises one or more processors and a memory storing one or more programs configured to be executed by one or more processors, wherein one or more programs include instructions for receiving natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing region of a three-dimensional (3D) scene; capturing image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors; and providing a first audio output corresponding to the first object via one or more audio output devices in response to receiving natural language input corresponding to the first object in the 3D scene, according to a determination that a visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene satisfies a condition; and providing a second audio output corresponding to the first object via one or more audio output devices according to a determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy a condition.

[0013] An example of a computer system is configured to communicate with one or more visual imaging sensors and one or more audio output devices. The computer system includes means for receiving natural language input corresponding to a first object in a 3D scene while the head of the computer system user has a head orientation corresponding to the forward-facing region of the 3D scene; means for capturing image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors; and means for providing a first audio output corresponding to the first object via one or more audio output devices in response to receiving natural language input corresponding to the first object in the 3D scene, according to a determination that a visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene satisfies a condition, and providing a second audio output corresponding to the first object via one or more audio output devices in accordance with a determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy a condition.

[0014] An example of a computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices. The one or more programs include instructions for receiving natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing region of a three-dimensional (3D) scene; capturing image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors; and, in response to receiving the natural language input corresponding to the first object in the 3D scene, providing a first audio output corresponding to the first object via one or more audio output devices according to a determination that a visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene satisfies a condition; and providing a second audio output corresponding to the first object via one or more audio output devices according to a determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy a condition.

[0015] By providing audio output based on whether the visibility metric meets the criteria, the computer system can more accurately and efficiently satisfy user requests regarding objects present in a 3D scene. For example, if the visibility metric does not meet the criteria, the captured image data may not depict the object related to the user request, and therefore the computer system may not be able to satisfy the user request based on the captured image data. As described herein, the computer system may therefore provide one or more audio outputs prompting the user to specify the object related to the user request and / or prompting the user to capture an image of the object with a different device, thereby enabling the computer system to accurately satisfy the user request based on new image data depicting the object in question. As another example, if the visibility metric meets the criteria, the captured image data may depict the object related to the user request, and therefore the computer system can satisfy the user request based on the captured image data. In that case, the computer system may provide an audio output that satisfies the user request. In this way, the accuracy and efficiency of the user-device interface are improved (for example, by enabling the device to respond accurately to user requests regarding objects in a 3D scene, by preventing the electronic device from providing incorrect responses to user requests regarding objects in a 3D scene, by reducing the number of user inputs required for the electronic device to fulfill user requests, and by reducing the number of inputs required to undo and / or cancel the results of misinterpreted user requests), resulting in reduced power consumption and improved battery life of the device, as users can use the device more quickly and efficiently.

[0016] In some examples, the computer system is a desktop computer with an associated display. In some examples, the computer system is a portable device (e.g., a handheld device such as a notebook computer, tablet computer, or smartphone). In some examples, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or head-mounted device). In some examples, the computer system has a touchpad. In some examples, the computer system has one or more cameras. In some examples, the computer system has a display generating component (e.g., a display device such as a head-mounted display, display, projector, touch-sensitive display (also known as a “touchscreen” or “touchscreen display”), or other devices or components that present visual content to the user on or within the display generating component itself, or that are generated from the display generating component and are visible elsewhere). In some examples, the computer system does not have a display generating component and does not present visual content to the user. In some examples, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some examples, the computer system has one or more eye-tracking components. In some examples, the computer system has one or more hand-tracking components. In some examples, the computer system has one or more output devices, the output devices include one or more tactile output generators and / or one or more audio output devices. In some examples, the computer system has one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing the various functions described herein.In some examples, the user interacts with the computer system through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands or body in space captured by cameras and other motion sensors, and / or voice input captured by one or more audio input devices. The executable instructions that perform these functions are optionally contained in temporary computer-readable storage media and / or non-temporary computer-readable storage media, or other computer program products configured to be executed by one or more processors.

[0017] It should be noted that the various examples described herein can be combined with any other examples described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention.

[0018] For a better understanding of the various embodiments described, please refer to the following “Modes for Carrying Out the Invention” in conjunction with the following drawings, where similar reference numbers refer to corresponding parts throughout those drawings. [Brief explanation of the drawing]

[0019] [Figure 1] This block diagram illustrates the operating environment of a computer system for interacting with a three-dimensional (3D) scene, using several examples.

[0020] [Figure 2] This is a block diagram of user-responsive components of a computer system, with several examples.

[0021] [Figure 3A] Here are some examples of block diagrams of computer system controllers.

[0022] [Figure 3B] FIG. 1 is a block diagram of an image evaluation unit of a computer system, according to some examples.

[0023] [Figure 4] FIG. 2 is a diagram illustrating an architecture of a foundation model, according to some examples.

[0024] [Figure 5A] FIG. 3 is a diagram illustrating image capture in a multi-device system, according to some examples. [Figure 5B] FIG. 4 is a diagram illustrating image capture in a multi-device system, according to some examples. [Figure 5C] FIG. 5 is a diagram illustrating image capture in a multi-device system, according to some examples. [Figure 5D] FIG. 6 is a diagram illustrating image capture in a multi-device system, according to some examples. [Figure 5E] FIG. 7 is a diagram illustrating image capture in a multi-device system, according to some examples.

[0025] [Figure 6] FIG. 8 is a flow diagram of a method for capturing an image in a multi-device system, according to some examples.

[0026] [Figure 7] FIG. 9 is a diagram illustrating a forward head posture determined based on the posture of a first device and the posture of a second device, according to some examples.

[0027] [Figure 8A] FIG. 10 is a diagram illustrating a region of interest determined based on a forward head posture, according to some examples. [Figure 8B] FIG. 11 is a diagram illustrating a region of interest determined based on a forward head posture, according to some examples.

[0028] [Figure 9A] This diagram illustrates the determination of visibility metrics for areas of interest using several examples. [Figure 9B] This diagram illustrates the determination of visibility metrics for areas of interest using several examples.

[0029] [Figure 10A] This diagram illustrates, with several examples, devices that perform various actions in response to receiving natural language input, according to a determined visibility metric. [Figure 10B] This diagram illustrates, with several examples, devices that perform various actions in response to receiving natural language input, according to a determined visibility metric. [Figure 10C] This diagram illustrates, with several examples, devices that perform various actions in response to receiving natural language input, according to a determined visibility metric. [Figure 10D] This diagram illustrates, with several examples, devices that perform various actions in response to receiving natural language input, according to a determined visibility metric. [Figure 10E] This diagram illustrates, with several examples, devices that perform various actions in response to receiving natural language input, according to a determined visibility metric. [Figure 10F] This diagram illustrates, with several examples, devices that perform various actions in response to receiving natural language input, according to a determined visibility metric. [Figure 10G] This diagram illustrates, with several examples, devices that perform various actions in response to receiving natural language input, according to a determined visibility metric.

[0030] [Figure 11]This is a flowchart illustrating a method for providing audio output in response to natural language input, using several examples. [Modes for carrying out the invention]

[0031] Figures 1-4 illustrate examples of computer systems and techniques for interacting with a 3D scene. Figures 5A-5E illustrate image capture in a multi-device system. Figure 6 is a flowchart of a method for capturing an image in a multi-device system, using several examples. The method in Figure 6 is explained using Figures 5A-5E. Figure 7 illustrates a forward-facing head orientation determined based on the orientation of the first device and the orientation of the second device. Figures 8A-8B illustrate a region of interest determined based on the forward-facing head orientation. Figures 9A-9B illustrate the determination of the visibility metric for the region of interest. Figures 10A-10G illustrate devices performing various actions according to the determined visibility metric and in response to receiving natural language input. Figure 11 is a flowchart of a method for providing audio output in response to natural language input. The method in Figure 11 is explained using Figures 7, 8A-8B, 9A-9B, and 10A-10G.

[0032] Furthermore, in any method described herein that is conditional on one or more conditions being met by one or more steps, it should be understood that the method described can be repeated in multiple iterations such that all the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires that a first step be performed if a condition is met, and a second step be performed if the condition is not met, a person skilled in the art will understand that the steps described in the claim are repeated in an unspecified order until the conditions are met and then not met. Thus, a method described in one or more steps that depends on one or more conditions being met can be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this is not required for claims of a system or computer-readable medium that include instructions for performing a contingency operation based on the satisfaction of the corresponding one or more conditions, and therefore it is possible to determine whether the contingency is met without explicitly repeating the steps of the method until all the conditions that the steps of the method are conditional on are met. Those skilled in the art will understand that, as with methods involving incidental steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all incidental steps are performed.

[0033] Figure 1 is a block diagram illustrating the operating environment of a computer system 101 for interacting with a 3D scene, in several examples. In Figure 1, the user interacts with the 3D scene 105 through an operating environment 100 which includes the computer system 101. In some examples, the computer system 101 includes a controller 110 (e.g., a processor in a portable electronic device or remote server), user-responsive components 120, one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, and / or other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, an audio sensor, etc.), and one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some examples, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the user-responsive component 120 (for example, within a head-mounted device or handheld device).

[0034] While Figure 1 shows features suitable for operating environment 100, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more suitable embodiments of the examples disclosed herein.

[0035] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various three-dimensional scenes. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system to provide audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). Alternatively, a head-mounted system may be configured to operate without displaying content, for example, by providing output to the user via haptic and / or auditory means. The head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to the person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, holographic medium, optical coupler, optical reflector, or any combination thereof. In one example, the transparent or translucent display may be configured to be selectively opaque.A projection-based system may employ retinal projection technology to project graphical images onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or onto physical surfaces.

[0036] In some examples, the user-responsive component 120 is configured to provide visual components of a three-dimensional scene. In some examples, the user-responsive component 120 includes a preferred combination of software, firmware, and / or hardware. The user-responsive component 120 is described in more detail below with reference to Figure 2. In some examples, the functionality of the controller 110 is provided by and / or combined with the user-responsive component 120. In some examples, the user-responsive component 120 provides the user with an augmented reality (XR) experience while the user is virtually and / or physically present in the scene 105.

[0037] In some examples, the user-responsive component 120 is worn on a part of the user's body (e.g., the user's head or the user's hand). In some examples, the user-responsive component 120 includes one or more XR displays provided for displaying XR content. In some examples, the user-responsive component 120 surrounds the user's field of view. In some examples, the user-responsive component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some examples, the handheld device is optionally placed in an enclosure worn on the user's head. In some examples, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some examples, the user-responsive component 120 is an XR chamber, enclosure, or room configured to present XR content in which the user is not wearing or holding the user-responsive component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may also be implemented on another type of hardware for displaying XR content (e.g., a head-mounted device (HMD) or other wearable computing devices). For example, a user interface showing interaction with XR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may also be implemented similarly to an HMD, where the interaction occurs in the space in front of the HMD and the XR content response is displayed through the HMD. Similarly, a user interface showing interaction with XR content triggered based on the movement of a handheld or tripod-mounted device relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may also be implemented similarly to an HMD, where the movement is triggered by the movement of the HMD relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).

[0038] Figure 2 is a block diagram of user-responsive components 120 in several examples. While certain specific features are illustrated, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more appropriate embodiments of the examples disclosed herein. Furthermore, Figure 2 is more intended to illustrate the function of various features that may be present in a particular implementation, in contrast to the structural outlines of the examples described herein. As will be recognized by those skilled in the art, the components shown separately can be combined, and some components can be separated. For example, some functional modules shown separately in Figure 2 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various examples. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some examples, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0039] In some examples, the user-responsive component 120 (e.g., HMD) includes one or more processing units 202 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 206, one or more communication interfaces 208 (e.g., USB, FIREWIRE®, THUNDERBOLT®, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH®, ZIGBEE®, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, one or more XR displays 212, one or more optional in-facing and / or out-facing image sensors 214, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0040] In some examples, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some examples, one or more I / O devices and sensors 206 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more biosensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).

[0041] In some examples, one or more XR displays 212 are configured to provide an XR experience to the user. In some examples, one or more XR displays 212 correspond to holographic, digital light processing (DLP), liquid crystal displays (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistors (OLET), organic light-emitting diodes (OLED), surface conduction electron emission displays (SED), field emission displays (FED), quantum dot light-emitting diodes (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some examples, one or more XR displays 212 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, a user-responsive component 120 (e.g., HMD) includes a single XR display. In another example, the user-responsive component 120 includes an XR display for each of the user's eyes. In some examples, one or more XR displays 212 can present XR content. In some examples, one or more XR displays 212 are omitted from the user-responsive component 120. For example, user-responsive component 120 does not include any components configured to display content (or any components configured to display XR content), and user-responsive component 120 provides output via audio and / or haptic output types.

[0042] In some examples, one or more image sensors 214 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may also be referred to as eye-tracking cameras). In some examples, one or more image sensors 214 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally, at least a portion of the user's arms (and may also be referred to as hand-tracking cameras). In some examples, one or more image sensors 214 are configured to face forward to acquire image data corresponding to a scene that the user would view if a user-responsive component 120 (e.g., an HMD) were not present (and may also be referred to as a scene camera). One or more optional image sensors 214 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.

[0043] Memory 220 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access semiconductor memory devices. In some examples, memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile semiconductor storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes non-temporary computer-readable storage media. In some examples, memory 220, or the non-temporary computer-readable storage media of memory 220, including an optional operating system 230 and XR experience module 240, stores the following programs, modules, and data structures, or subsets thereof:

[0044] The operating system 230 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some examples, the XR experience module 240 is configured to present XR content to the user via one or more XR displays 212 or one or more speakers. For this purpose, in various examples, the XR experience module 240 includes a data acquisition unit 242, an XR presentation unit 244, an XR map generation unit 246, and a data transmission unit 248.

[0045] In some examples, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. For this purpose, in various examples, the data acquisition unit 242 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0046] In some examples, the XR presentation unit 244 is configured to present XR content via one or more XR displays 212 or one or more speakers. For this purpose, in various examples, the XR presentation unit 244 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0047] In some examples, the XR map generation unit 246 is configured to generate an XR map (e.g., a 3D map of an augmented reality scene or a map of a physical environment in which computer-generated objects can be placed) based on media content data. For this purpose, in various examples, the XR map generation unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0048] In some examples, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, sensor data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various examples, the data transmission unit 248 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0049] Although the data acquisition unit 242, XR presentation unit 244, XR map generation unit 246, and data transmission unit 248 are shown residing on a single device (e.g., user-responsive component 120 in Figure 1), in other examples any combination of the data acquisition unit 242, XR presentation unit 244, XR map generation unit 246, and data transmission unit 248 may reside in separate computing devices.

[0050] Returning to Figure 1, the controller 110 is configured to manage and adjust the user experience with respect to the 3D scene. In some examples, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with respect to Figure 3A.

[0051] In some examples, the controller 110 is a computing device that is local to or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some examples, the controller 110 is communicably coupled to components of the computer system 101 (e.g., output device 155 and / or user-responsive component 120) configured to provide output to the user via one or more wired or wireless communication channels (e.g., BLUETOOTH®, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In some examples, the controller 110 is contained within an enclosure (e.g., a physical housing) of a component of the computer system 101 (e.g., user-responsive component 120) configured to provide output to the user, or shares the same physical enclosure or support structure as a component of the computer system 101 configured to provide output to the user.

[0052] In some examples, the various components and functions of the controller 110 described below with respect to Figures 3A-3B, 4, 5A-5E, 6, 7, 8A-8B, 9A-9B, 10A-10G, and 11 are distributed across multiple devices. For example, a first set of components (and their associated functions) of the controller 110 is implemented on a remote server system relative to scene 105, while a second set of components (and their associated functions) is local to scene 105. For example, the second set of components is implemented within a portable electronic device (e.g., a wearable device such as an HMD) located within scene 105. It will be understood that the specific manner in which the various components and functions of the controller 110 are distributed across various devices may vary based on the various implementations of the examples described herein.

[0053] Figure 3A is a block diagram of controller 110 in several examples. While certain specific features are illustrated, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more appropriate embodiments of the examples disclosed herein. Furthermore, Figure 3A is more intended to illustrate the function of various features that may be present in a particular implementation, in contrast to the structural schematics of the examples described herein. As will be recognized by those skilled in the art, the components shown separately can be combined, and some components can be separated. For example, some functional modules shown separately in Figure 3A can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various examples. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some examples, will depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0054] In some examples, the controller 110 includes one or more processing units 302 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, and / or similar), one or more input / output (I / O) devices 306, and one or more communication interfaces 308 (e.g., Universal Serial Bus (USB), FireWire®, Thunderbolt®, IEEE 802.3x, IEEE 802.11x, IEEE 802.11x, etc.). This includes 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth®, ZIGBEE®, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0055] In some examples, one or more communication buses 304 include circuits that interconnect system components and control communication between system components. In some examples, one or more I / O devices 306 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and / or the like.

[0056] Memory 320 includes high-speed random access memory such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access semiconductor memory devices. In some examples, memory 320 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile semiconductor storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some examples, memory 320, or the non-temporary computer-readable storage medium of memory 320, including an optional operating system 330 and a three-dimensional (3D) experience module 340, stores the following programs, modules, and data structures, or subsets thereof:

[0057] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks.

[0058] In some examples, the 3D Experience Module 340 is configured to manage and adjust the user experience provided by the computer system 101 with respect to a 3D scene. For example, the 3D Experience Module 340 is configured to acquire data corresponding to the 3D scene (e.g., data generated by the computer system 101 and / or data from the data acquisition unit 341 described later) and to cause the computer system 101 to perform actions for the user based on the data (e.g., provide suggestions, display content, etc.). For this purpose, in various examples, the 3D Experience Module 340 includes a data acquisition unit 341, a tracking unit 342, an adjustment unit 346, a data transmission unit 348, a digital assistant (DA) unit 350, an image evaluation unit 370, an area tracking unit 380, and a visibility analysis unit 390.

[0059] In some examples, the data acquisition unit 341 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of the user-responsive components 120, input devices 125, output devices 155, sensors 190, and peripheral devices 195. For this purpose, in various embodiments, the data acquisition unit 341 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0060] In some examples, the tracking unit 342 is configured to map scene 105 and track the location of the user (and / or any portable devices held or worn by the user). For this purpose, in various examples, the tracking unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0061] In some examples, the tracking unit 342 includes an eye-tracking unit 343. The eye-tracking unit 343 includes instructions and / or logic for tracking the position and movement of the user's gaze (or more broadly, the user's eyes, face, or head) using data acquired from the eye-tracking device 130. In some examples, the eye-tracking unit 343 tracks the position and movement of the user's gaze relative to the physical environment, relative to the user (e.g., the user's hands, face, or head), relative to a device worn or held by the user, and / or relative to content displayed by the user-responsive component 120.

[0062] The eye-tracking device 130 is controlled by the eye-tracking unit 343 and includes various hardware and / or software components configured to perform eye-tracking techniques. For example, the eye-tracking device 130 includes at least one eye-tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye-tracking camera may be directed toward the user's eyes to receive IR or NIR light directly from the reflected light source, or alternatively, it may be directed toward a mirror that reflects IR or NIR light from the eyes toward the eye-tracking camera. The eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames / second), analyzes the images to generate eye-tracking information, and communicates the eye-tracking information to the eye-tracking unit 343. In some examples, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some cases, only one of the user's eyes is tracked by a separate eye-tracking camera and lighting source.

[0063] In some examples, the tracking unit 342 includes a hand tracking unit 344. The hand tracking unit 344 includes instructions and / or logic for tracking the position and / or movement of one or more parts of the user's hand using hand tracking data acquired from the hand tracking device 140. The hand tracking unit 344 tracks the position and / or movement relative to the scene 105, relative to the user (e.g., the user's head, face, or eyes), relative to a device worn or held by the user, relative to content displayed by the user-responsive component 120, and / or relative to a coordinate system defined relative to the user's hand. In some examples, the hand tracking unit 344 analyzes the hand tracking data to identify hand gestures (e.g., pointing gestures, pinch gestures, clenching gestures, and / or grasping gestures) and / or identify content corresponding to the hand gestures (e.g., physical or virtual content), such as content selected by the hand gestures. In some examples, the hand gestures are air gestures. Air gestures are gestures detected without the user touching (or independently of) an input element that is part of a device (e.g., computer system 101, one or more input devices 125, hand tracking device 140, device 500, device 1000, and / or device 1062), and are based on detected movements of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including the user's body movement relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), the user's body movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the user's other hand, and / or the movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movements of a part of the user's body (e.g., a tap gesture including the movement of the hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture including a predetermined speed or a rotation amount of the part of the user's body).

[0064] The hand tracking device 140 is controlled by the hand tracking unit 344 and includes various hardware and / or software components configured to perform hand tracking and hand gesture recognition techniques. For example, the hand tracking device 140 includes one or more image sensors (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that capture three-dimensional information (e.g., a depth map) representing the hand of a human user. One or more image sensors capture an image of the hand with sufficient resolution to distinguish the fingers and their respective positions. In some examples, one or more image sensors project a speckled pattern onto the environment including the hand and capture an image of the projected pattern. In some examples, one or more image sensors capture a temporal sequence of hand tracking data (e.g., captured three-dimensional information and / or captured images of the projected pattern), and the hand tracking device 140 communicates the temporal sequence of hand tracking data to the hand tracking unit 344 for further analysis, for example, to identify hand gestures, hand poses, and / or hand movements.

[0065] In some examples, the hand tracking device 140 includes one or more hardware input devices configured to be worn and / or held (or otherwise attached) by each of the user's one or more hands. In such examples, the hand tracking unit 344 tracks the position, orientation, and / or movement of the user's hand based on tracking the position, orientation, and / or movement of each hardware input device. The hand tracking unit 344 tracks the position, orientation, and / or movement of each hardware input device optically (e.g., via one or more image sensors) and / or based on data obtained from sensors contained within the hardware input device (e.g., accelerometers, magnetometers, gyroscopes, inertial measurement units, etc.). In some examples, the hardware input device includes one or more physical controls (e.g., buttons, touch-sensitive surfaces, pressure-sensitive surfaces, knobs, joysticks, etc.). In some examples, instead of, or in addition to, performing a particular function in response to the detection of a distinct type of hand gesture, the computer system 101 also performs that particular function in response to user input selecting a distinct physical control of the hardware input device. For example, the computer system 101 interprets a pinch hand gesture input as a selection of the focused element, and / or the selection of a physical button on a hardware device as a selection of the focused element.

[0066] In some examples, the adjustment unit 346 is configured to manage and adjust the experience provided to the user via the user-responsive components 120, one or more output devices 155, and / or one or more peripheral devices 195. For this purpose, in various examples, the adjustment unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0067] In some examples, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to a user-responsive component 120, one or more input devices 125, an output device 155, a sensor 190, and / or peripheral devices 195. For this purpose, in various examples, the data transmission unit 348 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0068] The Digital Assistant (DA) unit 350 includes instructions and / or logic for providing DA functionality to the computer system 101. Thus, the DA unit 350 provides DA functionality to the user of the computer system 101 while the user and / or their avatar are present in a three-dimensional scene. For example, the DA performs various tasks related to the three-dimensional scene, either in advance or upon request from the user. In some examples, the DA unit 350 performs at least some of the following: converting speech input to text (e.g., using a speech-to-text (STT) processing unit 352); identifying the user's intent expressed in natural language input received from the user; actively extracting and acquiring information necessary to fully satisfy the user's intent (e.g., by removing ambiguity in terminology in natural language input and / or by acquiring information from a data acquisition unit 341); determining a task flow to satisfy the identified intent; and executing that task flow to satisfy the identified intent.

[0069] In some examples, the DA unit 350 includes a natural language processing (NLP) unit 351 configured to identify user intents. The NLP unit 351 retrieves n best text representation candidates (singular or plural) ("word sequences (singular or plural)" or "token sequences (singular or plural)") generated by the STT processing unit 352 and attempts to associate each of these text representation candidates with one or more user intents recognized by the DA. In some examples, the user intent represents a task, which is executable by the DA and has an associated task flow implemented in the task flow processing unit 353. This associated task flow is a set of programmed actions and steps that the DA takes to execute that task. The scope of the DA's capabilities depends, in some examples, on the number and types of task flows implemented in the task flow processing unit 353, in other words, on the number and types of user intents recognized by the DA.

[0070] In some examples, once the NLP unit 351 identifies a user intent based on a user request, the NLP unit 351 causes the task flow processing unit 353 to perform the necessary actions to satisfy the user request. For example, the task flow processing unit 353 executes a task flow corresponding to the identified user intent in order to perform a task that satisfies the user request. In some examples, performing a task includes causing the computer system 101 to provide output (e.g., graphic output, audio output, and / or haptic output) indicating the performed task.

[0071] As shown in Figure 3B, the image evaluation unit 370 receives the input image 372 and the user request 374 and determines the prompt 376 and / or response 378. In some examples, the image evaluation unit 370 is included in the DA unit 350. In some examples, some or all of the functions of the image evaluation unit 370 discussed below are performed by and / or in conjunction with the DA unit 350.

[0072] The image evaluation unit 370 uses one or more image sensors of device 101, such as the image sensor 214, to acquire (e.g., receive and / or capture) an input image 372 (e.g., an image including an environment, an image including an object, an image including a person, and / or an image including any combination of an environment, an object, and / or a person), and uses one or more sensors of device 101, such as a microphone, a touch-sensitive display, and / or other sensors capable of receiving user speech input and / or text input, to receive user requests 374 related to the input image 372. In some examples, the user request 374 includes a task to be performed based on the content contained in the input image 372. In some examples, the user request 374 includes a request for data related to the content contained in the input image 372.

[0073] In some examples, one or more image sensors are part of the same computer system and / or device as the image evaluation unit 370 (e.g., at least partially inside the computer system and / or directly connected to the computer system). In some examples, one or more image sensors are part of another computer system. In some examples, at least one image sensor is part of the computer system that includes the image evaluation unit 370. In some examples, at least one image sensor is part of another computer system. In some examples, at least one image sensor is a forward-facing camera of the computer system that includes the image evaluation unit 370 (e.g., the camera faces the front of the computer system). In some examples, at least one image sensor is a rear-facing camera of the computer system that includes the image evaluation unit 370 (e.g., the camera faces the rear of the computer system).

[0074] In some examples, the image sensor of the computer system and / or device including the image evaluation unit 370 is of lower quality than the image sensor of another computer system and / or device connected to and / or communicating with the computer system and / or device including the image evaluation unit 370. In some examples, the image sensor of the computer system and / or device including the image evaluation unit 370 has at least one characteristic (e.g., resolution, associated focal length, magnification, aperture, dynamic range, etc.) that is different from the image sensor of another computer system and / or device connected to and / or communicating with the computer system and / or device including the image evaluation unit 370.

[0075] In that case, the image evaluation unit 370 determines the quality of the input image 372 and determines whether the quality of the input image 372 meets the quality criteria (e.g., matches) or does not meet the quality criteria (e.g., does not match).

[0076] In some examples, the quality of the input image 372 is based on factors including blur, sharpness, clarity, noise, exposure, tone, contrast, distortion, vignetting, artifacts, and / or lens flare present in the input image 372. In some examples, the image evaluation unit 370 determines the quality of the input image 372 by processing the image to determine whether and to what extent the above factors are present. In some examples, the image evaluation unit 370 includes and / or uses one or more AI models to determine the quality of the input image 372.

[0077] In some cases, the quality criteria are based on user request 374 and / or the tasks included in user request 374 (for example, some tasks require higher quality images than others). In some cases, the image evaluation unit 370 selects a first quality criterion as the quality criterion based on the determination that user request 374 includes a first type of request, and the image evaluation unit 370 selects a second quality criterion as the quality criterion based on the determination that user request 374 includes a second type of request different from the first type. For example, if the image evaluation unit 370 determines that the task of user request 374 requires a large amount of information from the input image 372, the image evaluation unit 370 will select a quality criterion of relatively high image quality (e.g., not blurry, sharp, clear, without much noise, not distorted, etc.). However, if the image evaluation unit 370 determines that the task of user request 374 requires only a small amount of information from the input image 372, the image evaluation unit 370 will select a quality criterion of relatively low image quality (e.g., possibly somewhat blurry, not necessarily perfectly clear, possibly containing noise and / or some distortion, etc.).

[0078] In some examples, the image evaluation unit 370 provides the input image 372 to a large-scale language model (LLM) or other AI model and requests the LLM or other AI model to determine the quality of the input image 372. In some examples, the image evaluation unit 370 provides the input image 372 to a large-scale language model (LLM) or other AI model and requests the LLM or other AI model to determine whether the input image 372 is of sufficient quality to complete the task determined from the user request 374.

[0079] In some examples, the image evaluation unit 370 generates an embedding of the input image 372 and determines the quality of the input image 372 by comparing the embedding of the input image 372 with a set of trained embeddings that represent high-quality or low-quality images. In some examples, the image evaluation unit 370 determines the quality of the input image 372 by generating an embedding of the input image 372 and providing the embedding of the input image 372 to an LLM or other AI model. In this case, the image evaluation unit 370 requests the LLM or other AI model to determine the quality of the input image 372 by comparing the provided embedding with other embeddings of images of various qualities. In some examples, the image evaluation unit 370 selects a set of trained embeddings based on the type of request contained in the user request 374. For example, if the image evaluation unit 370 determines that the task of user request 374 requires a large amount of information from the input image 372, the image evaluation unit 370 selects a set of embeddings that represent high-quality images. However, if the image evaluation unit 370 determines that the task of user request 374 requires only a small amount of information from the input image 372, the image evaluation unit 370 selects a set of embeddings that represent low-quality images.

[0080] In some examples, the image evaluation unit 370 provides an input image 372 to an artificial intelligence (AI) model and / or other model in order to perform a user request 374. When the confidence level of the outcome of performing the user request 374, as determined by the model, is sufficiently high (e.g., the criteria for performing the task are met), the input image 372 is of sufficient quality to perform the task. When the confidence level of the outcome of performing the user request 374, as determined by the model, is not sufficiently high (e.g., the criteria for performing the task are not met), the input image 372 is not of sufficient quality to perform the task. In some examples, the confidence level of the outcome of performing the user request 374 is provided to the image evaluation unit 370, which uses the confidence level of the outcome to determine whether another photograph should be taken and / or whether the camera of another device should be opened (e.g., launched, activated, called, etc.).

[0081] When the image evaluation unit 370 determines that the quality of the input image 372 meets the quality criteria, the image evaluation unit 370 generates a response 378 to the user request 374 based on the input image 372 (for example, by utilizing the capabilities of the DA unit 350 to determine the user intent and perform one or more actions to satisfy the user request 374), and provides an output including the response 378. In some examples, the image evaluation unit 370 determines that the quality of the input image 372 meets the quality criteria when the quality of the first image is high enough for the computer system and / or digital assistant to determine the response to the user request.

[0082] In some examples, the response 378 to user request 374 includes an output indicating that the task is complete, a response to a request for information, and / or a follow-up prompt for further information related to user request 374. In some examples, the output of response 378 is an audio output and / or an output on a display communicating with the computer system.

[0083] When the image evaluation unit 370 determines that the quality of the input image 372 does not meet the quality standards, the image evaluation unit 370 generates a prompt 376 for capturing a second input image and causes another computer system and / or electronic device, whether or not it is physically connected to device 101, to provide the prompt 376 for capturing the second input image using the sensors of that other computer system and / or electronic device. In some examples, the image evaluation unit 370 determines that the quality of the input image 372 does not meet the quality standards when it determines that the quality of the input image 372 is too low for the computer system and / or digital assistant to decide how to respond to the user request 374.

[0084] In some examples, prompt 376 includes an output requesting that another image be captured by an image sensor (e.g., a camera) of another computer system. In some examples, prompt 376 is provided as an output by device 101 (e.g., the same computer system that includes the image evaluation unit 370). In some examples, prompt 376 is provided as an output by another computer system. In some examples, prompt 376 is provided as an audio output. In some examples, prompt 376 is provided as a visual output. In some examples, prompt 376 is provided by a digital assistant associated with both computer systems. In some examples, the output is provided within a user interface associated with that digital assistant. In some examples, the output is provided within a user interface of a camera application. In some examples, the two devices and / or computer systems are communicating with each other. In some examples, the two devices and / or computer systems are connected wirelessly via, for example, Wi-Fi, Bluetooth®, NFC, and / or other wireless communication protocols. In some examples, the two devices and / or computer systems are both associated with the same user and / or the same profile of that user. In some examples, the two devices are connected via wires and / or other physical connections.

[0085] In some examples, after providing prompt 376, and / or after another device and / or computer system has provided prompt 376, user input for capturing another input image is detected, and in response to the detection of user input, the other input image is acquired (e.g., received and / or captured). The image evaluation unit 370 then determines a response to user request 374 based on the other input image acquired using the other device and / or computer system, and provides an output including the response to user request 374. Thus, the user receives a response to user request 374 based on the information available to both devices and / or computer systems. In some examples, after detecting user input for capturing another input image, the display of prompt 376 and / or another user interface is discontinued.

[0086] In some examples, in response to detecting the input image 372 and the user request 374, the image evaluation unit 370 determines whether the context of device 101 (e.g., the device and / or computer system receiving, capturing, and / or acquiring the input image 372) indicates that the input image 372 does not meet quality criteria. The context of device 101 (e.g., the device and / or computer system receiving, capturing, and / or acquiring the input image 372) is determined using data received from one or more sensors of the device, which may include information representing the level of light around the device, the location of the device, the presence of an object in front of the device's image sensor, the movement of the device, the presence of text in front of the device, and / or other information related to the quality of the input image 372. For example, data from the device's sensors may indicate that the device is in a dark room or that there is an object in front of the device's camera, and therefore the input image 372 is too dark and / or out of focus to extract information.

[0087] In some cases, when the image evaluation unit 370 determines that the device context indicates that the input image 372 does not meet the quality criteria, the image evaluation unit 370 ceases determining whether the input image 372 meets the quality criteria and provides a prompt 376 to another computer system and / or the device that is communicating with the device.

[0088] In some examples, the 3D EXPERIENCE module 340 accesses one or more artificial intelligence (AI) models configured to perform various functions described herein. The AI ​​models are implemented at least in part on the controller 110 (for example, locally on a single device or in a distributed manner), and / or the controller 110 communicates with one or more external services that provide access to the AI ​​models. In some examples, one or more components and functions of the DA unit 350, the image evaluation unit 370, the area tracking unit 380, and / or the visibility analysis unit 390 are implemented using the AI ​​models. For example, the DA unit 350 implements one or more AI models to perform speech recognition, intent determination (e.g., natural language processing and / or image processing), object recognition, and / or response generation; the image evaluation unit 370 implements one or more AI models to determine whether an input image meets quality criteria for determining a response to a user request, and / or to determine prompts for other images on another device and / or computer system; and / or the visibility analysis unit 390 implements one or more AI models to determine (e.g., identify) occluded portions of captured image data.

[0089] In some cases, an AI model is based on (e.g., is or is built upon) one or more foundational models. Generally, a foundational model is a deep learning neural network that is trained on a large training dataset and can be adapted to perform specific functions. Thus, a foundational model can aggregate information learned from large (and optionally, multimodal) datasets and be adapted (e.g., fine-tuned) to perform a variety of downstream tasks that the foundational model may not have been originally designed to perform. Examples of such tasks include language translation, speech recognition, user intent determination (e.g., natural language processing), sentiment analysis, computer vision tasks (e.g., object recognition and scene understanding), question answering, image generation, audio generation, and the generation of computer executable instructions. A foundational model can accept a single type of input (e.g., text data) or multimodal input such as two or more of text data, image data, video data, audio data, sensor data, etc. In some cases, a foundational model is prompted to perform a particular task by providing a natural language description of that task. Examples of foundational models include Open AI, Inc.'s GPT-n series models (e.g., GPT-1, GPT-2, GPT-3®, and GPT-4®), DALL-E, and CLIP, Microsoft Corporation's Florence and Florence-2, Google LLC's BERT, and Meta Platforms, Inc.'s LLaMA, LLaMA-2, and LLaMA-3.

[0090] Figure 4 illustrates Architecture 400 for a foundational model with several examples. Architecture 400 is merely illustrative, and various modifications to it are possible. Thus, the components of Architecture 400 (and the functions associated with them) can be combined, the order of the components (and the functions associated with them) can be changed, components of Architecture 400 can be removed, and other components can be added to Architecture 400. Furthermore, although Architecture 400 is based on transformers, those skilled in the art will understand that Architecture 400 can additionally or alternatively implement other types of machine learning models, such as models based on convolutional neural networks (CNNs) and models based on recurrent neural networks (RNNs).

[0091] Architecture 400 is configured to process input data 402 and generate output data 480 corresponding to a desired task. Input data 402 includes one or more types of data, such as text data, image data, video data, audio data, sensor data (e.g., motion sensors, biosensors, temperature sensors, etc.), computer executable instructions, and structured data (e.g., in the form of XML files, JSON files, or other file types). In some examples, input data 402 includes data from data acquisition unit 341. Output data 480 includes one or more types of data, depending on the task being performed. For example, output data 480 includes one or more of text data, image data, audio data, and computer executable instructions. The input and output data types described above are merely examples, and it will be understood that Architecture 400 can be configured to accept various types of data as input and generate various types of data as output. Such data types can be diverse based on the specific functions configured for the underlying model to perform.

[0092] Architecture 400 includes an embedded module 404, an encoder 408, an embedded module 428, a decoder 424, and an output module 450, the functions of which are described below.

[0093] The embedding module 404 is configured to receive input data 402 and parse the input data 402 into one or more token sequences. The embedding module 404 is further configured to determine the embedding (e.g., vector representation) of each token representing each token in the embedding space, such that, for example, similar tokens are close together in the embedding space and dissimilar tokens are far apart. In some examples, the embedding module 404 includes a position encoder configured to encode position information within the embedding. The individual position information of a given embedding indicates its relative position within the sequence. The embedding module 404 is configured to output embedding data 406 of the input data by aggregating the token embeddings of the input data 402.

[0094] The encoder 408 is configured to map the embedded data 406 to an encoder representation 410. The encoder representation 410 represents contextual information for each token, showing learned information about how each token relates to (e.g., attends) each other. The encoder 408 includes an attention layer 412, a feedforward layer 416, normalization layers 414 and 418, and residual connections 420 and 422. In some examples, the attention layer 412 applies a self-attention mechanism to the embedded data 406 to compute an attention representation (e.g., in matrix form) of the relationships of each token in the sequence to each other. In some examples, the attention layer 412 is multi-headed to compute multiple different attention representations of the relationships of each token to each other, each different representation showing a different learned property of the token sequence. The attention layer 412 is configured to aggregate the attention representations to output attention data 460 showing the interrelationships between tokens from the input data 402. In some examples, the attention layer 412 further masks the attention data 460 to suppress data representing relationships between selected tokens. The encoder 408 then passes the attention data 460 (optionally masked) through the normalization layer 414, the feedforward layer 416, and the normalization layer 418 to generate the encoder representation 410. The residual connections 420 and 422 can help stabilize and shorten the training and / or inference process by allowing the output of the embedding module 404 (i.e., the embedding data 406) to be passed directly to the normalization layer 414, and the output of the normalization layer 414 to be passed directly to the normalization layer 418, respectively.

[0095] Figure 4 illustrates that architecture 400 includes a single encoder 408, but in other examples, architecture 400 includes multiple stacked encoders configured to output an encoder representation 410. Each of the stacked encoders can generate different attention data, which may enable architecture 400 to learn different types of interrelationships between tokens and generate output data 410 based on a more complete set of learned relationships.

[0096] Decoder 424 is configured to receive encoder representation 410 and previous output embedding 430 as input and generate output data 480. Embedding module 428 is configured to generate previous output embedding 430. Embedding module 428 is similar to embedding module 404. Specifically, embedding module 428 tokenizes previous output data 426 (e.g., output data 480 generated by a previous iteration), determines the embedding for each token, and optionally encodes positional information into each embedding to generate previous output embedding 430.

[0097] Decoder 424 includes attention layers 432 and 436, normalization layers 434, 438, and 442, a feedforward layer 440, and residual connections 462, 464, and 466. Attention layer 432 is configured to output attention data 470 indicating the interrelationships between tokens from the previous output data 426. Attention layer 432 is similar to attention layer 412. For example, attention layer 432 applies a multi-head self-attention mechanism to the previous output embedding 430 and optionally masks attention data 470 to suppress data representing relationships between selected tokens (e.g., relationships between one token and a future token) so that architecture 400 does not consider future tokens as context when generating output data 480. Decoder 424 then passes attention data 470 (optionally masked) to normalization layer 434 to generate normalized attention data 470-1.

[0098] The attention layer 436 receives the encoder representation 410 and normalized attention data 470-1 as input to generate encoder-decoder attention data 475. The encoder-decoder attention data 475 correlates the input data 402 to the previous output data 426 by representing the relationship between the output of the encoder 408 and the previous output of the decoder 424. The attention layer 436 allows the decoder 424 to increase the weights of the parts of the encoder representation 410 that have been learned as relatively reasonable for generating the output data 480. In some examples, the attention layer 436 applies a multi-head attention mechanism to the encoder representation 410 and normalized attention data 470-1 to generate the encoder-decoder attention data 475. In some examples, the attention layer 436 further masks the encoder-decoder attention data 475 to suppress the interrelationships between selected tokens.

[0099] Decoder 424 then passes encoder-decoder attention data 475 (optionally masked) through normalization layer 438, feedforward layer 440, and normalization layer 442 to generate further processed encoder-decoder attention data 475-1. Normalization layer 442 then provides the further processed encoder-decoder attention data 475-1 to output module 450. Similar to residual connections 420 and 422, residual connections 462, 464, and 466 can stabilize and shorten the training and / or inference process by allowing the output of the corresponding component to be passed directly as input to the corresponding component.

[0100] Figure 4 illustrates that architecture 400 includes a single decoder 424, but in other examples, architecture 400 includes multiple stacked decoders, each configured to learn / generate different types of encoder-decoder attention data 475. This allows architecture 400 to learn multiple different types of interrelationships between tokens from input data 402 and tokens from output data 480, thereby allowing architecture 400 to generate output data 480 based on a more complete set of learned relationships.

[0101] The output module 450 is configured to generate output data 480 from further processed encoder-decoder attention data 475-1. For example, the output module 450 includes one or more linear layers that apply a learned linear transformation to the further processed encoder-decoder attention data 475-1, and a softmax layer that generates a probability distribution over possible classes (e.g., words or symbols) of output tokens based on the linear transformation data. The output module 450 then selects (e.g., predicts) elements of the output data 480 based on the probability distribution. The architecture 400 then passes the output data 480 as the previous input data 426 to the embedding module 428 to start another iteration of the training and / or inference process for the architecture 400.

[0102] It will be understood that various different AI models can be built based on the components of Architecture 400. For example, some large-scale language models (LLMs) (e.g., GPT-2 and GPT-3®) are decoder-only (e.g., include one or more instances of decoder 424 and do not include encoder 408), some LLMs (e.g., BERT) are encoder-only (e.g., include one or more instances of encoder 408 and do not include decoder 424), and other foundational models (e.g., Florence-2) are encoder-decoder type (e.g., include one or more instances of encoder 408 and one or more instances of decoder 424). Furthermore, it will be understood that foundational models built based on the components of Architecture 400 can be fine-tuned based on reinforcement learning techniques and training data specific to those tasks in order to optimize specific tasks such as extracting meaningful information from image and / or video data, generating code, generating music, or providing meaningful suggestions to a particular user.

[0103] Figures 5A to 5E illustrate image capture in a multi-device system using several examples.

[0104] Devices 500 and 550 implement at least some of the components of the computer system 101. For example, devices 500 and 550 include one or more sensors configured to detect data (e.g., image data and / or audio data) corresponding to each scene. In some examples, devices 500 and / or device 550 are HMDs (e.g., XR headsets or smart glasses), and Figures 5A to 5E illustrate the user's view of each scene through the HMD. For example, Figures 5A to 5E illustrate a physical scene viewed via pass-through video, a physical scene viewed via direct optical see-through, or a virtual scene viewed through one or more displays of the HMD. In other examples, devices 500 and / or device 550 are other types of devices such as smartwatches, smartphones, tablet devices, laptop computers, glasses without displays, headphones, earphones, or projection-based devices.

[0105] The examples in Figures 5A to 5E illustrate the presence of the user and devices 500 and 550 within their respective scenes. For example, the scene is a physical or augmented reality scene, and the user and devices 500 and 550 are physically present within the scene. In other examples, the user's avatar is present within the scene. For example, if the scene is a virtual reality scene, the user's avatar is present within the virtual reality scene.

[0106] Figures 5A to 5E include devices 500 and 550, both capable of acquiring image data (e.g., detection and / or capture) and receiving user input, including user requests. In some examples, devices 500 and 550 communicate but are not physically connected. In some examples, devices 500 and 550 are wirelessly connected (e.g., via Bluetooth®, Wi-Fi, NFC, and / or other wireless communication protocols). In some examples, devices 500 and 550 are connected via wires and / or other physical connections. In some examples, devices 500 and 550 are associated with the same user and / or the same user profile. In some examples, devices 500 and 550 are located close to each other but are not physically connected.

[0107] In Figure 5A, device 500 (e.g., a head-mounted device, smartphone, tablet, wearable computer system, smart device, and / or shared computer system) acquires (e.g., captures and / or receives) image 502a using one or more image sensors communicating with device 500 (e.g., the camera of device 500, the camera of a device connected to device 500, and / or the camera of a device communicating with device 500). In some examples, as shown in Figure 5A, image 502a is displayed on and / or communicates with the display and / or display generating component of device 500. In some examples, device 500 does not have a display, image 502a is not displayed, and / or the scene is directly viewed by the user.

[0108] Device 500 also receives user requests 504a related to image 502a (e.g., detect, acquire, and / or capture). In some examples, user requests 504a are detected before image 502a is acquired. In some examples, user requests 504a are detected after image 502a is acquired. In some examples, user requests 504a are detected at the same time as, or substantially at the same time as, image 502a is acquired.

[0109] In response to acquiring image 502a and receiving user request 504a, device 500 uses the image evaluation unit 370 to determine whether the quality of image 502a meets the quality criteria, as discussed above with reference to Figure 3B. Device 500 determines that the quality of image 502a meets the quality criteria and therefore determines response 506a and provides response 506a as audio output. Specifically, based on user request 504a, "What is that?", device 500 determines that the user is trying to identify an object in image 502a. Device 500 (for example, using the image evaluation unit 370) determines that the quality of image 502a is high enough to identify the object in image 502a and processes image 502a accordingly to determine that image 102a contains a tree, and further determines that image 102a contains an oak tree. Device 500 then generates the response, "That is an oak tree," and provides this response as an audio output to respond to user request 504a.

[0110] Since device 500 determines that image 502a is of high quality and therefore meets the quality criteria, device 500 does not determine a prompt to capture another image, nor does it prompt device 550 to provide a prompt or perform any other task. Therefore, as shown in Figure 5A, device 550 does not change or alter its display while determining response 506a.

[0111] In some examples, device 500 provides image 502a to device 550 and / or another device to determine, using the image evaluation unit 370, whether the quality of image 502a meets the quality criteria, as discussed above with reference to Figure 3B. Device 550 determines that the quality of image 502a meets the quality criteria and therefore determines response 506a, causing device 500 to provide response 506a as an audio output. Specifically, based on the user request 504a, "What is that?", device 500 provides the user request 504a to device 550, and device 550 determines that the user is trying to identify an object in image 502a. Device 550 determines that the quality of image 502a is high enough (for example, using the image evaluation unit 370) to identify the object in image 502a, and processes image 502a accordingly, determining that image 102a contains a tree, and further, that image 102a contains an oak tree. Device 550 then generates the response, "That is an oak tree," and in order to respond to user request 504a, device 550 provides that response as an audio output.

[0112] In Figure 5B, device 500 acquires image 502b using one or more image sensors communicating with device 500. In some examples, image 502b is displayed on and / or communicated with the display and / or display generation components of device 500. In some examples, image 502b is not displayed, device 500 does not include a display, and / or the user views the scene directly.

[0113] Device 500 also receives user requests 504b related to image 502b (e.g., detect, acquire, and / or capture). In some examples, user requests 504b are detected before image 502b is acquired. In some examples, user requests 504b are detected after image 502b is acquired. In some examples, user requests 504b are detected at the same time as, or substantially at the same time as, image 502b is acquired.

[0114] In response to acquiring image 502b and receiving user request 504b, device 500 uses the image evaluation unit 370 to determine whether the quality of image 502b meets the quality criteria, as discussed above with reference to Figure 3B. Device 500 determines that the quality of image 502b does not meet the quality criteria and therefore determines prompt 508b. After determining (e.g., generating) prompt 508b, device 500 causes device 550 to provide prompt 508b. Specifically, since device 550 has a higher quality camera and is therefore more likely to capture a higher quality image that will contain information for determining the response to user request 504b, device 500 causes device 550 to display prompt 508b to capture another image.

[0115] In some cases, as discussed above with reference to Figure 3B, device 500 determines that the quality of image 502b does not meet the quality criteria because image 502b is blurry, unclear, contains artifacts, is obscured, contains noise, is distorted, and / or has other factors that degrade the quality of image 502b.

[0116] After prompting device 550 to display prompt 508b, device 550 and / or device 500 detect user input 510b on the "Yes" button of prompt 508b. In response to detecting input 510b, device 550 acquires (e.g., captures and / or receives) image 502c, as shown in Figure 5C. After acquiring image 502c, device 550 provides device 500 with image 502c and / or data representing image 502c so that device 500 can determine response 506c to user request 504b. Device 500 then determines response 506c and provides response 506c as an audio output.

[0117] In some examples, device 550 provides response 506c as an audio output instead of device 500. In some examples, device 500 and / or device 550 display response 506c on the displays of device 500 and / or device 550, and / or on display generating components communicating with device 500 and / or device 550. In some examples, as shown in Figure 5C, after detecting user input 510b, device 500 stops displaying image 502b.

[0118] In some examples, device 550 provides response 506c as an audio output instead of device 500. In some examples, device 500 and / or device 550 display response 506c on the displays of device 500 and / or device 550, and / or on display generating components communicating with device 500 and / or device 550. In some examples, as shown in Figure 5C, after detecting user input 510b, device 500 stops displaying image 502b.

[0119] In some examples, in response to acquiring image 502b and receiving user request 504b, device 500 provides image 502b and user request 504b to device 550, causing device 550 to use the image evaluation unit 370 to determine whether the quality of image 502b meets the quality criteria, as discussed above with reference to Figure 3B. Device 550 determines that the quality of image 502b does not meet the quality criteria and therefore determines prompt 508b. After determining (e.g., generating) prompt 508b, device 550 provides prompt 508b. Specifically, since device 550 has a higher quality camera and is therefore more likely to capture a higher quality image, which will contain information for determining the response to user request 504b, prompt 508b includes a prompt to capture another image.

[0120] In some cases, as discussed above with reference to Figure 3B, the device 550 determines that the quality of image 502b does not meet the quality criteria because image 502b is blurry, unclear, contains artifacts, is obscured, contains noise, is distorted, and / or has other factors that degrade the quality of image 502b.

[0121] After device 550 displays prompt 508b, device 550 and / or device 500 detect user input 510b on the "Yes" button of prompt 508b. In response to detecting input 510b, device 550 acquires (e.g., captures and / or receives) image 502c, as shown in Figure 5C. After acquiring image 502c, device 550 determines response 506c to user request 504b. Device 550 then determines response 506c and provides response 506c as an audio output and / or provides response 506c to device 500 to be provided as an audio output.

[0122] In Figure 5D, device 500 acquires image 502d using one or more image sensors communicating with device 500. In some examples, image 502d is displayed on and / or communicated with the display and / or display generation components of device 500. In some examples, device 500 does not have a display, image 502d is not displayed, and / or the scene is directly viewed by the user of device 500.

[0123] Device 500 also receives user requests 504d related to image 502d (e.g., detect, acquire, and / or capture). In some examples, user requests 504d are detected before acquiring image 502d. In some examples, user requests 504d are detected after acquiring image 502d. In some examples, user requests 504d are detected at the same time as, or substantially at the same time as, acquiring image 502d.

[0124] In response to acquiring image 502d and receiving user request 504d, device 500 uses the image evaluation unit 370 to determine that the context of device 500 indicates that image 502d does not meet the quality criteria, as discussed above with reference to Figure 3B. Specifically, device 500 determines that it is moving while capturing image 502d, based on data received from one or more of its sensors. Therefore, device 500 ceases determining whether the quality of image 502d meets the quality criteria and determines prompt 508d to open the camera application user interface. After determining (e.g., generating) prompt 508d, device 500 causes device 550 to provide prompt 508d by causing the camera application of device 550 to open and display the camera application user interface. Specifically, since device 550 has a higher quality camera and is therefore more likely to capture a higher quality image that will contain information for determining the response to user request 504d, device 500 causes device 550 to open a camera application (e.g., display prompt 508d) to capture another image.

[0125] After prompting device 550 to display prompt 508d, device 550 and / or device 500 detect user input 510d on the capture button of the camera user interface. In response to detecting input 510d, device 550 acquires (e.g., captures and / or receives) an image and provides the image and / or data representing the image to device 500 so that device 500 can determine a response 506e to user request 504d. Device 500 then determines a response 506e and provides the response 506e as an audio output, as shown in Figure 5E.

[0126] In some examples, device 550 provides response 506e as an audio output instead of device 500. In some examples, device 500 and / or device 550 display response 506e on the displays of device 500 and / or device 550, and / or on display generating components communicating with device 500 and / or device 550. In some examples, after detecting user input 510d, device 500 ceases to display image 502d, as shown in Figure 5E. In some examples, after detecting user input 510d, device 550 ceases to display prompt 508d and / or the camera user interface, and instead displays the lock screen, as shown in Figure 5E. Thus, in some examples, device 550 does not display the captured image in response to detecting user input 510d, but instead provides the captured image and / or data corresponding to the image without displaying the image.

[0127] In some examples, in response to acquiring image 502d and receiving user request 504d, device 500 provides image 502d and user request 504d to device 550, which uses the image evaluation unit 370 to determine that the context of device 500 indicates that image 502d does not meet the quality criteria, as discussed above with reference to Figure 3B. Specifically, device 550 determines that device 500 is moving while device 500 is capturing image 502d, based on data received from one or more sensors of device 500 and / or device 550. Therefore, device 550 ceases determining whether the quality of image 502d meets the quality criteria and determines prompt 508d to open the camera application user interface. After determining (e.g., generating) prompt 508d, device 550 provides prompt 508d by opening (e.g., launching, activating, and / or calling) the camera application of device 550, opening and displaying the user interface of the camera application. Specifically, if device 550 has a higher quality camera and is therefore more likely to capture a higher quality image that will contain information for determining the response to user request 504d, device 550 opens the camera application (e.g., display prompt 508d) to capture a different image.

[0128] After device 550 displays prompt 508d, device 550 and / or device 500 detect user input 510d on the capture button of the camera user interface. In response to detecting input 510d, device 550 acquires (e.g., captures and / or receives) an image and determines a response 504e to user request 506d. Device 550 then provides the response 506e as an audio output.

[0129] The above example is discussed in terms of device 500 receiving a first image and determining whether the quality of the first image meets quality criteria, but it will be understood that device 550 can also receive a first image and determine whether the quality of the first image meets quality criteria. Similarly, one device, such as device 500, can capture an image, and another device, such as device 550, can determine whether the quality of the image meets quality criteria. Thus, the steps of capturing an image, determining whether the quality meets the criteria, and / or causing a prompt to appear on another device can be performed by any of the devices in the system. Similarly, the above example discusses two devices in the system, but the system can include three, four, five, or any other number of devices connected and / or communicating wirelessly to exchange data including the captured image and the determination of whether the image quality meets quality criteria.

[0130] Further explanations regarding Figures 5A to 5E are provided below with reference to Method 600 described below with respect to Figure 6.

[0131] Figure 6 is a flowchart of method 600 for capturing images in a multi-device system. In some examples, method 600 is performed in a first computer system (e.g., computer system 101, device 500, and / or device 550 in Figure 1) communicating with one or more image sensors (e.g., image sensors, light sensors, and / or photosensors). In some examples, method 600 is stored in a non-temporary (or temporary) computer-readable storage medium and managed by instructions executed by one or more processors of the computer system, such as one or more processing units 302 of computer system 101 (e.g., controller 110 in Figure 1). In some examples, the operation of method 600 is distributed across multiple computer systems, such as a computer system and a separate server system. Some operations of method 600 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0132] In block 602, a first image (e.g., 372, 502a, 502b, and / or 502d) is acquired using one or more image sensors.

[0133] In block 604, user requests related to the first image (e.g., 374, 504a, 504b, and / or 504d) are received.

[0134] In block 608, in response to the acquisition of the first image and the receipt of a user request (606), a second computer system (e.g., computer system 101, device 500, and / or device 550 in Figure 1) is provided with prompts (e.g., 376, 508b, and / or 508d) to capture a second image (e.g., 502, 502a, 502b, and / or 372d) according to a determination (e.g., by the image evaluation unit 370) that the quality of the first image does not meet the quality criteria.

[0135] In block 612, in response to acquiring the first image and receiving a user request (606), a response to the user request (e.g., 378, 506a, 506c, and / or 506e) is generated based on the first image, according to a determination (e.g., by the image evaluation unit 370) that the quality of the first image meets the quality criteria (610).

[0136] In block 614, according to the determination that the quality of the first image meets the quality criteria (610), output is provided that includes a response to a user request based on the first image.

[0137] In some examples, the first computer system is a head-mounted electronic device, and the second computer system is a smartphone. In some examples, the image sensor of the first computer system is a lower-quality image sensor than the image sensor of the second computer system.

[0138] In some examples, method 600 further includes, in response to detecting a first image and receiving a user request, selecting a first quality criterion as a quality criterion according to a determination that the user request includes a request of a first type, and selecting a second quality criterion different from the first quality criterion as a quality criterion according to a determination that the user request includes a request of a second type different from the first type.

[0139] In some examples, method 600 further includes causing a second computer system to provide a prompt to capture a second image using the second computer system, detecting user input to capture a second image using the second computer system, generating a response to the user request based on the second image, and providing an output including the response to the user request based on the second image.

[0140] In some examples, determining whether the quality of a first image meets a quality criterion involves providing a Large Language Model (LLM) with a prompt that includes a requirement to determine whether the first image is of sufficient quality to complete a task determined from a user request.

[0141] In some examples, determining whether the quality of a first image meets a quality criterion involves generating an embedding for the first image and comparing the embedding for the first image to a set of trained embeddings that represent high-quality or low-quality images.

[0142] In some examples, method 600 further includes selecting a first set of trained embeddings as the set of trained embeddings according to the determination that the user request is a request of a first type, and selecting a second set of trained embeddings as the set of trained embeddings according to the determination that the user request is a request of a second type.

[0143] In some examples, method 600 further includes, in response to detecting a first image and receiving a user request, ceasing to determine whether the first image meets quality criteria based on the determination that the context of the first computer system indicates that the first image does not meet quality criteria, and prompting the second computer system to capture a second image using the second computer system.

[0144] In some examples, the determination that the context of the first computer system indicates that the first image does not meet quality criteria includes the determination that the first computer system is moving. In some examples, the determination that the context of the first computer system indicates that the first image does not meet quality criteria includes the determination that the illumination level of the environment of the first computer system is below the illumination threshold. In some examples, the determination that the context of the first computer system indicates that the first image does not meet quality criteria includes the determination that one or more image sensors are obscured. In some examples, the determination that the context of the first computer system indicates that the first image does not meet quality criteria includes the determination that the field of view of one or more image sensors contains text.

[0145] In some examples, method 600 further includes, in accordance with the determination that the quality of the first image does not meet quality standards, causing the camera user interface to be displayed using a display generation component that communicates with a second computer system.

[0146] In some examples, method 600 further includes making a camera user interface visible using a display generation component communicating with a second computer system, detecting user input for capturing a second image, and, in response to detecting user input for capturing a second image, ceasing to display the camera user interface using the display generation component communicating with the second computer system.

[0147] In some examples, the camera user interface is displayed on the lock screen using a display generation component that communicates with a second computer system.

[0148] In some examples, the second computer system and the first computer system are not physically connected. In some examples, the second computer system and the first computer system are physically connected by wires and are not located in the same enclosure.

[0149] Returning to Figure 3A, the region tracking unit 380 and the visibility analysis unit 390 are configured to determine the region of interest in the 3D scene and to determine a visibility metric that represents the amount of the region of interest represented by the captured image data. The visibility analysis unit 390 is further configured to cause the device (e.g., 1000 in Figures 10A to 10G) to provide various outputs that depend on the visibility metric, as will be described below with respect to Figures 10A to 10G.

[0150] The region tracking unit 380 is configured to determine and update regions of interest within the 3D scene in which the user is immersed (e.g., 802, 808, 906, 916, 1012, or 1092 in Figures 8A-8B, 9A-9B, and 10A-10G). The regions of interest may include one or more objects (e.g., physical or virtual objects) to which the user can issue corresponding user requests, such as "What is this?", "How much does this cost?", or "Add this to my shopping list."

[0151] In some cases, the region of interest is forward relative to the user's head posture (the position and orientation of the user's head). For example, the region of interest is in front of the user's head (e.g., in front of the user's face), and the user can view the region of interest without changing their current head posture. Because the region of interest is forward relative to the user's head posture, the region of interest (e.g., its position in the 3D scene) changes as the user's head posture changes, such as when the user rotates their head and / or when the user moves around. For example, if the user turns their head from looking straight ahead to looking up, the region of interest changes from the area directly in front of the user to the area upward relative to the user. As another example, if the user turns their head from looking right to looking left, the region of interest changes from being right-facing relative to the user to being left-facing relative to the user. As yet another example, if the user rotates 180 degrees while maintaining a neutral head position, the region of interest changes from the area previously in front of the user to the area currently in front of the user (and previously behind the user). When a user issues a request related to an object in a 3D scene, such as "How much does this cost?", they are likely to be referring to an object in front of their head, so their region of interest is forward of the user's head orientation.

[0152] In some examples, the region tracking unit 380 determines the region of interest based on the respective positions and orientations of two different devices (e.g., 1002 and 1004 in Figures 10A to 10G). The positions and orientations are calculated relative to the user (e.g., the user's head, face, and / or ears), to a coordinate system centered between the two different devices, and / or to a 3D scene (e.g., a reference point in the 3D scene (e.g., an object) or a surface in the 3D scene (e.g., a bottom)). In some examples, the two different devices are worn simultaneously by the user. For example, the two different devices are a first device (e.g., a first camera) worn on a first side of the user's head and a second device (e.g., a second camera) worn on a second side opposite the user's head. In some examples, the first device is worn on the user's first ear (e.g., the left or right ear) (e.g., by being physically housed in a wearable device), and the second device is worn on the user's other second ear (e.g., the left or right ear) (e.g., by being physically housed in a wearable device). In some examples, the first and second devices are physically housed in a single device worn by the user (e.g., a headset or glasses), such as being worn on the user's head. For example, the first device is a first camera positioned on the first side of the single device (e.g., positioned near the right or left side of the user's head when the single device is worn on the user's head), and the second device is a second camera positioned on the second side opposite the single device (e.g., positioned near the right or left side of the user's head when the single device is worn on the user's head).

[0153] Figure 7 illustrates a forward-facing head pose 702 determined based on the pose 704 (position and orientation) of the first device and the pose 706 (position and orientation) of the second device, in several examples. As described below, the region tracking unit 380 determines the region of interest based on the forward-facing head pose 702. The forward-facing head pose 702 corresponds to a user view cone 708 (e.g., 806 in Figure 8A) that represents at least a portion of the user's view of the 3D scene when the user's head has the forward-facing head pose 702. The pose 704 of the first device corresponds to a view cone 710 of the first device that represents the view of the 3D scene captured by the camera of the first device when the first device has the pose 704. The pose 706 of the second device corresponds to a view cone 712 of the second device that represents the view of the 3D scene captured by the camera of the second device when the second device has the pose 706.

[0154] In some examples, to determine the region of interest, the region tracking unit 380 relies on a 6-degree of freedom (6DOF) relationship between poses 704 and 706 and the forward-facing head pose 702. Specifically, based on the 6DOF information of the first device's pose 704 (e.g., values ​​for three spatial dimensions indicating position and three angular dimensions indicating orientation) and the 6DOF information of the second device's pose 706, the region tracking unit 380 determines (e.g., approximates) the 6DOF information of the forward-facing head pose 702. In some examples, the position of the forward-facing head pose 702 determined based on the 6DOF relationship corresponds (e.g., approximates) to a position at the center between the user's two eyes. The position at the center between the user's two eyes can provide an accurate reference point from which the forward-facing region of interest is determined. In some examples, determining a forward-facing head pose 702 using a 6DOF relationship involves calculating the relative relationship between a first device and a second device (e.g., representing how the first and second devices are oriented and positioned relative to each other), calculating the relative relationship between the first device and the user (e.g., the user's ears or the user's head) (e.g., representing how the first device is positioned and oriented relative to the user), and calculating the relative relationship between the second device and the user (e.g., the user's ears or the user's head) (e.g., representing how the second device is positioned and oriented relative to the user).

[0155] The orientations of the first and second devices (orientations of postures 704 and 706) may not correspond to the orientation of the forward-facing head posture 702 due to the various ways in which the first and second devices are worn. For example, in the default way of wearing the first and second devices (e.g., the default orientation of the first and second devices relative to the user's head and / or ears), the first and second cameras may point in approximately the same direction as the forward-facing head posture 702 (e.g., as shown by the field cones 710, 712, and 708 in Figure 7). However, the user may wear the first and / or second devices in a non-default way (e.g., rotated upward, downward, or sideways relative to the user's head and / or ears), resulting in the first and / or second cameras pointing in different directions than the orientation of the forward-facing head posture 702. Therefore, by calculating the relative relationship between the first device and the user and / or the relative relationship between the second device and the user, the area tracking unit 380 can determine whether the orientations of the first and second devices correspond to the orientation of the forward-facing head posture 702 (for example, when the user is facing straight ahead and the first and second cameras are also pointing straight ahead, or when the user is facing upward and the first and second cameras are also pointing upward) or do not correspond (for example, when the user is facing straight ahead but the first and / or second cameras are rotated upward, or when the user is facing upward but the first and / or second cameras are pointing straight ahead).

[0156] Figures 8A and 8B illustrate the regions of interest (802 in Figure 8A or 808 in Figure 8B) determined based on a forward-facing head posture 702, using several examples.

[0157] In Figure 8A, the user 800's head has a forward-facing head pose 702. The region tracking unit 380 defines the region of interest 802 by defining a forward-facing view cone 806 centered on the forward-facing head pose 702, with the position of the forward-facing head pose 702 as the starting point. In some examples, the forward-facing view cone 806 has predetermined dimensions defined by a first angular deviation (e.g., ±30°) in a first dimension (e.g., the up and down dimension in Figure 8A) orthogonal to the orientation of the forward-facing head pose 702 (e.g., represented by an arrow) and a second angular deviation (e.g., ±25°) in a second dimension (e.g., the left and right dimension in Figure 8A) orthogonal to both the orientation of the forward-facing head pose 702 and the first dimension, such that the slices of the forward-facing view cone 806 are circles or ellipses. In this way, the forward-facing view cone 806 represents the region of the 3D scene that is assumed to be in front of the user 800, and the boundary of the forward-facing view cone 806 approximates the boundary between the region of the 3D scene in front of the user 800 and the region of the 3D scene around the user 800. In some examples, the region tracking unit 380 defines the region of interest 802 (indicated by diagonal lines) as the portion of the forward-facing view cone 806 that is at least a predetermined distance 804 (e.g., 0.1 meters, 0.2 meters, 0.3 meters, 0.4 meters, or 0.5 meters) away from the user 800 (e.g., the user 800's head or face). The region of interest 802 is defined as described above because it is unlikely that the user 800 will issue a user request regarding an object that is too close to their face, and also unlikely that the user 800 will issue a user request regarding an object that is around them.

[0158] In Figure 8B, user 800's head has a forward-facing head posture 702. The region tracking unit 380 defines a region of interest 808 (indicated by diagonal lines) based on the forward-facing head posture 702 and a typical region for a handheld object. The typical region for a handheld object designates, relative to the arbitrary user's forward-facing head posture, the typical area (e.g., a 1-sigma confidence region, a 2-sigma confidence region, or a 3-sigma confidence region) where an arbitrary user would hold an object in their hand when issuing a request about the object, such as "Tell me about this object in my hand." In some examples, the typical region for a handheld object is determined through user studies and / or tests in which a group of users are asked to hold an object in their respective hands and issue user requests about the object. The typical region for a handheld object is a typical distance from the arbitrary user's head and / or face and has typical dimensions (e.g., length, width, depth, shape, or volume). The area tracking unit 380 defines a region of interest 808 (indicated by diagonal lines) based on typical distances and typical dimensions of a typical region for a handheld object. The area tracking unit 380 further positions the region of interest 808 based on the orientation of the forward-facing head posture 702 (represented by the arrow in Figure 8B) (for example, so as to intersect the arrow or have a predetermined angular deviation from the arrow). The region of interest 808 is defined as described above to take into account a scenario in which the user is holding the object in their hand when issuing a request for the object.

[0159] The visibility analysis unit 390 is configured to determine a visibility metric that represents the amount of region of interest (e.g., 802 and 808) depicted by the captured image data. For example, a high visibility metric indicates that a relatively large region of interest is depicted by the captured image data, while a low visibility metric indicates that a relatively small (or none at all) region of interest is depicted by the captured image data. As will be explained below with respect to Figures 10A to 10G, in response to receiving a user request (e.g., in the form of natural language input) related to an object in the 3D scene, the visibility analysis unit 390 causes device 1000 to perform one or more actions that depend on the visibility metric. For example, if the visibility metric is high, it is likely that the correct object for the user request is sufficiently visible (e.g., depicted) in the captured image data, and therefore device 1000 provides an output that satisfies the user request (e.g., provides information about the correct object and / or performs a task based on the correct object). However, if the visibility metric is low, the correct object for the user request may not be clearly visible in the captured image data. Therefore, before providing output that satisfies the user request, device 1000 takes various actions to ensure the correct object is clearly visible in the captured image data (for example, to avoid outputting incorrect audio output for incorrect objects and / or to avoid outputting errors resulting from failure to identify the correct object).

[0160] In some cases, the visibility metric determined depends on the type of natural language input, including user requests. In some cases, the visibility analysis unit 390 invokes the natural language processing capabilities of the DA unit 350 to determine whether the natural language input is of a first type (e.g., referring to an object in the user's hand) or a second type (e.g., not referring to an object in the user's hand). If the natural language input is of the first type, the visibility analysis unit 390 selects region of interest 808 (Figure 8B) as the region of interest and determines the visibility metric for region of interest 808. If the natural language input is of the second type, the visibility analysis unit 390 selects region of interest 802 (Figure 8A) as the region of interest and determines the visibility metric for region of interest 802. Thus, the visibility analysis unit 390 can favorably select the appropriate region of interest for determining the visibility metric based on the content of the natural language input. For example, when a user issues a request indicating that an object is in their hand, such as "How much does this object in my hand cost?", it can be advantageous to determine how much of the handheld object's area (e.g., 808) is visible. Even if a user issues a request about an object without indicating that it is in their hand (e.g., "How much does this cost?"), it can still be advantageous to determine how much of the forward-facing area (e.g., 802) is visible, as the object is likely located within the forward-facing area.

[0161] In some examples, image data is captured by two separate cameras, e.g., a first camera and a second camera as described above with respect to the region tracking unit 380 (e.g., captured simultaneously). In examples where two separate cameras are worn on the sides of the user's head (e.g., both sides) (e.g., one camera on each ear), it may be desirable to capture image data with two separate cameras in order to capture a complete view of the region of interest (e.g., 802 or 808). For example, the first camera alone cannot capture a complete view of the region of interest because part of the view of the 3D scene from the first camera is obscured by the corresponding side of the user's head. Similarly, the second camera alone cannot capture a complete view of the region of interest because part of the view of the 3D scene from the second camera is obscured by the corresponding side of the user's head. Therefore, in some examples, image data refers to a combination of two separate images (or two separate sets of images) each captured by different cameras.

[0162] Figures 9A and 9B illustrate the determination of the visibility metric for the region of interest 802 or 808 using several examples. Generally, the visibility analysis unit 390 determines the visibility metric by determining the amount of overlap between the unoccluded area of ​​the 3D scene depicted by the image data and the region of interest. For example, the visibility analysis unit 390 identifies pixels in the image data that represent an unoccluded view of the 3D scene and also indicate the region of interest. It will be understood that the image data may be occluded by the user's face, the user's hair, the user's clothing, the user's hands, dirt on the camera, etc. The visibility analysis unit 390 implements techniques known in the art to classify different parts of the image data as being occluded over unoccluded areas (e.g., to distinguish between occluded and unoccluded pixels).

[0163] Figure 9A illustrates an example in which the visibility analysis unit 390 determines a relatively low visibility metric for a region of interest 906 (e.g., 802 or 808). In Figure 9A, the camera region 902 represents the region of the 3D scene that is depicted by the image data, such as the region that is depicted by the image data when the image data is not occluded. In some examples, the visibility analysis unit 390 determines the camera region 902 based on the poses 704 and 706 of two different devices, such as two different cameras. For example, the visibility analysis unit 390 maps the 3D scene and determines the camera region 902 using the poses 704 and 706 as well as the camera's field of view.

[0164] In Figure 9A, the camera region 902 includes the occluded region 904 (indicated by horizontal hatching). The occluded region 904 represents the portion of the 3D scene that is occluded in the image data. In Figure 9A, the visibility analysis unit 390 projects the camera region 902 onto the region of interest 906 to determine the correlation between the camera region 902 and the region of interest 906 (for example, by using the orientations 702, 704, and 706, the known dimensions and position of the region of interest 906, and the known dimensions of the camera's field of view). For example, the visibility analysis unit 390 identifies the overlapping region 908 (indicated by vertical hatching) that represents the overlap between the camera region 902 and the region of interest 906 (for example, how much of the region of interest 906 would be depicted by the camera region 904 if the image data were not occluded). Next, the visibility analysis unit 390 identifies the unoccluded portion of the overlapping region 908, for example, the portion of the overlapping region 908 that has vertical hatching but no horizontal hatching, as the visibility region 910. Then, the visibility analysis unit 390 determines the visibility metric based on the pixels of the image data representing the visibility region 910.

[0165] In some examples, different pixels representing the viewing region 910 have different weights in determining the visibility metric. For example, the positive contribution of pixels in the viewing region 910 to the visibility metric decreases as the pixel moves away from the center of the region of interest 906 or 916. As a result, pixels representing the central portion of the region of interest 906 or 916 provide a larger positive contribution to the visibility score than pixels representing the edges of the region of interest 906 or 916. For example, assume that the same number of pixels depict the region of interest 906. If the same number of pixels primarily depict the central portion of the region of interest 906, the resulting visibility metric will be higher than if the same number of pixels primarily depict the edges of the region of interest 906.

[0166] In Figure 9A, the visibility analysis unit 390 determines a relatively low visibility metric because there are relatively few pixels representing the viewable area 910, and because the pixels representing the viewable area 910 have a relatively low weight in determining the visibility metric (for example, because the pixels represent the edge of the region of interest 906). The example in Figure 9A results in a relatively low visibility metric due to the significant occlusion of the camera region 902 (indicated by the size of the occluded region 904 within the camera region 902) and the relatively small amount of overlap between the camera region 902 and the region of interest 906. The small amount of overlap between the camera region 902 and the region of interest 906 may be because the camera orientation does not correspond to the orientation of the user's head in a forward-facing position. For example, the small amount of overlap between region 902 and region 906 is because the user is looking straight ahead and the camera is rotated upward to capture images of the region mainly above the user.

[0167] Figure 9B illustrates an example of how the visibility analysis unit 390 determines a relatively high visibility metric for a region of interest 916 (e.g., 802 or 808). In Figure 9B, the camera region 912 represents the region of the 3D scene that is depicted by the image data, such as the region that is depicted by the image data if the image data is not occluded. The camera region 912 is similar to the camera region 902 and is determined in a similar manner. The camera region 912 includes an occluded region 914 (indicated by horizontal hatching). The occluded region 914 represents the portion of the 3D scene that is occluded in the image data. In Figure 9B, as in Figure 9A, the visibility analysis unit 390 projects the camera region 912 onto the region of interest 916 and identifies an overlapping region 918 (indicated by vertical hatching) that represents the overlap between the camera region 912 and the region of interest 916 (e.g., how much of the region of interest 916 is depicted by the camera region 912 if the image data is not occluded). The visibility analysis unit 390 identifies the unoccluded portion of the overlapping region 918 (the portion of the overlapping region 918 that has vertical hatching but no horizontal hatching) as the visibility region 920. The visibility analysis unit 390 then determines the visibility metric based on the pixels of the image data representing the visibility region 920.

[0168] In Figure 9B, the visibility analysis unit 390 determines a relatively high visibility metric because there are a relatively large number of pixels representing the viewing area 920, and because some pixels representing the viewing area 920 have relatively high weights in determining the visibility metric (for example, because the pixels represent the central part of the region of interest 916). The example in Figure 9B results in a relatively high visibility metric due to the relatively small amount of occlusion within the camera area 912 (indicated by the size of the occluded area 914 within the camera area 912) and the relatively large amount of overlap between the camera area 912 and the region of interest 916. The large overlap between the camera area 912 and the region of interest 916 may be because the camera orientation corresponds to the orientation of the user's head in a forward-facing posture. For example, the large overlap between the camera area 912 and the region of interest 916 is because the user is looking straight ahead, and the camera is also oriented to face almost straight ahead (for example, relative to the user's head).

[0169] As explained below with respect to Figures 10A to 10G, the visibility analysis unit 390 is configured to cause device 1000 to perform various actions based on whether the determined visibility metric satisfies a condition (e.g., a threshold). In some examples, the visibility metric satisfies the condition if the visibility metric is greater than or equal to a threshold, and does not satisfy the condition if the visibility metric is less than a threshold.

[0170] Figures 10A to 10G illustrate, through several examples, a device 1000 that performs various actions in response to receiving natural language input, according to a determined visibility metric.

[0171] Device 1000 implements at least some of the components of computer system 101. In some examples, device 1000 includes one or more sensors configured to detect audio data (e.g., natural language user requests), one or more cameras configured to detect subject image data on which a visibility metric is determined, and one or more audio output devices (e.g., speakers) configured to provide audio output. In the examples in Figures 10A to 10G, device 1000 is worn on the head of user 1010. For example, device 1000 is an XR headset, smart glasses, headphones, or a set of earphones. In other examples, device 1000 is another type of electronic device such as a smartwatch, smartphone, tablet device, laptop computer, or projection-based device.

[0172] In the examples of Figures 10A to 10G, device 1000 includes camera 1002 (e.g., a set of one or more cameras) and camera 1004 (e.g., a set of one or more cameras). Camera 1002 is located close to a first side of user 1010's head (e.g., worn), and camera 1004 is located close to a second side of user 1010's head (e.g., worn). For example, camera 1002 is physically housed in a first device (e.g., a first earphone) worn on the first side of user 1010's head (e.g., worn on user 1010's first ear), and camera 1004 is physically housed in a second device (e.g., a second earphone) worn on a second side of user 1010's head (e.g., worn on user 1010's second ear).

[0173] In Figures 10A to 10G, the left portion illustrates a coordinate system defining the horizontal dimension x, height dimension y, and depth dimension z. Figures 10A to 10G illustrate a coordinate system defining the relative directions, i.e., up, down, right, left, front, and back, in the context of Figures 10A to 10G and with respect to user 1010. Specifically, if an object or region has a height coordinate y greater than the height coordinate y' of user 1010, the object or region is upward (e.g., up) relative to user 1010; if an object or region has a height coordinate y less than the height coordinate y' of user 1010, the object or region is downward (e.g., down) relative to user 1010; if an object or region has a horizontal coordinate x greater than the horizontal coordinate x' of user 1010, the object or region is to the right relative to user 1010; and If an object or region has a horizontal coordinate x smaller than the horizontal coordinate x' of user 1010, the object or region faces left relative to user 1010; if an object or region has a depth coordinate z larger than the depth coordinate z' of user 1010, the object or region faces forward relative to user 1010 (e.g., in front of user 1010); and if an object or region has a depth coordinate z smaller than the depth coordinate z' of user 1010, the object or region faces backward relative to user 1010 (e.g., behind user 1010). In some examples, the height, horizontal, and depth coordinates (x', y', z') of user 1010 are the height, horizontal, and depth coordinates of user 1010's head. In some examples, the height, horizontal, and depth coordinates of user 1010 are the height, horizontal, and depth coordinates of another part of user 1010, such as user 1010's face or chest.

[0174] In Figures 10A to 10G, the right-hand portion of the figures illustrates images captured by device 1000 (e.g., 1018, 1026, 1046, 1076, or 1093) or a user interface displayed by external device 1062 (e.g., 1068).

[0175] In Figure 10A, camera 1002 has a forward orientation 1002-1 (e.g., orientation of posture 706), camera 1004 has a forward orientation 1004-1 (e.g., orientation of posture 704), and user 1010's head (e.g., head posture) has a forward orientation 1006 (e.g., orientation of forward head posture 702). The forward orientation 1006 corresponds to the region of interest 1012 (e.g., 802). In Figure 10A, the orientations 1002-1 and 1004-1 of cameras 1002 and 1004, respectively, almost coincide with orientation 1006 without camera occlusion, so the image 1018 collectively captured by cameras 1002 and 1004 will depict a relatively large portion of the region of interest 1012.

[0176] In Figure 10A, the 3D scene is in front of user 1010 and includes object 1014 within region of interest 1012. In Figure 10A, device 1000 receives a natural language request 1016, “How much does this cost?” spoken by user 1010, because user 1010 wants to know the price of object 1014. Cameras 1002 and 1004 capture an image 1018 associated with request 1016. For example, camera 1002 captures a first individual image, camera 1004 simultaneously captures a second individual image, and device 1000 constructs image 1018 based on combining the first individual image and the second individual image. The first and second individual images are captured simultaneously with and / or in response to receiving request 1016.

[0177] Image 1018 includes a relatively small occluded region 1020 representing the portion of the 3D scene that is occluded in Image 1018 (for example, due to the occluding of cameras 1002 and / or 1004 by the face of user 1010, the hair of user 1010, the clothing of user 1010, dirt on cameras 1002 and / or 1004, etc.). Image 1018 further includes region 1022 (inside the dashed line) corresponding to a portion of the region of interest 1012 (for example, corresponding to overlapping region 908 or 918). In Figures 10A to 10G, the dashed lines in the images (e.g., 1018, 1026, 1046, 1076, or 1093) are for illustrative purposes only and are not present in the respective images. Due to the relatively small amount of occlusion, and because region 1022 corresponds to a large region of interest 1012, image 1018 contains a relatively complete depiction of the object 1014 that the user is asking about.

[0178] In Figure 10A, because image 1018 has relatively little occlusion and because image 1018 corresponds to most of the region of interest 1012, in response to receiving user request 1016, device 1000 determines a high visibility metric for the region of interest 1012 (for example, according to the technique described above with respect to Figures 9A-9B). Because the visibility metric is high (for example, above the threshold), device 1000 attempts to perform a task to satisfy request 1016, and device 1000 provides an audio output 1024. Specifically, the digital assistant (for example, provided by DA unit 350) processes request 1016 in conjunction with image 1018 to determine the price of object 1014, and device 1000 provides an audio output 1024 saying, "This costs $100."

[0179] In Figure 10B, similar to Figure 10A, camera 1002 has a forward orientation 1002-1 (e.g., orientation of posture 706), camera 1004 has a forward orientation 1004-1 (e.g., orientation of posture 704), and user 1010's head (e.g., head posture) has a forward orientation 1006 (e.g., orientation of forward head posture 702). The forward orientation 1006 corresponds to the region of interest 1012 (e.g., 802). In Figure 10B, the orientations 1002-1 and 1004-1 of cameras 1002 and 1004, respectively, almost coincide with orientation 1006 without camera occlusion, so the image 1026 collectively captured by cameras 1002 and 1004 will depict a relatively large portion of the region of interest 1012.

[0180] In Figure 10B, the 3D scene is in front of user 1010 and includes objects 1028 and 1030 within the region of interest 1012. In Figure 10B, device 1000 receives a natural language request 1032, “How much does this cost?” spoken by user 1010, because user 1010 wants to know the price of object 1028. Cameras 1002 and 1004 capture image 1026 associated with request 1032, similar to how cameras 1002 and 1004 capture image 1018 in Figure 10A, for example.

[0181] Image 1026 includes a relatively small occluded region 1034 representing the portion of the 3D scene that is occluded in Image 1026. Image 1026 further includes region 1036 (inside the dashed line) which corresponds to a portion of the region of interest 1012 (for example, to overlapping regions 908 or 918). Because the occlusion is relatively small and region 1036 corresponds to a large portion of the region of interest 1012, Image 1026 includes a relatively complete depiction of objects 1028 and 1030 of potential user interest.

[0182] In Figure 10B, because image 1026 has little occlusion and corresponds to a large number of regions of interest 1012, device 1000 determines a high visibility metric for region of interest 1012 in response to receiving request 1032 (for example, according to the techniques described above with respect to Figures 9A to 9B). Because the visibility metric is high (e.g., above the threshold), device 1000 attempts to perform a task to satisfy request 1032. Specifically, the digital assistant processes image 1026 together with request 1032, "How much does this cost?", and attempts to determine an answer to request 1032. The digital assistant determines that image 1026 contains multiple objects 1028 and 1030 (e.g., multiple objects are detected within region of interest 1012). Device 1000 then provides an audio output 1038, “Which object do you intend?”, requesting user 1010 to resolve the ambiguity between objects 1028 and 1030. After device 1010 provides the audio output 1038, device 1010 receives a response 1040, “The object on the left,” spoken by user 1010, which resolves the ambiguity between objects 1028 and 1030. In response to receiving response 1040, the digital assistant processes request 1032, response 1040, and image 1026 together to determine that the price of object 1028 is $150, and device 1000 provides an audio output 1042, “The price of the object on the left is $150.”

[0183] In some examples, before device 1000 receives request 1032, device 1000 captures one or more images of the 3D scene, and device 1000 detects object 1044 based on the captured images. In some examples, in response to receiving request 1032, device 1000 provides one or more audio outputs based on the detected object 1044. For example, one or more audio outputs refer to the respective positions of objects 1028 and / or 1030 relative to the detected object 1044, such as the detection of objects 1028, 1030, and 1044, the head pose of user 1010, and their respective positions determined based on a map of the 3D scene. As a specific example, if object 1044 is a green ball, audio output 1038 would instead be, "Do you mean the object closer to the green ball, or the object further away from the green ball?" and / or audio output 1042 would instead be, "The object closer to the green ball costs $150." In this way, when providing audio output, device 1000 uses the location and / or identification information of previously detected objects, which may help user 1010 provide device 1000 with an improved response (e.g., one that device 1000 can interpret more accurately), and may help device 1000 provide user 1010 with an improved (e.g., more informative and / or clearer) audio output.

[0184] In Figure 10C, similar to Figure 10B, camera 1002 has a forward orientation 1002-1 (e.g., orientation of posture 706), camera 1004 has a forward orientation 1004-1 (e.g., orientation of posture 704), and user 1010's head (e.g., head posture) has a forward orientation 1006 (e.g., orientation of forward head posture 702). The forward orientation 1006 corresponds to the region of interest 1012 (e.g., 802). In Figure 10C, the orientations 1002-1 and 1004-1 of cameras 1002 and 1004, respectively, almost coincide with orientation 1006 without camera occlusion, so the image 1046 collectively captured by cameras 1002 and 1004 will depict a relatively large portion of the region of interest 1012.

[0185] In Figure 10C, the 3D scene is in front of user 1010 and includes object 1048 within the region of interest 1012. In Figure 10C, device 1000 receives a natural language request 1050 “Add this to my shopping list” spoken by user 1010, because user 1010 wants to add object 1048 to his shopping list. Cameras 1002 and 1004 capture image 1046 associated with request 1050, similar to how cameras 1002 and 1004 capture image 1018 in Figure 10A, for example.

[0186] Image 1046 includes a relatively large occluded region 1052 representing the portion of the 3D scene that is occluded in Image 1046. In Figure 10C, Image 1046 includes an occluded region 1052 resulting from the occlusion of cameras 1002 and / or 1004 by the face of user 1010, user 1010's hair, and / or user 1010's clothing. Image 1046 further includes region 1054 (inside the dashed line) which corresponds to a portion of the region of interest 1012 (for example, corresponding to overlapping regions 908 or 918). Although Image 1046 corresponds to a large portion of the region of interest 1012, due to the large amount of occlusion, Image 1046 does not depict the object 1048 that user 1010 asks about.

[0187] In Figure 10C, even though image 1046 corresponds to a large portion of region of interest 1012, due to a large amount of occlusion, device 1000, in response to receiving request 1050, determines a low visibility metric for region of interest 1012 (for example, according to the techniques described above with respect to Figures 9A-9B). Because the visibility metric is low (e.g., below the threshold), device 1000 provides audio output 1056, "Which object do you mean?", requesting user 1010 to specify the object 1048 corresponding to request 1050 (e.g., to specify which object user 1010 is referring to). After device 1000 provides audio output 1056, device 1000 receives response 1058, "The object in front of me," spoken by user 1010. In response to receiving response 1058, the digital assistant processes response 1058 and image 1046 to determine whether image 1046 depicts object 1048. For example, based on the user 1010's head posture, the orientation of camera 1002 1002-1 and camera 1004 1004-1, and image 1048, the digital assistant determines whether an object in front of user 1010 can be detected (e.g., identified) from image 1048 with sufficient confidence. In Figure 10C, the digital assistant determines that image 1046 does not depict object 1048 (due to image occlusion), and therefore device 1000 provides audio output 1060 “The object is not visible. Please capture an image of the object with your phone” requesting user 1010 to capture an image of object 1048 using external device 1062 (Figures 10D-10E).

[0188] In Figure 10D, after device 1000 provides an audio output 1060, user 1010 holds external device 1062 to capture an image of a desired object 1048. External device 1062 implements at least some of the components of computer system 101 and includes one or more cameras. Figures 10D to 10E show that external device 1062 is a smartphone, but in other examples, external device 1062 is a different type of device, such as a laptop computer, tablet device, or smartwatch.

[0189] In Figure 10D, after providing (or simultaneously with) the audio output 1060, the external device 1062 displays a camera icon 1066 prompting the user 1010 to activate one or more cameras of the external device 1010. In Figure 10D, the external device 1062 receives user input 1064 (e.g., touch input, speech input, gesture input, motion input, gaze input, and / or input received via a peripheral device) to select the camera icon 1066.

[0190] In Figure 10E, in response to receiving user input 1066, the external device 1062 displays a camera user interface 1068. The camera user interface 1068 includes a live view of a 3D scene captured by one or more cameras of the external device 1062 and includes a selectable shutter button 1070 for capturing an image. The live view depicts object 1048 because user 1010 has pointed one or more cameras of the external device 1062 at object 1048 to capture an image of object 1048. In Figure 10E, the external device 1062 receives user input 1072 (e.g., touch input, speech input, gesture input, motion input, gaze input, and / or input received via a peripheral device) to select the shutter button 1070. In response to receiving user input 1072, the external device 1062 captures an image of object 1048. Next, the digital assistant processes the image of object 1048 in conjunction with request 1050, "Add this to my shopping list," and performs the requested task of adding object 1048 (e.g., canned sardines) to user 1010's shopping list, and device 1000 provides audio output 1074, "OK, canned sardines have been added to your shopping list." In this way, when the initial image 1046 does not provide sufficient information for device 1000 to satisfy user request 1050 (e.g., due to image occlusion, and / or the respective orientations 1002-1 and / or 1004-2 of cameras 1002 and / or 1004), device 1000 and external device 1062 can satisfy user request 1050 without requiring user 1010 to repeat the user request 1050.

[0191] The examples in Figures 10D to 10E illustrate how the external device 1062 displays the camera user interface 1068 in response to receiving user input 1064 selecting the camera icon 1066. In other examples, the external device 1062 displays the camera user interface 1068 in response to a different trigger event. For example, the external device 1062 may automatically display the camera user interface 1068 without user input at the same time that device 1000 provides audio output 1060, the external device 1062 may automatically display the camera user interface 1068 without user input after device 1000 provides audio output 1060 (for example, thus Figure 10C proceeds directly to Figure 10E), the external device 1062 may display the camera user interface in response to detection of motion input corresponding to a movement of lifting the external device 1062 (for example, a movement associated with taking the external device 1062 out of user 1010's pocket or bag) after device 1000 provides audio output 1060, or the external device 1062 may display the camera user interface 1068 in response to the user selecting an icon to launch the camera application after device 1000 provides audio output 1060.

[0192] In some examples, if device 1000 determines that the visibility metric does not meet the criteria (e.g., falls below a threshold), device 1000 provides an audio output 1060 that requests user 1010 to capture an image of object 1048 using external device 1062, without providing an audio output 1056 that asks user 1010 to specify object 1048 (e.g., to specify which object user 1010 is referring to) (and / or without determining whether image 1046 depicts object 1048). After device 1000 provides audio output 1060 (or simultaneously with device 1000 providing audio output 1060), external device 1062 displays camera user interface 1068 according to the techniques discussed above. Thus, in some examples, if device 1000 determines the visibility metric is low in response to receiving a user request, device 1000 directly prompts user 1010 to capture an image of the relevant object without asking user 1010 to specify the relevant object.

[0193] In Figure 10F, camera 1002 has a downward orientation 1002-2 (e.g., orientation of posture 706), camera 1004 has a downward orientation 1004-2 (e.g., orientation of posture 704), and user 1010's head (e.g., head posture) has a forward orientation 1006 (e.g., orientation of forward-facing head posture 702). The forward orientation 1006 corresponds to the region of interest 1012 (e.g., 802). In Figure 10F, since the orientations 1002-2 and 1004-2 of cameras 1002 and 1004 respectively do not coincide with orientation 1006, the image 1076 collectively captured by cameras 1002 and 1004 depicts a relatively small portion of the region of interest 1012 (or none of the region of interest 1012).

[0194] In Figure 10F, the 3D scene includes an object 1078 that is facing downwards relative to user 1010 (e.g., on the floor) and is not within the region of interest 1012. In Figure 10F, device 1000 receives a natural language request 1080, “How much does this cost?” spoken by user 1010, because user 1010 wants to know the price of object 1078. Figure 10F illustrates an example where user 1010 requests that a task be performed based on object 1078, which is not within the region of interest 1012. Specifically, while user 1010 is looking forward, user 1010 issues a request 1080 to ask about object 1078 which is on the floor (e.g., in the xz plane) (e.g., user 1010's head is facing forward, but user 1010 is looking downwards to ask about object 1078 which is on the floor).

[0195] In Figure 10F, cameras 1002 and 1004 capture image 1076 associated with request 1080, similar to how cameras 1002 and 1004 capture image 1018 in Figure 10A. Image 1076 includes a relatively small occluded region 1082 representing the portion of the 3D scene that is occluded in image 1076. Image 1076 further includes region 1084 (inside the dashed line) corresponding to a portion of the region of interest 1012 (e.g., corresponding to overlapping region 908 or 918). Region 1084 is relatively small because orientations 1002-2 and 1006 do not coincide, and orientations 1004-2 and 1006 do not coincide. For example, since cameras 1002 and 1004 are facing downwards relative to user 1010, only the top of image 1076 depicts the region of interest 1012 that is facing forwards relative to user 1010. Image 1076 depicts object 1078, which user 1010 is asking about.

[0196] In Figure 10F, although the occlusion in image 1076 is small, the portion of region of interest 1012 depicted by image 1076 is small, so in response to receiving request 1080, device 1000 determines a low visibility metric for region of interest 1012 (for example, according to the technique described above with respect to Figures 9A-9B). Because the visibility metric is low (for example, below the threshold), device 1000 provides audio output 1086 “Which object do you mean?” requesting user 1010 to specify the object 1078 corresponding to request 1080 (for example, to specify which object user 1010 is referring to). After device 1000 provides audio output 1086, device 1000 receives the response 1088 “The object on the floor below me” spoken by user 1010. In response to receiving response 1088, the digital assistant processes response 1088 and image 1076 to determine whether image 1076 depicts object 1078. For example, based on user 1010's head pose, camera 1002 orientation 1002-2 and camera 1004 orientation 1004-2, and image 1048, the digital assistant determines whether an object on the floor beneath user 1010 can be detected (e.g., identified) from image 1076 with sufficient confidence. In Figure 10F, the digital assistant determines that image 1076 depicts object 1078 (e.g., with a sufficient degree of confidence). Since image 1076 depicts object 1078, the digital assistant processes image 1076 together with request 1080, "How much does this cost?", to find the price of object 1078, and device 1000 provides audio output 1090, "The price of this object is $68."

[0197] In some examples, if device 1000 determines that the visibility metric does not meet the criteria (e.g., below a threshold), device 1000 provides an audio output (e.g., 1060 in Figure 10C) that requests user 1010 to capture an image of object 1078 using external device 1062, without providing an audio output 1086 that asks user 1010 to specify object 1078 (and / or without determining whether image 1076 depicts object 1078). In some examples, user 1010 then captures an image of object 1078 using external device 1062 so that device 1000 provides an audio output 1090 that satisfies request 1080, similar to those described with respect to Figures 10D-10E. Therefore, in some examples, even if the initial image 1076 may already depict the relevant object 1078, as in the case of Figure 10F, if the visibility metric does not meet the requirements, device 1000 prompts user 1010 to use external device 1062 to capture an image of the relevant object 1078.

[0198] In some examples, before device 1000 receives request 1080, device 1000 captures one or more images of the 3D scene, and device 1000 detects object 1091 based on the captured images. In some examples, in response to receiving request 1080, device 1000 provides one or more audio outputs based on the detected object 1091. For example, one or more audio outputs refer to the position of object 1078 relative to the position of detected object 1091. As a specific example, if object 1091 is a red ball, for example, similar to how device 1000 uses the position of previously detected object 1044 in Figure 10B, then audio output 1086 would instead be "Do you mean the object under the red ball?" and / or audio output 1090 would instead be "The price of the object under the red ball is $68."

[0199] In Figure 10G, camera 1002 has a forward orientation 1002-3 (e.g., orientation of posture 704), camera 1004 has a forward orientation 1004-3 (e.g., orientation of posture 706), and user 1010's head (e.g., head posture) has a forward orientation 1006 (e.g., orientation of forward head posture 702). The forward orientation 1006 corresponds to the region of interest 1092 (e.g., 808). In Figure 10G, the orientations 1002-3 and 1004-3 of cameras 1002 and 1004, respectively, almost coincide with orientation 1006 without camera occlusion, so the image 1093 collectively captured by cameras 1002 and 1004 will depict a relatively large portion of the region of interest 1092.

[0200] In Figure 10G, the 3D scene includes an object 1094 held in the hand of user 1010 and located in a region of interest 1092 (for example, a region of interest for handheld objects as described above with respect to Figure 8B). In Figure 10G, device 1000 receives a natural language request 1095 spoken by user 1010, "How much does this object in my hand cost?", because user 1010 wants to know the price of object 1094 (held in user 1010's hand).

[0201] In Figure 10G, cameras 1002 and 1004 capture image 1093 associated with request 1095, similar to how cameras 1002 and 1004 capture image 1018 in Figure 10A. Image 1093 includes a relatively small occluded region 1097 representing the portion of the 3D scene that is occluded in image 1093. Image 1093 further includes region 1096 (inside the dashed line) corresponding to region of interest 1092. Due to the relatively small amount of occlusion, and because region 1096 corresponds to a large portion of region of interest 1092 (e.g., all of region of interest 1092 if region 1096 were not occluded), image 1093 includes a relatively complete depiction of object 1094 in user 1010's hand.

[0202] In Figure 10G, device 1000 determines that request 1095 corresponds to an object held in user 1010's hand. Since request 1095 corresponds to an object held in user 1010's hand, device 1000 determines the visibility metric for region of interest 1092 (e.g., region of interest for handheld objects) (e.g., region of interest 808 as described in relation to Figure 8B). In contrast, in Figures 10A to 10F, device 1000 did not determine that each user request (e.g., 1016, 1032, 1050, and 1080) corresponds to an object held in user 1010's hand, so device 1000 instead determines the visibility metric for a different region of interest 1012 (e.g., region of interest 802 as described in relation to Figure 8A).

[0203] In Figure 10G, because image 1093 has relatively little occlusion and because image 1093 corresponds to most of the region of interest 1092, in response to receiving user request 1095, device 1000 determines a high visibility metric for the region of interest 1092 (for example, according to the techniques described above with respect to Figures 9A to 9B). Because the visibility metric is high (for example, above the threshold), device 1000 attempts to perform a task to satisfy request 1095, and device 1000 provides an audio output 1098. Specifically, the digital assistant processes request 1095 together with image 1093 to determine the price of object 1094, and device 1000 provides an audio output 1098 saying, "This costs $1,000."

[0204] Additional explanations regarding Figures 7, 8A-8B, 9A-9B, and 10A-10G are provided below with reference to Method 1100 described in relation to Figure 11.

[0205] Figure 11 is a flowchart of Method 1100 for providing audio output in response to natural language input, as shown in several examples. In some examples, Method 1100 is performed in a computer system (e.g., Device 1000) communicating with one or more visual imaging sensors (e.g., cameras, e.g., RGB cameras, infrared cameras, and / or depth cameras) and one or more audio output devices (e.g., speakers). In some embodiments, Method 1100 is governed by instructions stored in a non-temporary (or temporary) computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 302 of computer system 101 (e.g., Control 110 in Figure 1). In some examples, the operation of Method 1100 is distributed across multiple computer systems, such as a computer system and a separate server system. Some operations of Method 1100 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0206] Method 1100 includes receiving natural language input (e.g., 1016, 1032, 1050, 1080, or 1095) corresponding to a first object in a 3D scene (e.g., 1014, 1028, 1048, 1078, or 1094) (1102) while the head of a user of a computer system (e.g., 1010) has a head pose (e.g., 702) corresponding to a forward-facing region of a 3D scene (e.g., 802, 808, 906, 916, 1012, or 1092) (e.g., forward-facing relative to the head pose) (e.g., forward-facing region determined by the region tracking unit 380).

[0207] Method 1100 includes capturing image data (e.g., 1018, 1026, 1046, 1076, or 1093) associated with a natural language input corresponding to a first object in a 3D scene via one or more visual imaging sensors (1104).

[0208] Method 1100 includes, in response to receiving natural language input corresponding to a first object in a 3D scene (1106), providing a first audio output (e.g., 1024, 1038, or 1098) corresponding to the first object via one or more audio output devices (1108) according to a determination that a visibility metric (e.g., determined by the visibility analysis unit 390) representing the amount depicted by image data in the forward-facing region of the 3D scene does not satisfy a condition (e.g., is greater than a threshold), and providing a second audio output (e.g., 1056, 1086) (different from the first audio output) corresponding to the first object via one or more audio output devices (1110) according to a determination that a visibility metric representing the amount depicted by image data in the forward-facing region of the 3D scene does not satisfy a condition (e.g., is less than a threshold).

[0209] In some cases, the forward-facing region of the 3D scene is the first region of the 3D scene, based on the determination that the head posture is the first head posture, and the forward-facing region of the 3D scene is the second region of the 3D scene, different from the first region of the 3D scene, based on the determination that the head posture is the second head posture, different from the first head posture.

[0210] In some examples, capturing image data associated with natural language input via one or more visual imaging sensors includes capturing image data associated with natural language input via one or more visual imaging sensors while receiving natural language input.

[0211] In some examples, capturing image data associated with natural language input via one or more visual imaging sensors includes capturing image data associated with natural language input via one or more visual imaging sensors in response to receiving natural language input.

[0212] In some examples, a first audio output (e.g., 1024 or 1098) corresponding to a first object represents the result of a first task performed (e.g., by a digital assistant) based on the natural language input and the first object (e.g., a result that satisfies the user request contained in the natural language input).

[0213] In some examples, providing a first audio output corresponding to a first object via one or more audio output devices includes providing an audio output (e.g., 1038) that prompts the user to remove ambiguity between multiple detected objects, based on the determination that a forward region of a 3D scene contains multiple detected objects (e.g., 1028 and 1030) including a first object (e.g., 1028).

[0214] In some examples, a second audio output (e.g., 1056) corresponding to a first object (e.g., 1048) includes a request asking the user to specify the first object (e.g., specify the attributes of the first object (e.g., identification information, shape, size, color, orientation, and position)). In some examples, the second audio output corresponds to a request (e.g., 1060) asking the user to capture an image of the first object using an external device (e.g., 1062). In some examples, after the external device has captured an image of the first object, the computer system provides audio output (e.g., 1074) via one or more audio output devices indicating a task to be performed based on the natural language input and the captured image of the first object. In some examples, after receiving natural language input, the computer system provides audio output without receiving any further natural language input, and therefore the user does not need to repeat their initial natural language request in order for the requested task related to the object to be performed and for the results of the requested task to be output via one or more audio output devices.

[0215] In some examples, the computer system includes a first device (e.g., 1002) and a second device (e.g., 1004), wherein the first device is distinct from the second device, natural language input is received while the first device is worn by the user and while the second device is worn by the user, and the forward-facing region of the 3D scene is determined based on the position of the first device while it is worn by the user, the orientation of the first device while it is worn by the user (e.g., 1002-1, 1002-2, or 1002-3), the position of the second device while it is worn by the user, and the orientation of the second device while it is worn by the user (e.g., 1004-1, 1004-2, or 1004-3).

[0216] In some examples, one or more visual imaging sensors include a first visual imaging sensor (e.g., 1002) and a second visual imaging sensor (e.g., 1004) different from the first visual imaging sensor, and capturing image data associated with natural language input (e.g., 1018, 1026, 1046, 1076, or 1093) via one or more visual imaging sensors includes capturing first image data via the first visual imaging sensor and capturing second image data different from the first image data via the second visual imaging sensor (e.g., capturing the first and second image data simultaneously).

[0217] In some examples, first image data is captured while the first visual imaging sensor is worn on a first side of the user's head (e.g., the left or right side) (e.g., while the device containing the first visual imaging sensor is worn (e.g., while the device is inserted into the ear)), and second image data is captured while the second visual imaging sensor is worn on a second side of the user's head (e.g., the left or right side) (e.g., while the device containing the second visual imaging sensor is worn (e.g., while the device is inserted into the ear)), where the first side of the user's head is opposite to the second side of the user's head.

[0218] In some examples, according to the determination that first image data was captured while the first visual imaging sensor had a first orientation (e.g., 1002-1 in Figure 10A) (e.g., relative to a device containing the first visual imaging sensor and / or the user's head), and that second image data was captured while the second visual imaging sensor had a second orientation (e.g., 1004-1 in Figure 10A) (e.g., relative to a device containing the second visual imaging sensor and / or the user's head), the visibility metric, which represents the amount depicted by the image data (e.g., 1018) in the forward-facing region (e.g., 1012) of the 3D scene, has a first value (e.g., as described with respect to Figure 10A). In accordance with the determination that first image data was captured while the first visual imaging sensor was in a third orientation (e.g., 1002-2 in Figure 10F) different from the first orientation (e.g., relative to the device containing the first visual imaging sensor and / or the user's head), and the determination that second image data was captured while the second visual imaging sensor was in a fourth orientation (e.g., 1004-2 in Figure 10F) different from the second orientation (e.g., relative to the device containing the second visual imaging sensor and / or the user's head), the visibility metric, which represents the amount depicted by the image data (e.g., 1076) in the forward-facing region (e.g., 1012) of the 3D scene, has a second value different from the first value (e.g., as described with respect to Figure 10F) (e.g., the visibility metric depends on the respective orientations of the first visual imaging sensor when the first image data was captured and the respective orientations of the second visual imaging sensor when the second image data was captured).

[0219] In some cases, the visibility metric has a first value (for example, as described with respect to Figure 10F) according to the determination that the first and second image data depict a first amount of the forward-facing region of the 3D scene (e.g., 10¹²) (e.g., collectively). The visibility metric has a second value (for example, as described with respect to Figure 10A) according to the determination that the first and second image data depict a second amount of the forward-facing region of the 3D scene that is greater than the first amount of the forward-facing region of the 3D scene (e.g., collectively).

[0220] In some examples, the image data includes image regions (e.g., pixels) (e.g., 904, 914, 1020, 1034, 1052, 1082, or 1097) that represent occlusion of the forward-facing region of the 3D scene, and the visibility metric, which represents the amount of the forward-facing region of the 3D scene depicted by the image data, is based on the image regions that represent occlusion of the forward-facing region of the 3D scene.

[0221] In some examples, according to the determination that an image region representing occlusion of the forward-facing region of a 3D scene (e.g., 1020) has a first size, the visibility metric representing the amount depicted by image data in the forward-facing region of the 3D scene has a third value (e.g., as described with respect to Figure 10A), and according to the determination that an image region representing occlusion of the forward-facing region of a 3D scene (e.g., 1052) has a second size greater than the first size, the visibility metric representing the amount depicted by image data in the forward-facing region of the 3D scene has a fourth value less than the third value (e.g., as described with respect to Figure 10C).

[0222] In some examples, the visibility metric, which represents the amount of the forward-facing region of a 3D scene (e.g., 802, 808, 906, 916, 1012, or 1092) depicted by image data, is based on the amount of overlap between the unoccluded region of the 3D scene depicted by the image data and the forward-facing region of the 3D scene (e.g., represented by visibility region 910 or 920) (e.g., a greater amount of overlap results in a higher visibility metric, and a less amount of overlap results in a lower visibility metric).

[0223] In some examples, the forward-facing region of a 3D scene has predetermined (e.g., fixed) dimensions (e.g., length, width, depth, area, and / or volume) (e.g., dimensions independent of the user's head posture) (e.g., dimensions determined before natural language input is received and image data is captured).

[0224] In some examples, the forward-facing region of the 3D scene is at least a predetermined (e.g., fixed) distance (e.g., 804) from the user's head (e.g., the portion of the forward-facing region of the 3D scene closest to the user's head is at least a predetermined non-zero distance from the user's head).

[0225] In some examples, a handheld object region is determined, which is where each user holds each object in their hand while issuing queries about each object, and based on the handheld object region, a forward-facing region of the 3D scene (e.g., 808 or 1092) is determined.

[0226] In some examples, the forward region of the 3D scene is the third region of the 3D scene (e.g., 802 or 1012) according to the determination that the natural language input corresponding to the first object is of a first type of natural language input (e.g., 1016, 1032, 1050, or 1080). The forward region of the 3D scene is the fourth region of the 3D scene (e.g., 808 or 1092) according to the determination that the natural language input corresponding to the first object (e.g., 1095) is of a second type of natural language input, different from the first type of natural language input.

[0227] In some examples, Method 1100 provides a second audio output (e.g., 1056 or 1086) corresponding to a first object via one or more audio output devices, then receives a user input (e.g., 1058 or 1088) (e.g., speech input, gaze input, and / or gesture input) (e.g., a user input that responds to the second audio output, identifies the first object, and / or identifies the location of the first object), and in response to receiving the user input corresponding to the first object, makes a determination based on the user input (e.g., 1088) corresponding to the first object (e.g., that the image data (e.g., 1076) satisfies a predetermined condition with respect to the first object (e.g., 1078) (e.g., it is determined that the image data depicts the first object with at least a threshold of confidence). The further includes providing a third audio output (e.g., 1090) via one or more audio output devices indicating the result of a second task performed (e.g., by a digital assistant) based on natural language input and a first object, in accordance with the above, and providing a fourth audio output (e.g., 1060) via one or more audio output devices requesting the user to use an external device (e.g., 1062) (e.g., to use an external device to capture an image of the first object) in accordance with a determination based on user input (e.g., 1058) corresponding to the first object (e.g., the image data was not determined to depict the first object with at least a threshold of confidence) that the image data (e.g., 1046) does not satisfy a predetermined condition with respect to the first object (e.g., 1048).

[0228] In some examples, a fourth audio output is provided via one or more audio output devices, prompting the user to use an external device, after which the external device displays a camera user interface (e.g., 1068), and the external device captures an image of a first object (e.g., 1048) while displaying the camera user interface. In some examples, method 1100 includes the external device capturing an image of a first object while displaying the camera user interface (e.g., in response to receiving user input 1072), and then providing a fifth audio output (e.g., 1074) via one or more audio output devices, indicating the result of a third task performed (e.g., by a digital assistant) based on the image of the first object and natural language input (e.g., 1050).

[0229] In some examples, the external device displays the camera user interface in response to a selection (e.g., 1064) of a user interface element (e.g., 1066) displayed by the external device.

[0230] In some examples, method 1100 includes capturing third image data representing a 3D scene via one or more visual imaging sensors before receiving natural language input (e.g., 1032 or 1080), wherein a first audio output is based on a second object (e.g., 1044) detected based on the third image data representing the 3D scene (e.g., as described with respect to Figure 10B), and / or a second audio output is based on a second object (e.g., 1091) detected based on the third image data representing the 3D scene (e.g., as described with respect to Figure 10F).

[0231] The above is written with reference to specific embodiments for illustrative purposes. However, the above exemplary discussion is not intended to be exhaustive or to limit the invention to the exact form disclosed. Many modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the invention and its practical applications, thereby enabling other persons skilled in the art to best use the invention and the various described embodiments with various modifications suitable for specific applications that may be conceived.

[0232] As described above, one aspect of this technology involves collecting and using data available from various sources to facilitate user interaction with a 3D scene. This disclosure suggests that in some cases, the collected data may include personal information data that uniquely identifies a particular person, or personal information data that can be used to contact a particular person or locate them. Such personal information data may include demographic data, location-based data, telephone numbers, email addresses, Twitter IDs, home addresses, data or records relating to a user's health or fitness level (e.g., vital signs measurements, medication information, exercise information), date of birth, or any other identifying or personal information.

[0233] This disclosure acknowledges that such use of personal data in the technology may be for the benefit of the user. For example, personal data can be used to provide voice responses to assist the user. Furthermore, other uses of personal data that may benefit the user are conceivable in this disclosure. For example, health and fitness data can be used to provide insights into the user's overall wellness, or as positive feedback to individuals using the technology to pursue wellness goals.

[0234] This disclosure implies that entities involved in the collection, analysis, disclosure, transmission, storage, or other use of such personal data should adhere to a robust privacy policy and / or privacy practices. Specifically, such entities should implement and consistently use a privacy policy and practices that are generally recognized as meeting or exceeding industry or government requirements for the strict confidentiality of personal data. Such policies should be readily accessible to users and should be updated as data collection and / or use changes. Personal data from users should be collected for the lawful and legitimate use of the entity and should not be shared or sold for any other purpose. Furthermore, such collection / sharing should be carried out only after informing and obtaining the user's consent. In addition, such entities should consider taking all necessary steps to protect and secure access to such personal data and to ensure that others with access to the personal data faithfully adhere to those privacy policies and procedures. Furthermore, such entities may undergo third-party evaluations to demonstrate their compliance with widely accepted privacy policies and practices. Furthermore, policies and practices should be adapted to the specific types of personal data collected and / or accessed, and to applicable laws and standards, including jurisdiction-specific considerations. For example, in the United States, the collection or access to certain health data may be subject to federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA). Health data in other countries, on the other hand, may be subject to other regulations and policies and should be addressed accordingly. Therefore, different privacy practices should be maintained in each country with respect to different types of personal data.

[0235] Notwithstanding the foregoing, the Disclosure also envisions embodiments that allow a user to selectively prevent the use of or access to personal data. That is, the Disclosure envisions that hardware and / or software elements may be provided to prevent or prevent access to such personal data. For example, if a voice response is output for a user, the technology may be configured to allow the user to choose to “opt in” or “opt out” of participating in the collection of personal data during or at any time thereafter when registering for the service. In another example, the user may choose not to provide the personal data on which the voice response is generated. In yet another example, the user may choose to limit the length of time such data is retained. In addition to providing “opt-in” and “opt-out” options, the Disclosure envisions providing notices regarding access to or use of personal data. For example, the user may be notified when downloading an app that will access the user’s personal data, and then reminded again immediately before the app accesses the personal data.

[0236] Furthermore, the intent of this disclosure is that personal data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Risks can be minimized by limiting data collection and deleting data when it is no longer needed. In addition, where applicable in certain health-related applications, data anonymization can be used to protect user privacy. Anonymization can be facilitated by removing certain identifiers (e.g., date of birth), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data across users), and / or by other means, where necessary.

[0237] Therefore, while this disclosure broadly covers the use of personal data to implement one or more different disclosed embodiments, it is also conceivable that these different embodiments could be implemented without requiring access to such personal data. In other words, the different embodiments of the technology would not be rendered inoperable by the absence of all or part of such personal data. For example, a voice response could be generated based on non-personal data or a minimal amount of personal information, such as content requested by a device associated with the user, other non-personal information available for the service, or publicly available information.

Claims

1. It is a method, In a first computer system communicating with one or more image sensors, Acquiring a first image using one or more image sensors, Receiving a user request related to the first image, and, In response to acquiring the first image and receiving the user request, In accordance with the determination that the quality of the first image does not meet the quality standards, the second computer system is prompted to use the second computer system to capture a second image, and In accordance with the determination that the quality of the first image meets the quality criteria, To generate a response to the user request based on the first image, and The method includes providing an output that includes the response to the user request based on the first image.

2. The method according to claim 1, wherein the first computer system is a head-mounted electronic device and the second computer system is a smartphone.

3. The method according to claim 1 or 2, wherein the image sensor of the first computer system has a lower quality metric than the image sensor of the second computer system.

4. In response to detecting the first image and receiving the user request, In accordance with the determination that the user requirement includes a requirement of the first type, the first quality criterion is selected as the quality criterion. In accordance with the determination that the user requirement includes a second type of requirement different from the first type, a second quality standard different from the first quality standard is selected as the quality standard. The method according to any one of claims 1 to 3, further comprising:

5. After causing the second computer system to provide a prompt for capturing a second image using the second computer system, The second computer system is used to detect user input for capturing the second image, To generate a response to the user request based on the second image, To provide an output including the response to the user request based on the second image, The method according to any one of claims 1 to 4, further comprising:

6. The determination of whether the quality of the first image meets the quality criteria is, The method according to any one of claims 1 to 5, comprising providing a prompt to a Large Language Model (LLM) that includes a request for whether the first image is of sufficient quality to complete a task determined from the user request.

7. The determination of whether the quality of the first image meets the quality criteria is, To generate the embedding of the first image, The method according to any one of claims 1 to 5, comprising comparing the embedding of the first image with a set of trained embeddings that represent a high-quality image or a low-quality image.

8. In accordance with the determination that the user request is a request of the first type, the first set of learned embeddings is selected as the set of learned embeddings, In accordance with the determination that the user request is a request of the second type, the second set of learned embeddings is selected as the set of learned embeddings, The method according to claim 7, further comprising:

9. In response to detecting the first image and receiving the user request, In accordance with the determination that the context of the first computer system indicates that the first image does not meet the quality criteria, To discontinue the determination of whether the first image meets the quality criteria, The second computer system is to provide the prompt for capturing the second image using the second computer system, The method according to any one of claims 1 to 8, further comprising:

10. The method according to claim 9, wherein the determination that the context of the first computer system indicates that the first image does not meet the quality criteria includes the determination that the first computer system is moving.

11. The method according to claim 9 or 10, wherein the determination that the context of the first computer system indicates that the first image does not meet the quality criteria includes the determination that the lighting level of the environment of the first computer system is below a lighting threshold.

12. The method according to any one of claims 9 to 11, wherein the determination that the context of the first computer system indicates that the first image does not meet the quality criteria includes the determination that one or more image sensors are obscured.

13. The method according to any one of claims 9 to 12, wherein the determination that the context of the first computer system indicates that the first image does not meet the quality criteria includes the determination that the field of view of one or more image sensors contains text.

14. In accordance with the determination that the quality of the first image does not meet the quality standard, the camera user interface is displayed using a display generation component that communicates with the second computer system. The method according to any one of claims 1 to 13, further comprising:

15. After the camera user interface is displayed using the display generation component that is communicating with the second computer system, Detecting user input for capturing the second image, In response to detecting the user input for capturing the second image, the camera user interface is stopped from being displayed using the display generation component that is communicating with the second computer system, The method according to claim 14, further comprising:

16. The method according to claim 14, wherein the camera user interface is displayed on the lock screen using the display generation component that communicates with the second computer system.

17. The method according to any one of claims 1 to 16, wherein the second computer system is not physically connected to the first computer system.

18. The method according to any one of claims 1 to 16, wherein the second computer system is physically connected to the first computer system by wires, and the second computer system and the first computer system are not located in the same enclosure.

19. A non-temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system communicating with one or more image sensors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 16.

20. A first computer system configured to communicate with one or more image sensors, wherein the first computer system is One or more processors, A first computer system comprising: one or more memories storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 16.

21. A first computer system configured to communicate with one or more image sensors, wherein the first computer system is A first computer system comprising means for performing the method described in any one of claims 1 to 16.

22. A computer program product comprising one or more programs configured to be executed by one or more processors of a first computer system communicating with one or more image sensors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 16.

23. A non-temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system communicating with one or more image sensors, wherein the one or more programs are Acquiring a first image using one or more image sensors, Receiving a user request related to the first image, and, In response to acquiring the first image and receiving the user request, In accordance with the determination that the quality of the first image does not meet the quality standards, the second computer system is prompted to use the second computer system to capture a second image, and In accordance with the determination that the quality of the first image meets the quality criteria, To generate a response to the user request based on the first image, and To provide an output including the response to the user request based on the first image, A non-temporary computer-readable storage medium containing instructions for [a specific purpose].

24. A computer system configured to communicate with one or more image sensors, wherein one or more of the computer systems One or more processors, The system comprises one or more memories storing one or more programs configured to be executed by the one or more processors, Acquiring a first image using one or more image sensors, Receiving a user request related to the first image, and, In response to acquiring the first image and receiving the user request, In accordance with the determination that the quality of the first image does not meet the quality standards, the second computer system is prompted to use the second computer system to capture a second image, and In accordance with the determination that the quality of the first image meets the quality criteria, To generate a response to the user request based on the first image, and To provide an output including the response to the user request based on the first image, A computer system that includes instructions for use.

25. A computer system configured to communicate with one or more image sensors, wherein the computer system is Means for acquiring a first image using one or more image sensors, Means for receiving user requests related to the first image, and In response to acquiring the first image and receiving the user request, In accordance with the determination that the quality of the first image does not meet the quality standards, the second computer system is prompted to use the second computer system to capture a second image, and In accordance with the determination that the quality of the first image meets the quality criteria, To generate a response to the user request based on the first image, and To provide an output including the response to the user request based on the first image, A computer system equipped with means for that purpose.

26. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more image sensors, wherein the one or more programs are Acquiring a first image using one or more image sensors, Receiving a user request related to the first image, and, In response to acquiring the first image and receiving the user request, In accordance with the determination that the quality of the first image does not meet the quality standards, the second computer system is prompted to use the second computer system to capture a second image, and In accordance with the determination that the quality of the first image meets the quality criteria, To generate a response to the user request based on the first image, and To provide an output including the response to the user request based on the first image, A computer program product that includes instructions for use.

27. It is a method, In a computer system that communicates with one or more visual imaging sensors and one or more audio output devices, The computer system receives natural language input corresponding to a first object in the 3D scene while the user's head has a head posture corresponding to the forward-facing region of the 3D scene. Capture image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors, and In response to receiving the natural language input corresponding to the first object in the 3D scene, In accordance with the determination that the visibility metric, which represents the amount depicted by the image data in the forward-facing region of the 3D scene, satisfies the conditions, a first audio output corresponding to the first object is provided via one or more audio output devices, and A method comprising providing a second audio output corresponding to the first object via one or more audio output devices, in accordance with the determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy the condition.

28. In accordance with the determination that the head posture is the first head posture, the forward-facing region of the 3D scene is the first region of the 3D scene. In accordance with the determination that the head posture is a second head posture different from the first head posture, the forward-facing region of the 3D scene is a second region of the 3D scene different from the first region of the 3D scene. The method according to claim 27.

29. Capturing the image data associated with the natural language input via one or more visual imaging sensors is The method according to claim 27 or 28, comprising capturing the image data associated with the natural language input via one or more visual imaging sensors while receiving the natural language input.

30. Capturing the image data associated with the natural language input via one or more visual imaging sensors is The method according to any one of claims 27 to 29, comprising capturing the image data associated with the natural language input via one or more visual imaging sensors in response to receiving the natural language input.

31. The method according to any one of claims 27 to 30, wherein the first audio output corresponding to the first object indicates the result of a first task performed based on the natural language input and the first object.

32. Providing the first audio output corresponding to the first object via one or more audio output devices is The method according to claim 31, further comprising providing an audio output that prompts the user to remove ambiguity among the plurality of detected objects, in accordance with the determination that the forward-facing region of the 3D scene includes the first object.

33. The method according to any one of claims 27 to 32, wherein the second audio output corresponding to the first object includes a request to the user to specify the first object.

34. The computer system includes a first device and a second device, the first device being different from the second device, The natural language input is received while the first device is being worn by the user and while the second device is being worn by the user. The forward-facing region of the 3D scene is determined based on the position of the first device while it is being worn by the user, the orientation of the first device while it is being worn by the user, the position of the second device while it is being worn by the user, and the orientation of the second device while it is being worn by the user. The method according to any one of claims 27 to 33.

35. The one or more visual imaging sensors include a first visual imaging sensor and a second visual imaging sensor different from the first visual imaging sensor. Capturing the image data associated with the natural language input via one or more visual imaging sensors is The first image data is captured via the first visual imaging sensor, This includes capturing a second image data, different from the first image data, via the second visual imaging sensor. The method according to any one of claims 27 to 34.

36. The first image data is captured while the first visual imaging sensor is worn on the first side of the user's head. The second image data is captured while the second visual imaging sensor is worn on the second side of the user's head, and the first side of the user's head is opposite to the second side of the user's head. The method according to claim 35.

37. In accordance with the determination that the first image data was captured while the first visual imaging sensor had a first orientation, and the determination that the second image data was captured while the second visual imaging sensor had a second orientation, the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene has a first value. In accordance with the determination that the first image data was captured while the first visual imaging sensor was in a third orientation different from the first orientation, and the determination that the second image data was captured while the second visual imaging sensor was in a fourth orientation different from the second orientation, the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene has a second value different from the first value. The method according to any one of claims 35 to 36.

38. According to the determination that the first image data and the second image data depict a first amount of the forward-facing region of the 3D scene, the visibility metric has a first value. According to the determination that the first image data and the second image data depict a second amount of the forward-facing region of the 3D scene that is larger than the first amount of the forward-facing region of the 3D scene, the visibility metric has a second value that is larger than the first value. The method according to any one of claims 35 to 37.

39. The method according to any one of claims 27 to 38, wherein the image data includes an image region representing the occlusion of the forward-facing region of the 3D scene, and the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene is based on the image region representing the occlusion of the forward-facing region of the 3D scene.

40. In accordance with the determination that the image region representing the occlusion of the forward-facing region of the 3D scene has a first size, the visibility metric representing the amount depicted by the image data within the forward-facing region of the 3D scene has a third value. In accordance with the determination that the image region representing the occlusion of the forward-facing region of the 3D scene has a second size that is larger than the first size, the visibility metric representing the amount depicted by the image data within the forward-facing region of the 3D scene has a fourth value that is smaller than the third value. The method according to claim 39.

41. The method according to any one of claims 27 to 40, wherein the visibility metric, which represents the amount depicted by the image data in the forward-facing region of the 3D scene, is based on the amount of overlap between the unobstructed region of the 3D scene depicted by the image data and the forward-facing region of the 3D scene.

42. The method according to any one of claims 27 to 41, wherein the forward-facing region of the 3D scene has predetermined dimensions.

43. The method according to any one of claims 27 to 42, wherein the forward-facing region of the 3D scene is at least a predetermined distance from the user's head.

44. The handheld object area is determined, The aforementioned handheld object area is where each user holds each object in their hand while issuing queries about that object. The forward-facing region of the 3D scene is determined based on the handheld object region. The method according to any one of claims 27 to 43.

45. In accordance with the determination that the natural language input corresponding to the first object is a first type of natural language input, the forward-facing region of the 3D scene is a third region of the 3D scene. In accordance with the determination that the natural language input corresponding to the first object is a second type of natural language input different from the first type of natural language input, the forward-facing region of the 3D scene is a fourth region of the 3D scene different from the third region of the 3D scene. The method according to any one of claims 27 to 44.

46. After providing the second audio output corresponding to the first object via one or more audio output devices, Receiving user input corresponding to the first object, and In response to receiving the user input corresponding to the first object, To provide, via one or more audio output devices, a third audio output indicating the result of a second task performed based on the natural language input and the first object, in accordance with a determination based on the user input corresponding to the first object that the image data satisfies predetermined conditions with respect to the first object, and To provide a fourth audio output via one or more audio output devices that requests the user to use an external device, in accordance with a determination based on the user input corresponding to the first object that the image data does not satisfy the predetermined conditions with respect to the first object. The method according to any one of claims 27 to 45, further comprising:

47. After providing the fourth audio output via one or more audio output devices, which prompts the user to use the external device, the external device displays a camera user interface, and while the external device is displaying the camera user interface, the method captures an image of the first object, The method according to claim 46, further comprising capturing the image of the first object while the external device is displaying the camera user interface, and then providing a fifth audio output via one or more audio output devices indicating the result of a third task performed based on the image of the first object and the natural language input.

48. The method according to claim 47, wherein the external device displays the camera user interface in response to the selection of a user interface element displayed by the external device.

49. The method further includes capturing a third image data representing the 3D scene via one or more visual imaging sensors before receiving the natural language input, The first audio output is based on a second object detected based on the third image data representing the 3D scene, and / or The second audio output is based on the second object detected based on the third image data representing the 3D scene, The method according to any one of claims 27 to 48.

50. A non-temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs include instructions for performing the method according to any one of claims 27 to 49.

51. A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, One or more processors, A computer system comprising one or more memories storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 27 to 49.

52. A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, A computer system comprising means for performing the method described in any one of claims 27 to 49.

53. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs include instructions for performing the method according to any one of claims 27 to 49.

54. A non-temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs are The computer system receives natural language input corresponding to a first object in the 3D scene while the user's head has a head posture corresponding to the forward-facing region of the 3D scene. Capture image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors, and In response to receiving the natural language input corresponding to the first object in the 3D scene, In accordance with the determination that the visibility metric, which represents the amount depicted by the image data in the forward-facing region of the 3D scene, satisfies the conditions, a first audio output corresponding to the first object is provided via one or more audio output devices, and In accordance with the determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy the conditions, a second audio output corresponding to the first object is provided via one or more audio output devices. A non-temporary computer-readable storage medium containing instructions for [a specific purpose].

55. A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, One or more processors, The system comprises one or more memories storing one or more programs configured to be executed by the one or more processors, The computer system receives natural language input corresponding to a first object in the 3D scene while the user's head has a head posture corresponding to the forward-facing region of the 3D scene. Capture image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors, and In response to receiving the natural language input corresponding to the first object in the 3D scene, In accordance with the determination that the visibility metric, which represents the amount depicted by the image data in the forward-facing region of the 3D scene, satisfies the conditions, a first audio output corresponding to the first object is provided via one or more audio output devices, and In accordance with the determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy the conditions, a second audio output corresponding to the first object is provided via one or more audio output devices. A computer system that includes instructions for use.

56. A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, Means for receiving natural language input corresponding to a first object in a 3D scene while the user's head of the computer system has a head posture corresponding to a forward-facing region of the 3D scene, Means for capturing image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors, In response to receiving the natural language input corresponding to the first object in the 3D scene, In accordance with the determination that the visibility metric, which represents the amount depicted by the image data in the forward-facing region of the 3D scene, satisfies the conditions, a first audio output corresponding to the first object is provided via one or more audio output devices, and In accordance with the determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy the conditions, a second audio output corresponding to the first object is provided via one or more audio output devices. A computer system equipped with means for that purpose.

57. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs are The computer system receives natural language input corresponding to a first object in the 3D scene while the user's head has a head posture corresponding to the forward-facing region of the 3D scene. Capture image data associated with the natural language input corresponding to the first object in the 3D scene via one or more visual imaging sensors, and In response to receiving the natural language input corresponding to the first object in the 3D scene, In accordance with the determination that the visibility metric, which represents the amount depicted by the image data in the forward-facing region of the 3D scene, satisfies the conditions, a first audio output corresponding to the first object is provided via one or more audio output devices, and In accordance with the determination that the visibility metric representing the amount depicted by the image data in the forward-facing region of the 3D scene does not satisfy the conditions, a second audio output corresponding to the first object is provided via one or more audio output devices. A computer program product that includes instructions for use.