Responding to user requests related to images
Patent Information
- Application Number
- KR1020260034736
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2026-01-26
- Filing Date
- 2026-02-25
- Publication Date
- 2026-09-04
Smart Images

Figure PAT00007_ABST
Abstract
Description
Technology Field
[0001] Cross-reference of related applications
[0002] This application relates to the following co-pending provisional application: U.S. Patent Application No. 63 / 765,276 filed on February 28, 2025, under the title of the invention "MULTIDEVICE CAMERA SELECTION," the contents of which are incorporated by reference in their entirety.
[0003] Technology field
[0004] The present disclosure generally relates to responding to user requests related to images. Background Technology
[0005] The development of computer systems for interacting with and / or providing three-dimensional (3D) scenes has expanded significantly in recent years. Exemplary three-dimensional scenes (e.g., environments) include physical scenes and extended reality (XR) scenes.
[0006] Exemplary methods are disclosed herein. An exemplary method comprises: a first computer system communicating with one or more image sensors, using one or more image sensors to acquire a first image; receiving a user request associated with the first image; and in response to acquiring the first image and receiving the user request: a second computer system providing a prompt to capture a second image to the second computer system based on a determination that the quality of the first image does not satisfy a quality criterion; and a response to the user request based on the first image based on a determination that the quality of the first image satisfies a quality criterion; and providing an output including the response to the user request based on the first image.
[0007] Exemplary non-transient computer-readable storage media are disclosed herein. An exemplary non-transient computer-readable storage medium stores one or more programs. One or more programs are configured to be executed by one or more processors of a first computer system communicating with one or more image sensors. One or more programs include instructions for acquiring a first image using one or more image sensors; receiving a user request related to the first image; and in response to acquiring the first image and receiving the user request: providing a prompt to a second computer system to capture a second image to the second computer system upon a determination that the quality of the first image does not satisfy a quality criterion; generating a response to the user request based on the first image upon a determination that the quality of the first image satisfies a quality criterion; and providing an output including the response to the user request based on the first image.
[0008] Exemplary computer systems are disclosed herein. An exemplary first computer system is configured to communicate with one or more image sensors. The first computer system comprises one or more processors; and a memory for storing one or more programs configured to be executed by the one or more processors, wherein the one or more programs use one or more image sensors to acquire a first image; receive a user request related to the first image; and in response to acquiring the first image and receiving the user request: provide a prompt to the second computer system to capture a second image according to a determination that the quality of the first image does not satisfy a quality criterion; generate a response to the user request based on the first image according to a determination that the quality of the first image satisfies a quality criterion; and provide an output including the response to the user request based on the first image.
[0009] An exemplary first computer system is configured to communicate with one or more image sensors. The first computer system includes means for acquiring a first image using one or more image sensors; means for receiving a user request related to the first image; and means for, in response to acquiring the first image and receiving the user request: providing a prompt to the second computer system to capture a second image based on a determination that the quality of the first image does not satisfy a quality criterion; generating a response to the user request based on the first image based on a determination that the quality of the first image satisfies a quality criterion; and providing an output including the response to the user request based on the first image.
[0010] An exemplary computer program product comprises one or more programs configured to be executed by one or more processors of a first computer system communicating with one or more image sensors. The one or more programs include instructions for acquiring a first image using one or more image sensors; receiving a user request related to the first image; and in response to acquiring the first image and receiving the user request: providing a prompt to a second computer system to capture a second image to the second computer system upon a determination that the quality of the first image does not satisfy a quality criterion; generating a response to the user request based on the first image upon a determination that the quality of the first image satisfies a quality criterion; and providing an output including the response to the user request based on the first image.
[0011] Having the second computer system provide a prompt to capture a second image when the first image does not meet quality standards allows for more fluid interaction between the user and multiple devices of the system. Specifically, the second computer system may indicate to the user that it can automatically take another photo that meets quality standards while the second computer system is in the process of taking a photo. This provides the user with more information and reduces the number of inputs the user must provide to capture an image in order to complete the requested task. In this way, user-device interaction becomes more efficient and accurate (e.g., by reducing the number of inputs required to capture a suitable image and by providing the user with additional information regarding device capabilities), which consequently reduces power consumption and improves the battery life of the devices by enabling the user to use the devices more quickly and efficiently.
[0012] Exemplary methods are disclosed herein. An exemplary method comprises: receiving a natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing region of a 3D scene; capturing image data associated with the natural language input corresponding to the first object in the 3D scene through one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: providing a first audio output corresponding to the first object through one or more audio output devices according to a determination that a visibility metric representing the amount of the forward-facing region of the 3D scene described by the image data satisfies a condition; and providing a second audio output corresponding to the first object through one or more audio output devices according to a determination that a visibility metric representing the amount of the forward-facing region of the 3D scene described by the image data does not satisfy a condition.
[0013] Exemplary non-transient computer-readable storage media are disclosed herein. An exemplary non-transient computer-readable storage medium stores one or more programs. One or more programs are configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices. One or more programs receive natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing area of a 3D scene; capture image data associated with the natural language input corresponding to the first object in the 3D scene through one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: provide a first audio output corresponding to the first object through one or more audio output devices according to a determination that a visibility metric representing the amount of the forward-facing area of the 3D scene described by the image data satisfies a condition; Instructions for providing a second audio output corresponding to a first object through one or more audio output devices, based on a determination that a visibility metric representing the amount of a forward-facing area of a 3D scene described by image data does not satisfy the condition.
[0014] Exemplary computer systems are disclosed herein. An exemplary computer system is configured to communicate with one or more visual imaging sensors and one or more audio output devices. The computer system comprises one or more processors; and a memory for storing one or more programs configured to be executed by the one or more processors, wherein the one or more programs receive a natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing area of a 3D scene; capture image data associated with the natural language input corresponding to the first object in the 3D scene through one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: provide a first audio output corresponding to the first object through one or more audio output devices, according to a determination that a visibility metric representing the amount of the forward-facing area of the 3D scene described by the image data satisfies a condition; Instructions for providing a second audio output corresponding to a first object through one or more audio output devices, based on a determination that a visibility metric representing the amount of a forward-facing area of a 3D scene described by image data does not satisfy the condition.
[0015] An exemplary computer system is configured to communicate with one or more visual imaging sensors and one or more audio output devices. The computer system comprises: means for receiving natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing area of a 3D scene; means for capturing image data associated with the natural language input corresponding to the first object in the 3D scene through one or more visual imaging sensors; and means for, in response to receiving the natural language input corresponding to the first object in the 3D scene: providing a first audio output corresponding to the first object through one or more audio output devices according to a determination that a visibility metric representing the amount of a forward-facing area of the 3D scene described by the image data satisfies a condition; and providing a second audio output corresponding to the first object through one or more audio output devices according to a determination that a visibility metric representing the amount of a forward-facing area of the 3D scene described by the image data does not satisfy a condition.
[0016] An exemplary computer program product comprises one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices. The one or more programs include instructions for receiving natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing area of a 3D scene; capturing image data associated with the natural language input corresponding to the first object in the 3D scene through one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: providing a first audio output corresponding to the first object through one or more audio output devices according to a determination that a visibility metric representing the amount of a forward-facing area of the 3D scene described by the image data satisfies a condition; and providing a second audio output corresponding to the first object through one or more audio output devices according to a determination that a visibility metric representing the amount of a forward-facing area of the 3D scene described by the image data does not satisfy a condition.
[0017] Providing audio outputs based on whether a visibility metric satisfies a condition allows the computer system to satisfy user requests regarding objects present in a 3D scene more accurately and efficiently. For example, if the visibility metric does not satisfy the condition, the captured image data may not depict the object related to the user request, and therefore, the computer system may not be able to satisfy the user request based on the captured image data. As described herein, the computer system may therefore provide one or more audio outputs that prompt the user to identify the object related to the user request and / or prompt the user to capture an image of the object with a different device, thereby allowing the computer system to accurately satisfy the user request based on new image data that depicts the relevant object. As another example, if the visibility metric satisfies the condition, the captured image data may depict the object related to the user request, and therefore, the computer system may satisfy the user request based on the captured image data. The computer system may then provide an audio output that satisfies the user request. In this way, the user-device interface becomes more accurate and efficient (e.g., by allowing devices to respond accurately to user requests regarding objects in a 3D scene, by preventing electronic devices from providing inaccurate responses to user requests regarding objects in a 3D scene, by reducing the number of user inputs required for electronic devices to satisfy user requests, and by reducing the number of inputs otherwise required to undo and / or cancel the results of inaccurately interpreted user requests), which consequently reduces power consumption and improves the battery life of the devices by enabling users to use the devices more quickly and efficiently.
[0018] In some examples, the computer system is a desktop computer having an associated display. In some examples, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device, such as a smartphone). In some examples, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or a head-mounted device (HMD)). In some examples, the computer system has a touchpad. In some examples, the computer system has one or more cameras. In some examples, the computer system has a display generating component (e.g., a display device, such as a head-mounted display, a display, a projector, a touch-sensitive display (also known as a "touch screen" or "touch screen display"), or, for example, another device or component that presents visual content visible to the user on or within the display generating component itself or generated from the display generating component and located elsewhere). In some examples, the computer system does not have a display generating component and does not provide visual content to the user. In some examples, the computer system has a touch-sensitive display (also known as a “touch screen” or “touch screen display”). In some examples, the computer system has one or more eye-tracking components. In some examples, the computer system has one or more hand-tracking components. In some examples, the computer system has one or more output devices, and the output devices include one or more haptic output generators and / or one or more audio output devices. In some examples, the computer system has one or more processors, memory, and one or more sets of modules, programs, or instructions stored in memory to perform the many functions described herein.In some examples, the user interacts with the computer system through stylus and / or finger contacts and gestures on a touch-sensitive surface, movements of the user's eyes and hands in the user's body or space as captured by cameras and other motion sensors, and / or voice inputs as captured by one or more audio input devices. Executable instructions for performing these functions are, optionally, contained in a transient and / or non-transient computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0019] It should be noted that the various examples described above may be combined with any other examples described herein. The features and benefits described herein are not all included, and in particular, many additional features and benefits will be apparent to those skilled in the art from the drawings, the specification, and the claims. Furthermore, it should be noted that the expressions used herein have been chosen primarily for readability and educational purposes and may not have been chosen to describe or limit the subject matter of the claims of the invention. Brief explanation of the drawing
[0020] For a better understanding of the various described examples, specific details for carrying out the invention below should be referred to in conjunction with the following drawings, in which similar reference numerals refer to corresponding parts throughout the drawings. FIG. 1 is a block diagram illustrating the operating environment of a computer system for interacting with three-dimensional (3D) scenes according to some examples. FIG. 2 is a block diagram of a user-facing component of a computer system according to some examples. FIG. 3a is a block diagram of a controller of a computer system according to some examples. FIG. 3b is a block diagram of an image evaluation unit of a computer system according to some examples. Figure 4 illustrates an architecture for a foundation model according to some examples. FIGS. 5a through 5e illustrate capturing images in a multi-device system according to some examples. FIG. 6 is a flowchart of a method for capturing images in a multi-device system according to some examples. FIG. 7 illustrates a forward-facing head pose determined based on the pose of the first device and the pose of the second device according to some examples. FIGS. 8a and FIGS. 8b illustrate regions of interest determined based on a forward-facing head pose according to some examples. FIGS. 9a and 9b illustrate the determination of visibility metrics for a region of interest according to some examples. FIGS. 10a through 10g illustrate a device performing various actions in response to receiving natural language input and according to a determined visibility metric, according to some examples. FIG. 11 is a flowchart of a method for providing audio outputs in response to natural language input, according to some examples. FIGS. 1 through 4 provide descriptions of exemplary computer systems and techniques for interacting with three-dimensional scenes. FIGS. 5a through 5e illustrate capturing images in a multi-device system. FIG. 6 is a flowchart of a method for capturing images in a multi-device system. FIGS. 5a through 5e are used to explain the method of FIG. 6. FIG. 7 illustrates a forward-facing head pose determined based on the pose of a first device and the pose of a second device. FIGS. 8a and 8b illustrate regions of interest determined based on the forward-facing head pose. FIGS. 9a and 9b illustrate the determination of a visibility metric for a region of interest. FIGS. 10a through 10g illustrate the device performing various actions according to the determined visibility metric and in response to receiving natural language input. FIG. 11 is a flowchart of a method for providing audio outputs in response to natural language input. FIGS. 7, FIGS. 8a and 8b, FIGS. 9a and 9b, and FIGS. 10a to 10g are used to explain the method of FIG. 11. Furthermore, in the methods described herein in which one or more steps are conditioned upon one or more conditions being satisfied, it should be understood that the described method may be repeated in a number of iterations, so that during the iterations, all conditions conditioned upon by the steps of the method may be satisfied in different iterations of the method. For example, if the method requires performing a first step when a condition is satisfied and a second step when a condition is not satisfied, those skilled in the art will recognize that the claimed steps are repeated in no particular order until the condition is satisfied and then not satisfied. Accordingly, a method described in which one or more steps are conditioned upon one or more conditions being satisfied may be rewritten as a method repeated until each of the conditions described in the method is satisfied. However, this is not required in claims of a system or computer-readable medium in which the system or computer-readable medium includes instructions for performing contingent actions based on the satisfaction of one or more corresponding conditions, and accordingly, can determine whether a contingency is satisfied or not without explicitly repeating the steps of the method until all conditions conditioned upon the steps of the method are satisfied. Those skilled in the art will also understand that, similar to a method having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all conditional steps have been performed. FIG. 1 is a block diagram illustrating an operating environment of a computer system (101) for interacting with three-dimensional scenes according to some examples. In FIG. 1, a user interacts with a three-dimensional scene (105) through an operating environment (100) including the computer system (101). In some examples, the computer system (101) includes a controller (110) (e.g., processors of a portable electronic device or a remote server), a user-facing component (120), one or more input devices (125) (e.g., an eye-tracking device (130), a hand-tracking device (140), and / or other input devices (150)), one or more output devices (155) (e.g., speakers (160), haptic output generators (170), and other output devices (180)), one or more sensors (190) (e.g., image sensors, light sensors, depth sensors, haptic sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, speed sensors, audio sensors, etc.), and one or more peripheral devices (195) (e.g., home appliances, wearable devices, etc.). In some examples, one or more of the input devices (125), output devices (155), sensors (190), and peripheral devices (195) are integrated with the user-facing component (120) (e.g., in a head-mounted device or a handheld device). Although relevant features of the operating environment (100) are illustrated in FIG. 1, those skilled in the art will recognize that various other features have not been illustrated in this disclosure for brevity and to avoid obscuring more relevant aspects of the examples disclosed herein. Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with three-dimensional scenes. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to rest over the human eye (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated within the head-mounted system to provide audio output. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accommodate an external opaque display (e.g., a smartphone). Alternatively, the head-mounted system may be configured to operate without displaying content, for example, so that the head-mounted system provides output to the user through tactile and / or auditory means. The head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment, and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display rather than an opaque display. The transparent or translucent display may have a medium through which light representing images is directed toward the human eyes.The display may utilize digital light projection, organic light-emitting diodes (OLEDs), LEDs, uLEDs, liquid crystal on silicon (LCoS), laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one example, a transparent or translucent display may be configured to optionally become opaque. Projection-based systems may utilize retinal projection technology to project graphic images onto the human retina. Projection systems may also be configured to project virtual objects into a physical environment, for example, as holograms or onto a physical surface. In some examples, the user-facing component (120) is configured to provide visual components of a three-dimensional scene. In some examples, the user-facing component (120) includes a suitable combination of software, firmware, and / or hardware. The user-facing component (120) is described in more detail below in relation to FIG. 2. In some examples, the functions of the controller (110) are provided by and / or combined with the user-facing component (120). In some examples, the user-facing component (120) provides an extended reality (XR) experience to the user while the user is virtually and / or physically present within the scene (105). In some examples, the user-facing component (120) is worn on a part of the user's body (e.g., his / her head, his / her hand, etc.). In some examples, the user-facing component (120) includes one or more XR displays provided to display XR content. In some examples, the user-facing component (120) surrounds the user's field of view. In some examples, the user-facing component (120) is a handheld device (e.g., a smartphone or tablet) configured to present XR content, and the user holds the device having a display oriented toward the user's field of view and a camera oriented toward the scene (105). In some examples, the handheld device is optionally placed within an enclosure worn on the user's head. In some examples, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some examples, the user-facing component (120) is an XR chamber, enclosure, or room configured to present XR content, in which the user does not wear or hold the user-facing component (120). Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., a head-mounted device (HMD) or other wearable computing device). For example, a user interface showing interactions with XR content triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be similarly implemented using an HMD in which interactions occur in the space in front of the HMD and responses of the XR content are displayed through the HMD.Similarly, a user interface showing interactions with XR content triggered based on the movement of a handheld or tripod-mounted device relative to a physical environment (e.g., scene (105) or part of the user's body (e.g., user's eye(s), head or hand)) can be similarly implemented using an HMD in which the movement is caused by the movement of the HMD relative to a physical environment (e.g., scene (105) or part of the user's body (e.g., user's eye(s), head or hand)). FIG. 2 is a block diagram of a user-facing component (120) according to some examples. While certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for brevity and to avoid obscuring more relevant aspects of the examples disclosed herein. Furthermore, FIG. 2 is intended more as a functional description of various features that may exist in a particular embodiment, in contrast to the structural schematic diagrams of the examples described herein. As will be recognized by those skilled in the art, individually illustrated components may be combined and some components may be separated. For example, in various examples, some functional modules individually illustrated in FIG. 2 may be implemented as a single module, and various functions of a single functional block may be implemented by one or more functional blocks. The actual number of modules, the division of specific functions, and how features are allocated among them will vary from embodiment to embodiment and, in some examples, depend in part on the specific combination of hardware, software, and / or firmware selected for the particular embodiment. In some examples, the user-facing component (120) (e.g., HMD) comprises one or more processing units (202) (e.g., microprocessors, ASICs (application-specific integrated-circuits), FPGAs (field-programmable gate arrays), GPUs (graphics processing units), CPUs (central processing units), processing cores, etc.), one or more input / output (I / O) devices and sensors (206), one or more communication interfaces (208) (e.g., USB (universal serial bus), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM (global system for mobile communications), CDMA (code division multiple access), TDMA (time division multiple access), GPS (global positioning system), infrared (IR), Bluetooth, Zigbee, and / or similar types of interfaces), one It includes the above programming (e.g., I / O) interfaces (210), one or more XR displays (212), one or more optional internal and / or external opposing image sensors (214), memory (220), and one or more communication buses (204) for interconnecting these and various other components. In some examples, one or more communication buses (204) include circuitry that interconnects and controls communications between system components. In some examples, one or more I / O devices and sensors (206) include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more biometric sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.). In some examples, one or more XR displays (212) are configured to provide an XR experience to a user. In some examples, one or more XR displays (212) correspond to holographic, DLP (digital light processing), LCD (liquid-crystal display), LCoS, OLET (organic light-emitting field-effect transistor), OLED, SED (surface-conduction electron-emitter display), FED (field-emission display), QD-LED (quantum-dot light-emitting diode), MEMS (micro-electro-mechanical system), and / or similar display types. In some examples, one or more XR displays (212) correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, a user-facing component (120) (e.g., HMD) includes a single XR display. In other examples, the user-facing component (120) includes an XR display for each eye of the user. In some examples, one or more XR displays (212) can present XR content. In some examples, one or more XR displays (212) are omitted from the user-facing component (120). For example, the user-facing component (120) does not include any component configured to display content (or does not include any component configured to display XR content), and the user-facing component (120) provides output through audio and / or haptic output types. In some examples, one or more image sensors (214) are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some examples, one or more image sensors (214) are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as a hand-tracking camera). In some examples, one or more image sensors (214) are configured to face forward to acquire image data corresponding to the scene the user would have seen if the user-facing component (120) (e.g., HMD) had not existed (and may be referred to as a scene camera). One or more optional image sensors (214) may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor CMOS image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, etc. Memory (220) includes high-speed random access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDR RAM), or other random access solid-state memory devices. In some examples, memory (220) includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory (220) optionally includes one or more storage devices located remotely from one or more processing units (202). Memory (220) includes a non-transient computer-readable storage medium. In some examples, memory (220) or the non-transient computer-readable storage medium of memory (220) stores the following programs, modules, and data structures, or a subset thereof including an optional operating system (230) and an XR experience module (240). The operating system (230) includes instructions for processing various basic system services and for performing hardware-dependent tasks. In some examples, the XR experience module (240) is configured to present XR content to a user through one or more XR displays (212) or one or more speakers. To this end, in various examples, the XR experience module (240) includes a data acquisition unit (242), an XR presentation unit (244), an XR map generation unit (246), and a data transmission unit (248). In some examples, the data acquisition unit (242) is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller (110) of FIG. 1. To this end, in various examples, the data acquisition unit (242) includes instructions and / or logic therefor, and heuristics and metadata therefor. In some examples, the XR presentation module (244) is configured to present XR content through one or more XR displays (212) or one or more speakers. To this end, in various examples, the XR presentation unit (244) includes commands and / or logic therefor, and heuristics and metadata therefor. In some examples, the XR map generation unit (246) is configured to generate an XR map (e.g., a map of a physical environment where computer-generated objects may be placed or a 3D map of an extended reality scene) based on media content data. To this end, in various examples, the XR map generation unit (246) includes instructions and / or logic therefor, and heuristics and metadata therefor. In some examples, the data transmission unit (248) is configured to transmit data (e.g., presentation data, location data, sensor data, etc.) to at least one of the controller (110) and optionally input devices (125), output devices (155), sensors (190), and / or peripheral devices (195). To this end, in various examples, the data transmission unit (248) includes commands and / or logic therefor, and heuristics and metadata therefor. Although the data acquisition unit (242), XR presentation unit (244), XR map generation unit (246), and data transmission unit (248) are depicted as existing on a single device (e.g., user-facing component (120) of FIG. 1), in other examples, any combination of the data acquisition unit (242), XR presentation unit (244), XR map generation unit (246), and data transmission unit (248) may exist on separate computing devices. Returning to Fig. 1, the controller (110) is configured to manage and coordinate the user's experience in relation to a three-dimensional scene. In some examples, the controller (110) includes a suitable combination of software, firmware, and / or hardware. The controller (110) is described in more detail below in relation to Fig. 3a. In some examples, the controller (110) is a computing device that is local or remote to the scene (105) (e.g., physical environment). For example, the controller (110) is a local server located within the scene (105). In other examples, the controller (110) is a remote server located outside the scene (105) (e.g., cloud server, central server, etc.). In some examples, the controller (110) is coupled to communicate with component(s) of a computer system (101) (e.g., output devices (155) and / or user-facing component (120)) configured to provide output to a user via one or more wired or wireless communication channels (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In some examples, the controller (110) is contained within an enclosure (e.g., physical housing) of a component(s) (e.g., user-facing component (120)) of a computer system (101) configured to provide output to a user, or shares the same physical enclosure or support structure as the component(s) of the computer system (101) configured to provide output to a user. In some examples, the various components and functions of the controller (110) described below in relation to FIGS. 3a and 3b, FIGS. 4, FIGS. 5a through 5e, FIGS. 6, FIG. 7, FIGS. 8a and 8b, FIGS. 9a and 9b, FIGS. 10a through 10g, and FIG. 11 are distributed across multiple devices. For example, a first set of components of the controller (110) (and their associated functions) is implemented on a server system that is remote from the scene (105), while a second set of components of the controller (110) (and their associated functions) is local to the scene (105). For example, the second set of components is implemented within a portable electronic device (e.g., a wearable device such as an HMD) present within the scene (105). It will be recognized that the specific manner in which the various components and functions of the controller (110) are distributed across various devices may vary based on different embodiments of the examples described herein. FIG. 3a is a block diagram of a controller (110) according to some examples. While certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for brevity and to avoid obscuring more relevant aspects of the examples disclosed herein. Furthermore, FIG. 3a is intended more as a functional description of various features that may exist in a particular embodiment, in contrast to the structural schematic diagrams of the examples described herein. As will be recognized by those skilled in the art, individually illustrated components may be combined and some components may be separated. For example, in various examples, some functional modules individually illustrated in FIG. 3a may be implemented as a single module, and various functions of a single functional block may be implemented by one or more functional blocks. The actual number of modules, the division of specific functions, and how features are assigned among them will vary from embodiment to embodiment and, in some examples, depend in part on the specific combination of hardware, software, and / or firmware selected for the particular embodiment. In some examples, the controller (110) includes one or more processing units (302) (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices (306), one or more communication interfaces (308) (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared (IR), Bluetooth, Zigbee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces (310), memory (320), and one or more communication buses (304) for interconnecting these and various other components. In some examples, one or more communication buses (304) include circuitry that interconnects and controls communications between system components. In some examples, one or more I / O devices (306) include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc. Memory (320) includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some examples, memory (320) includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory (320) optionally includes one or more storage devices located remotely from one or more processing units (302). Memory (320) includes a non-transient computer-readable storage medium. In some examples, memory (320) or the non-transient computer-readable storage medium of memory (320) stores the following programs, modules and data structures, or a subset thereof including an optional operating system (330) and a three-dimensional (3D) experience module (340). The operating system (330) includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some examples, the 3D experience module (340) is configured to manage and coordinate the user experience provided by the computer system (101) in relation to the 3D scene. For example, the 3D experience module (340) is configured to acquire data corresponding to the 3D scene (e.g., data generated by the computer system (101) and / or data from the data acquisition unit (341) described below) so that the computer system (101) can perform actions on the user based on the data (e.g., providing suggestions, displaying content, etc.). To this end, in various examples, the 3D experience module (340) includes a data acquisition unit (341), a tracking unit (342), a coordination unit (346), a data transmission unit (348), a digital assistant (DA) unit (350), an image evaluation unit (370), an area tracking unit (380), and a visibility analysis unit (390). In some examples, the data acquisition unit (341) is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of the user-facing component (120), input devices (125), output devices (155), sensors (190), and peripheral devices (195). To this end, in various examples, the data acquisition unit (341) includes commands and / or logic therefor, and heuristics and metadata therefor. In some examples, the tracking unit (342) is configured to map the scene (105) and to track the position / location of the user (and / or the portable device the user is holding or wearing). To this end, in various examples, the tracking unit (342) includes commands and / or logic therefor, and heuristics and metadata therefor. In some examples, the tracking unit (342) includes an eye tracking unit (343). The eye tracking unit (343) includes commands and / or logic for tracking the position and movement of the user's gaze (or more broadly, the user's eyes, face, or head) using data obtained from the eye tracking device (130). In some examples, the eye tracking unit (343) tracks the position and movement of the user's gaze toward the physical environment, toward the user (e.g., the user's hands, face, or head), toward the device worn or held by the user, and / or toward content displayed by the user-facing component (120). The eye tracking device (130) is controlled by the eye tracking unit (343) and includes various hardware and / or software components configured to perform eye tracking techniques. For example, the eye tracking device (130) includes at least one eye tracking camera (e.g., infrared (IR) or near-IR (NIR) cameras) and light sources (e.g., IR or NIR light sources such as an array or ring of LEDs) that emit light (e.g., IR or NIR light) toward the user's eyes. The eye tracking cameras may be oriented toward the user's eyes to receive the reflected IR or NIR light from the light sources directly from the eyes, or alternatively, may be oriented toward mirrors that reflect the IR or NIR light from the eyes toward the eye tracking cameras. The eye tracking device (130) optionally captures images of the user's eyes (e.g., as a video stream captured at a frame rate of 60 to 120 frames per second), analyzes the images to generate eye tracking information, and communicates the eye tracking information to the eye tracking unit (343). In some examples, the user's two eyes are tracked individually by their respective eye tracking cameras and light sources. In some examples, only one of the user's eyes is tracked by their respective eye tracking cameras and light sources. In some examples, the tracking unit (342) includes a hand tracking unit (344). The hand tracking unit (344) includes commands and / or logic for tracking the position of one or more parts of the user's hands and / or the motion of one or more parts of the user's hands using hand tracking data obtained from the hand tracking device (140). The hand tracking unit (344) tracks the position and / or motion with respect to a coordinate system defined for the scene (105), for the user (e.g., the user's head, face, or eyes), for the device worn or held by the user, for the content displayed by the user-facing component (120), and / or the user's hands. In some examples, the hand tracking unit (344) analyzes hand tracking data to identify hand gestures (e.g., pointing gesture, pinching gesture, clenching gesture and / or grabbing gesture) and / or content corresponding to the hand gesture (e.g., physical content or virtual content), e.g., content selected by the hand gesture. In some examples, the hand gesture is an air gesture.An air gesture is a gesture detected without the user touching (or independently thereof) an input element that is part of a device (e.g., computer system (101), one or more input devices (125), hand tracking device (140), device (500), device (1000), and / or device (1062)), and includes motion of the user's body relative to an absolute reference (e.g., angle of the user's arm relative to the ground or distance of the user's hand relative to the ground), motion of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one user's hand relative to the user's other hand, and / or movement of the user's finger relative to another finger or part of the user's hand), and / or absolute motion of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shaking gesture including a predetermined rotational speed or amount of rotation of a part of the user's body) through air (e.g., head, one or more arms, one or more hands, one or more It is based on the detected motion of the fingers and / or one or more legs. A hand tracking device (140) is controlled by a hand tracking unit (344) and includes various hardware and / or software components configured to perform hand tracking and hand gesture recognition techniques. For example, the hand tracking device (140) includes one or more image sensors (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that capture three-dimensional information (e.g., depth map) representing a human user's hand. One or more image sensors capture hand images with a resolution sufficient to distinguish fingers and their respective positions. In some examples, one or more image sensors project a pattern of spots onto an environment containing the hand and capture an image of the projected pattern. In some examples, one or more image sensors capture a temporal sequence of hand tracking data (e.g., captured images of a projected pattern and / or captured three-dimensional information), and the hand tracking device (140) communicates to a hand tracking unit (344) for further analysis of the temporal sequence of hand tracking data, e.g., to identify hand gestures, hand poses, and / or hand movements. In some examples, the hand tracking device (140) includes one or more hardware input devices configured to be worn and / or held (or otherwise attached to the hands) by each of the user's one or more hands. In such examples, the hand tracking unit (344) tracks the position, pose, and / or motion of the user's hand based on tracking the position, pose, and / or motion of each hardware input device. The hand tracking unit (344) tracks the position, pose, and / or motion of each hardware input device optically (e.g., through one or more image sensors) and / or based on data obtained from sensor(s) included within the hardware input device (e.g., accelerometer(s), magnetometer(s), gyroscope(s), inertial measurement unit(s), etc.). In some examples, the hardware input device includes one or more physical controls (e.g., buttons(s), touch-sensitive surfaces(s), pressure-sensitive surfaces(s), knob(s), joystick(s), etc.). In some examples, instead of performing a specific function in response to detecting each type of hand gesture, or in addition to this, the computer system (101) similarly performs a specific function in response to user input selecting each physical control of the hardware input device. For example, the computer system (101) interprets a pinching hand gesture input as a selection of an in-focus element and / or a selection of a physical button of the hardware device as a selection of an in-focus element. In some examples, the coordination unit (346) is configured to manage and coordinate the experience provided to the user through the user-facing component (120), one or more output devices (155), and / or one or more peripheral devices (195). To this end, in various examples, the coordination unit (346) includes commands and / or logic therefor, and heuristics and metadata therefor. In some examples, the data transmission unit (348) is configured to transmit data (e.g., presentation data, location data, etc.) to a user-facing component (120), one or more input devices (125), output devices (155), sensors (190), and / or peripheral devices (195). To this end, in various examples, the data transmission unit (348) includes commands and / or logic therefor, and heuristics and metadata therefor. A digital assistant (DA) unit (350) includes instructions and / or logic for providing DA functions to a computer system (101). Thus, the DA unit (350) provides DA functions to a user of the computer system (101) and / or their avatar while the user is present in a three-dimensional scene. For example, the DA performs various tasks related to the three-dimensional scene, either proactively or at the request of the user. In some examples, the DA unit (350) performs at least some of the following: converting voice input into text (e.g., using a speech-to-text (STT) processing unit (352)); identifying the user's intent expressed in natural language input received from the user; actively inducing and acquiring information necessary to fully satisfy the user's intent (e.g., by resolving ambiguities of terms in the natural language input and / or acquiring information from a data acquisition unit (341); and determining a task flow for implementing the identified intent. and execute the task flow to fulfill the identified intent. In some examples, the DA unit (350) includes a natural language processing (NLP) unit (351) configured to identify user intents. The NLP unit (351) takes n best candidate text expressions (word sequence(s) or token sequence(s)) generated by the STT processing unit (352) and attempts to associate each of the candidate text expressions with one or more user intents recognized by the DA. In some examples, a user intent represents a task that can be performed by the DA and has an associated task flow implemented in the task flow processing unit (353). The associated task flow is a series of programmed actions and steps that the DA takes to perform the task. In some examples, the range of capabilities of the DA depends on the number and types of task flows implemented in the task flow processing unit (353), or in other words, the number and types of user intents recognized by the DA. In some examples, once the NLP unit (351) identifies a user intent based on a user request, the NLP unit (351) causes the task flow processing unit (353) to perform actions necessary to satisfy the user request. For example, the task flow processing unit (353) executes a task flow corresponding to the identified user intent to perform a task to satisfy the user request. In some examples, performing the task includes causing the computer system (101) to provide an output (e.g., graphics, audio, and / or haptic output) indicating the performed task. The image evaluation unit (370) receives an input image (372) and a user request (374), as illustrated in FIG. 3b, and determines a prompt (376) and / or a response (378). In some examples, the image evaluation unit (370) is included in the DA unit (350). In some examples, some or all of the functions of the image evaluation unit (370) discussed below are performed by and / or together with the DA unit (350). An image evaluation unit (370) acquires (e.g., receives and / or captures) an input image (372) (e.g., an image containing an environment, an image containing an object, an image containing a person, and / or any combination of an environment, an object, and / or a person) using one or more image sensors of the device (101), such as image sensors (214), and receives a user request (374) related to the input image (372) using one or more sensors of the device (101), such as a microphone, a touch-sensitive display, and / or other sensors capable of receiving user voice and / or text input. In some examples, the user request (374) includes a task to be executed based on the content contained in the input image (372). In some examples, the user request (374) includes a request for data related to the content contained in the input image (372). In some examples, one or more image sensors are part of the same computer system and / or device as the image evaluation unit (370) (e.g., at least partially inside the computer system and / or directly connected to the computer system). In some examples, one or more image sensors are part of another computer system. In some examples, at least one image sensor is part of a computer system containing the image evaluation unit (370). In some examples, at least one image sensor is part of another computer system. In some examples, at least one image sensor is a forward-facing camera (e.g., the camera faces the front of the computer system) of a computer system containing the image evaluation unit (370). In some examples, at least one image sensor is a rear-facing camera (e.g., the camera faces the rear of the computer system) of a computer system containing the image evaluation unit (370). In some examples, the image sensor of a computer system and / or device including the image evaluation unit (370) is of lower quality than the image sensor of another computer system and / or device connected to and / or communicating with the computer system and / or device including the image evaluation unit (370). In some examples, the image sensor of a computer system and / or device including the image evaluation unit (370) includes at least one different characteristic (e.g., resolution, associated focal length, magnification, aperture, dynamic range, etc.) from the image sensor of another computer system and / or device connected to and / or communicating with the computer system and / or device including the image evaluation unit (370). Next, the image evaluation unit (370) determines the quality of the input image (372) and determines whether the quality of the input image (372) satisfies the quality criteria (e.g., whether it satisfies) or does not satisfy the quality criteria (e.g., whether it does not satisfy). In some examples, the quality of the input image (372) is based on factors including blurriness, sharpness, clarity, noise, exposure, tone, contrast, distortion, vignetting, artifacts, and / or lens flare present in the input image (372). In some examples, the image evaluation unit (370) determines the quality of the input image (372) by processing the image to determine whether and to what extent the factors discussed above are present. In some examples, the image evaluation unit (370) includes and / or uses one or more artificial intelligence (AI) models to determine the quality of the input image (372). In some examples, the quality standard is based on the user request (374) and / or the tasks included in the user request (374) (e.g., some tasks require higher quality images than others). In some examples, depending on the determination that the user request (374) includes a first type of request, the image evaluation unit (370) selects the first quality standard as the quality standard, and depending on the determination that the user request (374) includes a second type of request different from the first type, the image evaluation unit (370) selects the second quality standard different from the first quality standard as the quality standard. For example, when the image evaluation unit (370) determines that the task of the user request (374) is a task requiring a large amount of information from the input image (372), the image evaluation unit (370) selects a quality standard that the quality of the image must be relatively high (e.g., not blurry, sharp, clear, not noisy, not distorted, etc.), but when the image evaluation unit (370) determines that the task of the user request (374) is a task requiring a small amount of information from the input image (372), the image evaluation unit (370) selects a quality standard that the quality of the image must be relatively low (e.g., may be somewhat blurry, does not need to be completely clear, may contain noise and / or slight distortion, etc.). In some examples, the image evaluation unit (370) provides the input image (372) to a large language model (LLM) or other AI model and requests the LLM or other AI model to determine the quality of the input image (372). In some examples, the image evaluation unit (370) provides the input image (372) to a large language model (LLM) or other AI model and requests the LLM or other AI model to provide a determination of whether the input image (372) is of sufficient quality to complete a task determined from a user request (374). In some examples, the image evaluation unit (370) generates an embedding of the input image (372) and determines the quality of the input image (372) by comparing the embedding of the input image (372) with a set of learned embeddings representing a high-quality image or a low-quality image. In some examples, the image evaluation unit (370) generates an embedding of the input image (372) and determines the quality of the input image (372) by providing the embedding of the input image (372) to an LLM or other AI model. Subsequently, the image evaluation unit (370) requests the LLM or other AI model to determine the quality of the input image (372) by comparing the provided embedding with other embeddings of images of various quality. In some examples, the image evaluation unit (370) selects a set of learned embeddings based on the type of request included in the user request (374). For example, when the image evaluation unit (370) determines that the task of the user request (374) is a task that requires a large amount of information from the input image (372), the image evaluation unit (370) selects an embedding set that represents high-quality images, but when the image evaluation unit (370) determines that the task of the user request (374) is a task that requires a small amount of information from the input image (372), the image evaluation unit (370) selects an embedding set that represents low-quality images. In some examples, an image evaluation unit (370) provides an input image (372) to an artificial intelligence (AI) model and / or another model to execute a user request (374). When the confidence of the result of executing the user request (374), as determined by the model, is sufficiently high (e.g., when it satisfies the criteria for performing the task), the input image (372) is of sufficient quality to perform the task. When the confidence of the result of executing the user request (374), as determined by the model, is not sufficiently high (e.g., when it does not satisfy the criteria for performing the task), the input image (372) is not of sufficient quality to perform the task. In some examples, the confidence of the result of executing the user request (374) is provided to the image evaluation unit (370), and the image evaluation unit (370) uses the confidence of the result to determine whether another photo needs to be taken and / or whether the camera of another device needs to be opened (e.g., whether to execute, activate, call, etc.). When the image evaluation unit (370) determines that the quality of the input image (372) satisfies the quality criteria, the image evaluation unit (370) generates a response (378) to the user request (374) based on the input image (372) (e.g., by utilizing the capabilities of the DA unit (350) to determine the user intent and perform one or more actions to satisfy the user request (374), and provides an output containing the response (378). In some examples, the image evaluation unit (370) determines that the quality of the input image (372) satisfies the quality criteria when the computer system and / or digital assistant can determine the response to the user request because the quality of the first image is high. In some examples, the response (378) to the user request (374) includes an output that the task is completed, a response to an information request, and / or a follow-up prompt for additional information related to the user request (374). In some examples, the output of the response (378) is an audio output and / or an output on a display communicating with a computer system. When the image evaluation unit (370) determines that the quality of the input image (372) does not meet the quality criteria, the image evaluation unit (370) generates a prompt (376) to capture a second input image and provides a prompt (376) to allow another computer system and / or electronic device, whether physically connected to or not connected to the device (101), to capture the second input image with a sensor of the other computer system and / or electronic device. In some examples, the image evaluation unit (370) determines that the quality of the input image (372) does not meet the quality criteria when the image evaluation unit (370) determines that the computer system and / or digital assistant cannot determine a response to a user request (374) because the quality of the input image (372) is too low. In some examples, the prompt (376) includes an output containing a request to have another image captured by an image sensor (e.g., a camera) of another computer system. In some examples, the prompt (376) is provided as an output by the device (101) (e.g., the same computer system including the image evaluation unit (370)). In some examples, the prompt (376) is provided as an output by another computer system. In some examples, the prompt (376) is provided as an audio output. In some examples, the prompt (376) is provided as a visual output. In some examples, the prompt (376) is provided by a digital assistant associated with both computer systems. In some examples, the output is provided to a user interface associated with the digital assistant. In some examples, the output is provided to a user interface for a camera application. In some examples, two devices and / or computer systems are communicating. In some examples, two devices and / or computer systems are connected wirelessly (e.g., via Wi-Fi, Bluetooth, NFC, and / or other wireless communication protocols). In some examples, two devices and / or computer systems are both associated with the same user and / or the same user profile. In some examples, two devices are connected via wires and / or other physical connections. In some examples, after providing a prompt (376) and / or causing another device and / or computer system to provide the prompt (376), user input to capture another input image is detected, and in response to the detection of user input, another input image is acquired (e.g., received and / or captured). Subsequently, the image evaluation unit (370) determines a response to a user request (374) based on the other input image acquired using another device and / or computer system, and provides an output containing the response to the user request (374). Thus, the user receives a response to the user request (374) based on information available to both the devices and / or computer systems. In some examples, after detecting user input to capture another input image, the prompt (376) and / or other user interface are stopped from being displayed. In some examples, in response to detecting an input image (372) and a user request (374), the image evaluation unit (370) determines whether the context of the device (101) (e.g., a device and / or computer system that receives, captures, and / or acquires the input image (372)) indicates that the input image (372) does not satisfy quality criteria. The context of the device (101) (e.g., a device and / or computer system that receives, captures, and / or acquires the input image (372)) is determined using data received from one or more sensors of the device, which may include information representing the level of light around the device, the location of the device, the presence of objects in front of the device's image sensor, the movement of the device, the presence of text in front of the device, and / or other information related to the quality of the input image (372). For example, data from the device's sensors may indicate that the device is in a dark room or that there is an object right in front of the device's camera, so the input image (372) is too dark and / or out of focus and information cannot be extracted from it. In some examples, when the image evaluation unit (370) determines that the device context indicates that the input image (372) does not satisfy the quality criteria, the image evaluation unit (370) relinquishes the determination of whether the input image (372) satisfies the quality criteria and provides a prompt (376) to another computer system and / or device communicating with the device. In some examples, the 3D experience module (340) accesses one or more artificial intelligence (AI) models configured to perform the various functions described herein. The AI model(s) are implemented at least partially on the controller (110) (e.g., locally on a single device or in a distributed manner) and / or the controller (110) communicates with one or more external services that provide access to the AI model(s). In some examples, one or more components and functions of the DA unit (350), image evaluation unit (370), area tracking unit (380), and / or visibility analysis unit (390) are implemented using the AI model(s). For example, the DA unit (350) implements one or more AI models to perform speech recognition, intent determination (e.g., natural language processing and / or image processing), object recognition, and / or response generation, the image evaluation unit (370) implements one or more AI models to determine whether an input image satisfies quality criteria for determining a response to a user request and / or to determine a prompt for another image on another device and / or computer system, and / or the visibility analysis unit (390) implements one or more AI models to determine (e.g., identify) obscured parts of the captured image data. In some examples, AI model(s) are based on one or more foundational models (e.g., these or built from them). Generally, a foundational model is a deep learning neural network that is trained on a large training dataset and can be adapted to perform specific functions. Accordingly, the foundational model can aggregate information learned from a large (and optionally, multimodal) dataset and be adapted (e.g., fine-tuned) to perform various downstream tasks that the foundational model may not have originally been designed to perform. Examples of such tasks include language translation, speech recognition, user intent determination (e.g., natural language processing), sentiment analysis, computer vision tasks (e.g., object recognition and scene understanding), question answering, image generation, audio generation, and the generation of computer-executable instructions. Foundational models may accept a single type of input (e.g., text data) or multimodal inputs such as two or more of text data, image data, video data, audio data, sensor data, etc. In some examples, the underlying model is prompted to perform a specific task by providing it with a natural language description of the task. Exemplary underlying models include the GPT-n series models from Open AI, Inc. (e.g., GPT-1, GPT-2, GPT-3, and GPT-4), DALL-E, and CLIP, Florence and Florence-2 from Microsoft Corporation, BERT from Google LLC, and LLaMA, LLaMA-2, and LLaMA-3 from Meta Platforms, Inc. FIG. 4 illustrates an architecture (400) for a basic model according to some examples. The architecture (400) is merely illustrative, and various modifications to the architecture (400) are possible. Accordingly, the components of the architecture (400) (and their associated functions) may be combined, the order of the components (and their associated functions) may be changed, the components of the architecture (400) may be removed, and other components may be added to the architecture (400). Additionally, although the architecture (400) is Transformer-based, those skilled in the art will understand that the architecture (400) may additionally or alternatively implement other types of machine learning models, such as convolutional neural network (CNN)-based models and recurrent neural network (RNN)-based models. The architecture (400) is configured to process input data (402) to generate output data (480) corresponding to a desired task. The input data (402) includes one or more types of data, such as text data, image data, video data, audio data, sensor data (e.g., motion sensor, biometric sensor, temperature sensor, etc.), computer-executable instructions, structured data (e.g., XML file, JSON file, or other file types). In some examples, the input data (402) includes data from a data acquisition unit (341). The output data (480) includes one or more types of data that depend on the task to be performed. For example, the output data (480) includes one or more of the following: text data, image data, audio data, and computer-executable instructions. The input and output data types described above are merely exemplary, and it will be recognized that the architecture (400) may be configured to accept various types of data as input and generate various types of data as output. Such data types may vary based on specific functions configured for the underlying model to perform. The architecture (400) includes an embedding module (404), an encoder (408), an embedding module (428), a decoder (424), and an output module (450), the functions of which are now discussed below. The embedding module (404) is configured to receive input data (402) and parse the input data (402) into one or more sequences of tokens. The embedding module (404) is further configured to determine the embedding (e.g., vector representation) of each token representing each token in the embedding space, such that, for example, similar tokens have a closer distance in the embedding space and dissimilar tokens have a further distance. In some examples, the embedding module (404) includes a position encoder configured to encode position information into the embeddings. Each position information for an embedding indicates the relative position of the embedding in the sequence. The embedding module (404) is configured to output embedding data (406) of the input data by aggregating the embeddings for the tokens of the input data (402). The encoder (408) is configured to map embedding data (406) to an encoder representation (410). The encoder representation (410) represents contextual information for each token, indicating learned information about how each token relates to each other token (e.g., how attention is paid to them). The encoder (408) includes an attention layer (412), a feed-forward layer (416), normalization layers (414, 418), and residual connections (420, 422). In some examples, the attention layer (412) applies a self-attention mechanism to the embedding data (406) to calculate an attention representation (e.g., in the form of a matrix) of the relationship between each token and each other token in the sequence. In some examples, the attention layer (412) is multi-headed to compute multiple different attention expressions of the relationship between each token and each other token, each different expression representing a different learned attribute of the token sequence. The attention layer (412) is configured to aggregate the attention expressions to output attention data (460) representing cross-relationships between tokens from the input data (402). In some examples, the attention layer (412) further masks the attention data (460) to suppress the data representing relationships between selected tokens. Subsequently, the encoder (408) passes the (optionally masked) attention data (460) through the normalization layer (414), the feedforward layer (416), and the normalization layer (418) to generate an encoder expression (410).Residual connections (420, 422) can help stabilize and shorten the training and / or inference process by allowing the output of the embedding module (404) (i.e., embedding data (406)) to be directly passed to the normalization layer (414) and the output of the normalization layer (414) to be directly passed to the normalization layer (418). FIG. 4 illustrates an architecture (400) comprising a single encoder (408), but in other examples, the architecture (400) comprises a plurality of stacked encoders configured to output an encoder representation (410). Each of the stacked encoders can generate different attention data, which may allow the architecture (400) to learn different types of cross-relationships between tokens and generate output data (410) based on a more complete set of learned relationships. The decoder (424) is configured to receive the encoder representation (410) and the previous output embedding (430) as inputs to generate output data (480). The embedding module (428) is configured to generate the previous output embedding (430). The embedding module (428) is similar to the embedding module (404). Specifically, the embedding module (428) tokenizes the previous output data (426) (e.g., output data (480) generated by the previous iteration), determines embeddings for each token, and optionally encodes position information into each embedding to generate the previous output embedding (430). The decoder (424) includes attention layers (432, 436), normalization layers (434, 438, 442), a feedforward layer (440), and residual links (462, 464, 466). The attention layer (432) is configured to output attention data (470) that indicates cross-relationships between tokens from previous output data (426). The attention layer (432) is similar to the attention layer (412). For example, the attention layer (432) applies a multi-headed self-attention mechanism to the previous output embedding (430) and optionally masks the attention data (470) to suppress data expressing relationships between selected tokens (e.g., relationship(s) between a token and future token(s)), so that the architecture (400) does not consider future tokens as context when generating output data (480). Next, the decoder (424) passes the (optionally masked) attention data (470) through the normalization layer (434) to generate normalized attention data (470-1). The attention layer (436) takes the encoder representation (410) and normalized attention data (470-1) as inputs to generate encoder-decoder attention data (475). The encoder-decoder attention data (475) correlates the input data (402) with the previous output data (426) by expressing the relationship between the output of the encoder (408) and the previous output of the decoder (424). The attention layer (436) allows the decoder (424) to increase the weights of the parts of the encoder representation (410) that are learned to be more relevant to generating the output data (480). In some examples, the attention layer (436) generates the encoder-decoder attention data (475) by applying a multi-headed attention mechanism to the encoder representation (410) and to the normalized attention data (470-1). In some examples, the attention layer (436) further masks the encoder-decoder attention data (475) to suppress cross-relations between selection tokens. Next, the decoder (424) passes the (optionally masked) encoder-decoder attention data (475) through the normalization layer (438), the feedforward layer (440), and the normalization layer (442) to generate additionally processed encoder-decoder attention data (475-1). Next, the normalization layer (442) provides the additionally processed encoder-decoder attention data (475-1) to the output module (450). Similar to the residual connections (420, 422), the residual connections (462, 464, 466) can stabilize and shorten the training and / or inference process by allowing the output of a corresponding component to be directly passed as input to the corresponding component. FIG. 4 illustrates an architecture (400) comprising a single decoder (424), but in other examples, the architecture (400) comprises a plurality of stacked decoders, each configured to learn / generate different types of encoder-decoder attention data (475). This allows the architecture (400) to learn different types of cross-relations between tokens from input data (402) and tokens from output data (480), which may allow the architecture (400) to generate output data (480) based on a more complete set of learned relationships. The output module (450) is configured to generate output data (480) from additionally processed encoder-decoder attention data (475-1). For example, the output module (450) includes one or more linear layers that apply a learned linear transformation to the additionally processed encoder-decoder attention data (475-1), and a softmax layer that generates a probability distribution for possible classes of output tokens (e.g., words or symbols) based on the linear transformation data. Subsequently, the output module (450) selects (e.g., predicts) elements of the output data (480) based on the probability distribution. Subsequently, the architecture (400) passes the output data (480) to the embedding module (428) as previous input data (426) to start another iteration of the training and / or inference process for the architecture (400). It will be recognized that various different AI models can be built based on the components of the architecture (400). For example, some large language models (LLMs) (e.g., GPT-2 and GPT-3) are decoder-only (e.g., including one or more instances of decoder (424) and not including encoder (408)), some LLMs (e.g., BERT) are encoder-only (e.g., including one or more instances of encoder (408) and not including decoder (424)), and other base models (e.g., Florence-2) are encoder-decoder (e.g., including one or more instances of encoder (408) and one or more instances of decoder (424)). Additionally, it will be recognized that the basic models built based on the components of the architecture (400) can be fine-tuned based on specific task-specific training data and reinforcement learning techniques for optimization for specific tasks, such as extracting relevant semantic information from image and / or video data, generating code, generating music, and providing suggestions related to specific users. FIGS. 5a through 5e illustrate capturing images in a multi-device system according to some examples. The devices (500, 550) implement at least some of the components of the computer system (101). For example, the devices (500, 550) include one or more sensors configured to detect data (e.g., image data and / or audio data) corresponding to each scene. In some examples, the device (500) and / or the device (550) is an HMD (e.g., an XR headset or smart glasses), and FIGS. 5a through 5e illustrate a user's view of each scene through the HMD. For example, FIGS. 5a through 5e illustrate physical scenes viewed through pass-through video, physical scenes viewed through direct optical see-through, or virtual scenes viewed through one or more displays of the HMD. In other examples, the device (500) and / or the device (550) is a different type of device such as a smart watch, smartphone, tablet device, laptop computer, a pair of glasses without a display, headphones, earbuds, or a projection-based device. Examples of FIGS. 5a through 5e illustrate that users and devices (500, 550) exist within each scene. For example, the scenes are physical or extended reality scenes, and users and devices (500, 550) are physically present within the scenes. In other examples, a user's avatar exists within the scenes. For example, when the scenes are virtual reality scenes, a user's avatar exists within the virtual reality scenes. FIGS. 5a through 5e include a device (500) and a device (550), both of which can acquire image data (e.g., detect and / or capture) and receive user inputs, including user requests. In some examples, the device (500) and the device (550) communicate but are not physically connected. In some examples, the device (500) and the device (550) are wirelessly connected (e.g., via Bluetooth, Wi-Fi, NFC, and / or other wireless communication protocols). In some examples, the device (500) and the device (550) are connected via wires and / or other physical connections. In some examples, the device (500) and the device (550) are associated with the same user and / or the same user profile. In some examples, the device (500) and the device (550) are located near each other without being physically connected. In FIG. 5a, a device (500) (e.g., head-mounted device, smartphone, tablet, wearable computer system, smart device, and / or co-computer system) acquires (e.g., capture and / or receive) an image (502a) using one or more image sensors that communicate with the device (500) (e.g., camera of the device (500), camera of a device connected to the device (500), and / or camera of a device communicating with the device (500). In some examples, as illustrated in FIG. 5a, the image (502a) is displayed on a display and / or display generating component of the device (500) and / or communicating with the device (500). In some examples, the device (500) does not have a display, the image (502a) is not displayed, and / or the scene is viewed directly by the user. The device (500) also receives (e.g., detects, acquires, and / or captures) a user request (504a) related to the image (502a). In some examples, the user request (504a) is detected before acquiring the image (502a). In some examples, the user request (504a) is detected after acquiring the image (502a). In some examples, the user request (504a) is detected simultaneously with or substantially simultaneously with acquiring the image (502a). In response to acquiring an image (502a) and receiving a user request (504a), the device (500) uses an image evaluation unit (370) as discussed above with reference to FIG. 3b to determine whether the quality of the image (502a) satisfies a quality standard. The device (500) determines that the quality of the image (502a) satisfies a quality standard, and accordingly determines a response (506a) and provides the response (506a) as an audio output. In particular, based on the user request (504a) "What is that?", the device (500) determines that the user is attempting to identify an object in the image (502a). The device (500) determines (e.g., using an image evaluation unit (370)) that the quality of the image (502a) is sufficiently high so that the objects in the image (502a) can be identified, and accordingly, processes the image (502a) to determine that it contains trees and additionally contains oak trees. Then, the device (500) responds to a user request (504a) by generating a response that says "That is an oak tree" and providing it as an audio output. Because the device (500) determines that the quality of the image (502a) is high and thus meets the quality criteria, the device (500) does not determine a prompt to capture another image and does not cause the device (550) to provide a prompt or perform any other task. Therefore, as shown in FIG. 5a, the device (550) does not change or alter its display during the determination of the response (506a). In some examples, the device (500) provides the image (502a) to the device (550) and / or another device to determine whether the quality of the image (502a) satisfies a quality criterion using an image evaluation unit (370) as discussed above with reference to FIG. 3b. The device (550) determines that the quality of the image (502a) satisfies a quality criterion and thus determines a response (506a) and causes the device (500) to provide the response (506a) as an audio output. In particular, based on a user request (504a) of “What is that?”, the device (500) provides the user request (504a) to the device (550), and the device (550) determines that the user is attempting to identify an object in the image (502a). The device (550) determines (e.g., using an image evaluation unit (370)) that the quality of the image (502a) is sufficiently high so that the objects in the image (502a) can be identified, and accordingly, processes the image (502a) to determine that it contains a tree and additionally contains an oak tree. Then, the device (550) generates a response saying "That is an oak tree" and the device (550) provides it as an audio output to respond to a user request (504a). In FIG. 5b, the device (500) acquires an image (502b) using one or more image sensors that communicate with the device (500). In some examples, the image (502b) is displayed on a display and / or display generating component of the device (500) and / or communicating with the device (500). In some examples, the image (502b) is not displayed, the device (500) does not include a display, and / or the user views the scene directly. The device (500) also receives (e.g., detects, acquires, and / or captures) a user request (504b) related to the image (502b). In some examples, the user request (504b) is detected before acquiring the image (502b). In some examples, the user request (504b) is detected after acquiring the image (502b). In some examples, the user request (504b) is detected simultaneously with or substantially simultaneously with acquiring the image (502b). In response to acquiring an image (502b) and receiving a user request (504b), the device (500) uses an image evaluation unit (370) as discussed above with reference to FIG. 3b to determine whether the quality of the image (502b) satisfies a quality standard. The device (500) determines that the quality of the image (502b) does not satisfy a quality standard and therefore determines a prompt (508b). After determining (e.g., generating) the prompt (508b), the device (500) causes the device (550) to provide the prompt (508b). In particular, the device (500) causes the device (550) to display a prompt (508b) to capture a different image because the device (550) has a higher quality camera and is more likely to capture a higher quality image containing information to determine the response to the user request (504b). In some examples, as discussed above with reference to FIG. 3b, the device (500) determines that the quality of the image (502b) does not meet quality standards because the image (502b) is blurry, not sharp, has artifacts, is obscured, has noise, is distorted, and / or has other factors that degrade the quality of the image (502b). After the device (550) causes the prompt (508b) to display, the device (550) and / or the device (500) detect user input (510b) for the "Yes" button of the prompt (508b). In response to detecting the input (510b), the device (550) acquires (e.g., captures and / or receives) an image (502c) as illustrated in FIG. 5c. After acquiring the image (502c), the device (550) provides the image (502c) and / or data representing the image (502c) to the device (500) so that the device (500) can determine a response (506c) to the user request (504b). Subsequently, the device (500) determines the response (506c) and provides the response (506c) as an audio output. In some examples, the device (550) provides the response (506c) as an audio output instead of the device (500). In some examples, the device (500) and / or the device (550) display the response (506c) on the display of the device (500) and / or the device (550) and / or on a display generating component communicating with the device (500) and / or the device (550). In some examples, after detecting user input (510b), the device (500) stops displaying the image (502b) as shown in FIG. 5c. In some examples, the device (550) provides the response (506c) as an audio output instead of the device (500). In some examples, the device (500) and / or the device (550) display the response (506c) on the display of the device (500) and / or the device (550) and / or on a display generating component communicating with the device (500) and / or the device (550). In some examples, after detecting user input (510b), the device (500) stops displaying the image (502b) as shown in FIG. 5c. In some examples, in response to acquiring an image (502b) and receiving a user request (504b), the device (500) provides the image (502b) and the user request (504b) to the device (550), and allows the device (550) to determine whether the quality of the image (502b) satisfies a quality criterion using an image evaluation unit (370) as discussed above with reference to FIG. 3b. The device (550) determines that the quality of the image (502b) does not satisfy a quality criterion and accordingly determines a prompt (508b). After determining (e.g., generating) the prompt (508b), the device (550) provides the prompt (508b). In particular, the prompt (508b) includes a prompt to capture a different image because the device (550) has a higher quality camera and is more likely to capture a higher quality image containing information to determine the response to the user request (504b). In some examples, as discussed above with reference to FIG. 3b, the device (550) determines that the quality of the image (502b) does not meet quality standards because the image (502b) is blurry, not sharp, has artifacts, is obscured, has noise, is distorted, and / or has other factors that degrade the quality of the image (502b). After the device (550) displays the prompt (508b), the device (550) and / or the device (500) detect user input (510b) for the "Yes" button of the prompt (508b). In response to detecting the input (510b), the device (550) acquires (e.g., captures and / or receives) an image (502c) as illustrated in FIG. 5c. After acquiring the image (502c), the device (550) determines a response (506c) to the user request (504b). Subsequently, the device (550) determines the response (506c), provides the response (506c) as an audio output, and / or provides the response (506c) to the device (500) so that it is provided as an audio output. In FIG. 5d, the device (500) acquires an image (502d) using one or more image sensors that communicate with the device (500). In some examples, the image (502d) is displayed on a display and / or display generating component of the device (500) and / or communicating with the device (500). In some examples, the device (500) does not have a display, the image (502d) is not displayed, and / or the scene is viewed directly by the user of the device (500). The device (500) also receives (e.g., detects, acquires, and / or captures) a user request (504d) related to the image (502d). In some examples, the user request (504d) is detected before acquiring the image (502d). In some examples, the user request (504d) is detected after acquiring the image (502d). In some examples, the user request (504d) is detected simultaneously with or substantially simultaneously with acquiring the image (502d). In response to acquiring an image (502d) and receiving a user request (504d), the device (500) determines, using an image evaluation unit (370) as discussed above with reference to FIG. 3b, that the context of the device (500) indicates that the image (502d) does not meet quality criteria. In particular, the device (500) determines that the device (500) is moving based on data received from one or more sensors of the device (500) while capturing the image (502d). Accordingly, the device (500) decides to forgo determining whether the quality of the image (502d) meets quality criteria and determines a prompt (508d) to open a camera application user interface. After determining (e.g., generating) the prompt (508d), the device (500) causes the device (550) to provide the prompt (508d) by causing the device (550)’s camera application to open and display the user interface of the camera application. In particular, the device (500) causes the device (550) to open the camera application (e.g., displaying the prompt (508d)) to capture a different image because the device (550) has a higher quality camera and is more likely to capture a higher quality image containing information to determine the response to the user request (504d). After the device (550) causes the prompt (508b) to display, the device (550) and / or the device (500) detect user input (510d) for the capture button of the camera user interface. In response to detecting the input (510d), the device (550) acquires an image (e.g., captures and / or receives) and provides the image and / or data representing the image to the device (500) so that the device (500) can determine a response (506e) to the user request (504d). Subsequently, the device (500) determines the response (506e) and provides the response (506e) as an audio output as illustrated in FIG. 5e. In some examples, the device (550) provides the response (506e) as an audio output instead of the device (500). In some examples, the device (500) and / or the device (550) displays the response (506e) on the display of the device (500) and / or the device (550) and / or on a display generating component communicating with the device (500) and / or the device (550). In some examples, after detecting user input (510d), the device (500) stops displaying the image (502d) as shown in FIG. 5e. In some examples, after detecting user input (510d), the device (550) stops displaying the prompt (508d) and / or stops displaying the camera user interface and instead displays the lock screen as shown in FIG. 5e. Accordingly, in some examples, the device (550) does not display the captured image in response to detecting user input (510d), but instead provides the captured image and / or data corresponding to the image without displaying the image. In some examples, in response to acquiring an image (502d) and receiving a user request (504d), the device (500) provides the image (502d) and the user request (504d) to the device (550), and the device (550) determines that the context of the device (500) indicates that the image (502d) does not meet quality criteria using the image evaluation unit (370) as discussed above with reference to FIG. 3b. In particular, the device (550) determines that the device (500) is moving while the device (500) is capturing the image (502d) based on data received from the device (500) and / or one or more sensors of the device (550). Accordingly, the device (550) gives up on determining whether the quality of the image (502d) meets quality criteria and determines a prompt (508d) to open the camera application user interface. After determining (e.g., generating) the prompt (508d), the device (550) provides the prompt (508d) by opening (e.g., running, activating, and / or calling) the camera application of the device (550) and displaying the user interface of the camera application. In particular, the device (550) captures a different image by opening the camera application (e.g., displaying the prompt (508d)) because the device (550) has a higher quality camera and is more likely to capture a higher quality image containing information to determine the response to the user request (504d). After the device (550) displays the prompt (508d), the device (550) and / or the device (500) detect user input (510d) for the capture button of the camera user interface. In response to detecting the input (510d), the device (550) acquires an image (e.g., captures and / or receives) and determines a response (506e) to the user request (504d). Subsequently, the device (550) provides the response (506e) as an audio output. Although the above examples have been discussed in terms of a device (500) receiving a first image and determining whether the quality of the first image satisfies a quality criterion, it will be understood that a device (550) can also receive a first image and determine whether the quality of the first image satisfies a quality criterion. Similarly, one device such as device (500) can capture an image, and another device such as device (550) can determine whether the quality of the image satisfies a quality criterion. Thus, steps of capturing images, determining whether the quality satisfies a criterion, and / or causing a prompt to be displayed on another device may be performed by any device among the devices in the system. Similarly, although the above examples discuss two devices in the system, the system may include three, four, five, or any other number of devices connected and / or communicating wirelessly to exchange data including captured images and determinations of whether the quality of the images satisfies a quality criterion. Further explanations regarding FIGS. 5a through 5e are provided below with reference to the method (600) described below in relation to FIG. 6. FIG. 6 is a flowchart of a method (600) for capturing images in a multi-device system. In some examples, the method (600) is performed in a first computer system (e.g., computer system (101) of FIG. 1, device (500), and / or device (550)) that communicates with one or more image sensors (e.g., image sensors, light sensors, and / or photo sensors). In some examples, the method (600) is controlled by instructions stored in a non-transient (or transient) computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processing unit(s) (302) of the computer system (101) (e.g., controller (110) of FIG. 1). In some examples, the operations of the method (600) are distributed across multiple computer systems, such as a computer system and a separate server system. Some of the operations of the method (600) are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted. In block (602), a first image (e.g., 372, 502a, 502b, and / or 502d) is acquired using one or more image sensors. In block (604), a user request related to the first image (e.g., 374, 504a, 504b, and / or 504d) is received. In block (608), in response to acquiring a first image and receiving a user request (606), a prompt (e.g., 376, 508b, and / or 508d) is provided to the second computer system (e.g., computer system (101) of FIG. 1, device (500), and / or device (550)) to capture a second image (e.g., 372, 502a, 502b, and / or 502d). In block (612), in response to acquiring a first image and receiving a user request (606), a response to the user request (e.g., 378, 506a, 506c, and / or 506e) is generated based on the first image, in accordance with a determination (610) that the quality of the first image satisfies a quality standard (e.g., by an image evaluation unit (370)). In block (614), an output including a response to a user request based on the first image is provided, upon determination (610) that the quality of the first image satisfies the quality criteria. In some examples, the first computer system is a head-mounted electronic device, and the second computer system is a smartphone. In some examples, the image sensor of the first computer system is a lower quality image sensor than the image sensor of the second computer system. In some examples, the method (600) further comprises, in response to detecting a first image and receiving a user request: selecting a first quality standard as a quality standard based on a determination that the user request includes a first type of request; and selecting a second quality standard as a quality standard based on a determination that the user request includes a second type of request different from the first type. In some examples, the method (600) further comprises: detecting user input to capture the second image to the second computer system after the second computer system has provided a prompt to the second computer system to capture the second image; generating a response to the user request based on the second image; and providing an output including the response to the user request based on the second image. In some examples, determining whether the quality of a first image satisfies a quality criterion involves providing a prompt to a large language model (LLM), wherein the prompt includes a request for whether the first image is of sufficient quality to complete a task determined from a user request. In some examples, determining whether the quality of a first image satisfies a quality criterion includes generating an embedding of the first image; and comparing the embedding of the first image with a set of learned embeddings representing a high-quality image or a low-quality image. In some examples, the method (600) further includes the step of selecting a first learned embedding set as a learned embedding set based on a determination that the user request is a first type of request; and the step of selecting a second learned embedding set as a learned embedding set based on a determination that the user request is a second type of request. In some examples, the method (600) further comprises: a step of detecting a first image and, in response to receiving a user request: a decision that the context of the first computer system indicates that the first image does not satisfy a quality criterion; a step of abandoning the determination of whether the first image satisfies a quality criterion; and a step of causing the second computer system to provide a prompt to the second computer system to capture the second image. In some examples, the determination that the context of the first computer system indicates that the first image does not satisfy a quality criterion includes the determination that the first computer system is moving. In some examples, the determination that the context of the first computer system indicates that the first image does not satisfy a quality criterion includes the determination that the lighting level of the environment of the first computer system is below a lighting threshold. In some examples, the determination that the context of the first computer system indicates that the first image does not satisfy a quality criterion includes the determination that one or more image sensors are obscured. In some examples, the determination that the context of the first computer system indicates that the first image does not satisfy a quality criterion includes the determination that the field of view of one or more image sensors contains text. In some examples, the method (600) further includes the step of displaying a camera user interface to a display generating component communicating with a second computer system, based on a determination that the quality of the first image does not satisfy a quality standard. In some examples, the method (600) further comprises: detecting a user input to capture a second image after the camera user interface has been displayed to a display generating component communicating with a second computer system; and stopping the camera user interface from being displayed to a display generating component communicating with a second computer system in response to the detection of the user input to capture a second image. In some examples, the camera user interface is displayed on the lock screen as a display generating component that communicates with a second computer system. In some examples, the second computer system and the first computer system are not physically connected. In some examples, the second computer system and the first computer system are physically connected by a wire and are not located in the same housing. Returning to FIG. 3a, the region tracking unit (380) and the visibility analysis unit (390) are configured to determine a region of interest within a 3D scene and to determine a visibility metric representing the amount of the region of interest represented by the captured image data. The visibility analysis unit (390) is further configured to allow the device (e.g., 1000 of FIG. 10a through 10g) to provide various outputs that depend on the visibility metric, as described below in relation to FIG. 10a through 10g. The region tracking unit (380) is configured to determine and update a region of interest (e.g., 802, 808, 906, 916, 1012, or 1092 in FIGS. 8a and 8b, FIGS. 9a and 9b, and FIGS. 10a through 10g) within a 3D scene in which the user is immersed. The region of interest may include one or more objects (e.g., physical objects or virtual objects) to which the user can issue corresponding user requests, such as "What is this?", "What is the price of this?", "Please add this to my shopping list," etc. In some examples, the region of interest is oriented forward relative to the user's head pose (the position and orientation of the user's head). For instance, the region of interest is located in front of the user's head (e.g., in front of the user's face), and the user can view the region of interest without changing their current head pose. Because the region of interest is oriented forward relative to the user's head pose, the region of interest (e.g., the position of the region of interest within the 3D scene) changes as the user's head pose changes, for instance, as the user turns their head, and / or as the user moves around. For instance, when the user turns their head from looking straight ahead to looking upward, the region of interest changes from an area in front of the user to an area upward relative to the user. As another example, when the user turns their head from looking to the right to looking to the left, the region of interest changes from an area to the right relative to the user to an area to the left relative to the user. As another example, if a user rotates 180 degrees while maintaining a neutral head position, the region of interest changes from the area previously in front of the user to the area currently in front of the user (and previously behind the user). The region of interest is oriented forward relative to the user's head pose because users are likely to refer to objects in front of their heads when issuing requests related to objects in the 3D scene, such as "What is the price of this?". In some examples, the region tracking unit (380) determines a region of interest based on the respective positions and orientations of two different devices (e.g., 1002 and 1004 in FIG. 10a through 10g). The positions and orientations are computed for a user (e.g., the user's head, the user's face, and / or the user's ear), for a coordinate system centered between the two different devices, and / or for a 3D scene (e.g., a reference point (e.g., an object) within the 3D scene or for a surface (e.g., a floor surface) within the 3D scene). In some examples, the two different devices are worn simultaneously by the user. For example, the two different devices are a first device (e.g., a first camera) worn on a first side of the user's head and a second device (e.g., a second camera) worn on a second side opposite the user's head. In some examples, the first device is worn on the user's first ear (e.g., left ear or right ear) (e.g., by being physically housed in a device to be worn), and the second device is worn on the user's different second ear (e.g., left ear or right ear) (e.g., by being physically housed in a device to be worn). In some examples, the first device and the second device are physically housed in a single device worn by the user, for example, on the user's head (e.g., a headset or a pair of glasses). For example, the first device is a first camera positioned on the first side of the single device (e.g., positioned near the right or left side of the user's head when the single device is worn on the user's head), and the second device is a second camera positioned on the opposite second side of the single device (e.g., positioned near the right or left side of the user's head when the single device is worn on the user's head). FIG. 7 illustrates a forward-facing head pose (702) determined based on the pose (704) (position and orientation) of the first device and the pose (706) (position and orientation) of the second device, according to some examples. As described below, the region tracking unit (380) determines the region of interest based on the forward-facing head pose (702). The forward-facing head pose (702) corresponds to a user view cone (708) (e.g., 806 in FIG. 8a) which represents at least a portion of the user's view of the 3D scene when the user's head has the forward-facing head pose (702). The pose (704) of the first device corresponds to a first device view cone (710) which represents the view of the 3D scene captured by the camera of the first device when the first device has the pose (704). The pose (706) of the second device corresponds to a second device view cone (712) that represents a view of a 3D scene captured by the camera of the second device when the second device has the pose (706). In some examples, to determine the region of interest, the region tracking unit (380) relies on a 6-degrees of freedom (6DOF) relationship between the poses (704, 706) and the forward-facing head pose (702). Specifically, based on 6DOF information for the pose (704) of the first device (e.g., values for three spatial dimensions indicating position and three angular dimensions indicating orientation) and 6DOF information for the pose (706) of the second device, the region tracking unit (380) determines (e.g., approximates) 6DOF information for the forward-facing head pose (702). In some examples, the position of the forward-facing head pose (702) determined based on the 6DOF relationship corresponds (e.g., approximates) to a position centered between the user's two eyes. A position centered between the user's two eyes can provide an accurate reference point for determining the forward-facing region of interest. In some examples, determining the forward-facing head pose (702) using 6DOF relationships includes computing the relative relationship between the first device and the second device (e.g., how the first device and the second device are oriented and positioned relative to each other), computing the relative relationship between the first device and the user (e.g., the user's ear or the user's head) (e.g., how the first device is positioned and oriented relative to the user), and computing the relative relationship between the second device and the user (e.g., the user's ear or the user's head) (e.g., how the second device is positioned and oriented relative to the user). Sometimes, the respective orientations of the first device and the second device (orientations of the poses (704, 706)) do not correspond to the orientation of the forward-facing head pose (702) due to different ways in which the first device and the second device are worn. For example, the default way of wearing the first device and the second device (e.g., the default orientation of the first device and the second device relative to the user's head and / or ears) may result in the first camera and the second camera facing approximately the same direction as the forward-facing head pose (702) (e.g., as illustrated by the view cones (710, 712, 708) of FIG. 7). However, the user sometimes wears the first device and / or the second device in a non-default manner (e.g., rotated upward, downward, or sideways relative to the user's head and / or ears), resulting in the first camera and / or the second camera facing in respective directions different from the direction of the forward-facing head pose (702). Accordingly, computing the relative relationship between the first device and the user and / or the relative relationship between the second device and the user allows the area tracking unit (380) to determine whether the respective orientations of the first device and the second device correspond to the orientation of the forward-facing head pose (702) (e.g., when the user is looking forward and the first camera and the second camera are also facing forward, when the user is looking upward and the first camera and the second camera are also facing upward, etc.) or not (e.g., when the user is looking forward but the first camera and / or the second camera are rotated to face upward, when the user is looking upward but the first camera and / or the second camera are rotated to face forward, etc.). FIGS. 8A and FIGS. 8B illustrate regions of interest (802 in FIG. 8A or 808 in FIG. 8B) determined based on a forward-facing head pose (702) according to some examples. In FIG. 8a, the head of the user (800) has a forward-facing head pose (702). The area tracking unit (380) defines the area of interest (802) based on defining a forward-facing view cone (806) that starts at the position of the forward-facing head pose (702) and is centered on the forward-facing head pose (702). In some examples, the forward-facing view cone (806) has predetermined dimensions defined by a first angular deviation amount (e.g., ±30°) in a first dimension orthogonal to the orientation of the forward-facing head pose (702) (e.g., the dimension representing up and down in FIG. 8a) and a second angular deviation amount (e.g., ±25°) in a second dimension orthogonal to the orientation of the forward-facing head pose (702) and the first dimension (e.g., the dimension representing left and right in FIG. 8a), so that, for example, the slice of the forward-facing view cone (806) becomes a circle or an ellipse. In this way, the forward-facing view cone (806) represents an area of the 3D scene considered to be in front of the user (800), and the boundary of the forward-facing view cone (806) approximates the boundary between the area of the 3D scene in front of the user (800) and the area of the 3D scene around the user (800). In some examples, the area tracking unit (380) defines the area of interest (802) (exemplified by hatched lines) as a part of the forward-facing view cone (806) located at least a predetermined distance (804) (e.g., 0.1 meters, 0.2 meters, 0.3 meters, 0.4 meters, or 0.5 meters) from the user (800) (e.g., the head of the user (800) or the face of the user (800)). The region of interest (802) is defined in the manner described above, because the user (800) is unlikely to issue a user request regarding an object that is too close to their face, and the user (800) is unlikely to issue a user request regarding an object that is in their vicinity. In FIG. 8b, the head of the user (800) has a forward-facing head pose (702). The area tracking unit (380) defines a region of interest (808) (exemplified by hatched lines) based on the forward-facing head pose (702) and typical regions for handheld objects. The typical region for handheld objects specifies, for any user's forward-facing head pose, the typical region (e.g., 1-sigma region of confidence, 2-sigma region of confidence, or 3-sigma region of confidence) where any user holds an object in their hand when issuing a request regarding an object, e.g., "Tell me about this object in my hand." In some examples, the typical region for handheld objects is determined through user studies and / or testing, where a group of users is asked to hold objects in their respective hands and issue user requests regarding the objects. A typical area for handheld objects is located at a typical distance from any user's head and / or face and has typical dimensions (e.g., length, width, depth, shape, or volume). The area tracking unit (380) defines an area of interest (808) (exemplified by hatched lines) based on the typical distance and typical dimensions of the typical area for handheld objects. The area tracking unit (380) further positions the area of interest (808) based on the orientation of the forward-facing head pose (702) (as represented by the arrow in FIG. 8b) (e.g., to intersect the arrow, or to have a predetermined amount of angular deviation from the arrow). The area of interest (808) is defined in the manner described above to account for scenarios in which the user is holding the object in their hand when issuing a request regarding the object. The visibility analysis unit (390) is configured to determine a visibility metric that represents the amount of a region of interest (e.g., 802, 808) depicted by the captured image data. For example, a high visibility metric indicates that a relatively large amount of the region of interest is depicted by the captured image data, and a low visibility metric indicates that a relatively small amount (or no amount at all) of the region of interest is depicted by the captured image data. As described below in relation to FIGS. 10a through 10g, in response to receiving a user request (e.g., in the form of natural language input) related to an object in a 3D scene, the visibility analysis unit (390) causes the device (1000) to perform one or more actions that depend on the visibility metric. For example, when the visibility metric is high, it is likely that the correct object for the user request is sufficiently visible (e.g., depicted) in the captured image data, so the device (1000) provides an output that satisfies the user request (e.g., providing information about the correct object and / or performing a task based on the correct object). However, when the visibility metric is low, the correct object for the user request may not be sufficiently visible in the captured image data, so the device (1000) performs various actions to ensure the correct object is sufficiently visible in the captured image data before providing an output that satisfies the user request (e.g., to avoid outputting an incorrect audio output for an incorrect object and / or to avoid outputting an error due to failure to identify the correct object). In some examples, the visibility metric determined depends on the type of natural language input containing the user request. In some examples, the visibility analysis unit (390) calls the natural language processing capabilities of the DA unit (350) to determine whether the natural language input is of a first type (e.g., referring to an object in the user's hand) or a second type (e.g., not referring to an object in the user's hand). If the natural language input is of the first type, the visibility analysis unit (390) selects a region of interest (808) (Fig. 8b) as a region of interest and determines a visibility metric for the region of interest (808). If the natural language input is of the second type, the visibility analysis unit (390) selects a region of interest (802) (Fig. 8a) as a region of interest and determines a visibility metric for the region of interest (802). Accordingly, the visibility analysis unit (390) can advantageously select a relevant region of interest to determine the visibility metric based on the content of the natural language input. For example, when a user issues a request indicating that an object is in their hand, such as "What is the price of this object in my hand?", it may be advantageous to determine how much of the area for handheld objects (e.g., 808) is visible. And even when a user issues a request about an object (e.g., "What is the price of this?") without indicating that the object is in their hand, it may still be advantageous to determine how much of the area for the front (e.g., 802) is visible, for example, because the object is likely to be positioned within the area for the front. In some examples, image data is captured (e.g., simultaneously) by two separate cameras, e.g., the first camera and the second camera described above in relation to the area tracking unit (380). In examples where two separate cameras are worn on both sides of the user's head (e.g., opposite sides) (e.g., one camera is worn on each ear), it may be desirable to capture image data with two separate cameras to capture a complete view of the area of interest (e.g., 802 or 808). For example, the first camera alone cannot capture a complete view of the area of interest because part of the view of the 3D scene by the first camera is obscured by the corresponding side of the user's head. Similarly, the second camera alone cannot capture a complete view of the area of interest because part of the view of the 3D scene by the second camera is obscured by the corresponding side of the user's head. Accordingly, in some examples, image data refers to a combination of two distinct images (or two distinct sets of images) each captured by a different camera. FIGS. 9a and 9b illustrate the determination of a visibility metric for a region of interest (802 or 808) according to some examples. Generally, the visibility analysis unit (390) determines the visibility metric by determining the amount of overlap between the region of interest and the non-occluded region of the 3D scene depicted by the image data. For example, the visibility analysis unit (390) identifies pixels of the image data that represent an unoccluded view of the 3D scene and also depict the region of interest. It will be recognized that the image data may be obscured by the user's face, the user's hair, the user's clothes, the user's hands, smudges on the camera(s), etc. The visibility analysis unit (390) implements techniques known in the art to classify various parts of the image data as obscured or unobscured (e.g., to identify obscured pixels and unobscured pixels). FIG. 9a illustrates an example in which a visibility analysis unit (390) determines a relatively low visibility metric for a region of interest (906) (e.g., 802 or 808). In FIG. 9a, the camera area (902) represents an area of a 3D scene depicted by image data, for example, when the image data is not obscured. In some examples, the visibility analysis unit (390) determines the camera area (902) based on the poses (704, 706) of two different devices, for example, two different cameras. For example, the visibility analysis unit (390) maps the 3D scene and determines the camera area (902) using the poses (704, 706) and the fields of view of the cameras. In FIG. 9a, the camera area (902) includes a occluded area (904) (exemplified by horizontal hatching). The occluded area (904) represents a portion of the 3D scene that is obscured in the image data. In FIG. 9a, the visibility analysis unit (390) projects the camera area (902) onto the area of interest (906) (e.g., by using poses (702, 704, 706), known dimensions and positions of the area of interest (906), and known dimensions of the fields of view of the cameras) to determine the correlation between the camera area (902) and the area of interest (906). For example, the visibility analysis unit (390) identifies an overlapping area (908) (as exemplified by vertical hatching) that represents the overlap between the camera area (902) and the area of interest (906) (e.g., how much of the area of interest (906) is depicted by the camera area (904) when the image data is not obscured). Subsequently, the visibility analysis unit (390) identifies a portion of the overlapping area (908) that is not obscured by the visibility area (910), e.g., a portion of the overlapping area (908) that has vertical hatching and does not have horizontal hatching. Subsequently, the visibility analysis unit (390) determines a visibility metric based on the pixels of the image data representing the visibility area (910). In some examples, different pixels representing the visibility area (910) have different weights in relation to determining the visibility metric. For example, the positive magnitude of the contribution of a pixel of the visibility area (910) to the visibility metric decreases as the pixel moves further away from the center of the region of interest (906 or 916), and accordingly, a pixel representing the central part of the region of interest (906 or 916) provides a greater positive contribution to the visibility score than a pixel representing the edge part of the region of interest (906 or 916). For example, assume that the same number of pixels depict the region of interest (906). When the same number of pixels mainly depict the central part of the region of interest (906), the resulting visibility metric is higher than when the same number of pixels mainly depict the edge part of the region of interest (906). In FIG. 9a, the visibility analysis unit (390) determines a relatively low visibility metric because there are relatively few pixels representing the visibility area (910) and because the pixels representing the visibility area (910) have a relatively low weight in relation to determining the visibility metric (e.g., because the pixels represent the edge portions of the region of interest (906). The example in FIG. 9a results in a relatively low visibility metric due to significant occlusion of the camera area (902) (as exemplified by the size of the occlusion area (904) within the camera area (902)) and a relatively small amount of overlap between the camera area (902) and the region of interest (906). The small amount of overlap between the camera area (902) and the region of interest (906) may be because the orientation(s) of the camera do not correspond to the orientation of the user's head in a forward-facing pose. For example, the small amount of overlap between the regions (902, 906) is because the user is facing forward and the camera(s) are rotated upward to capture an image of the region mainly above the user. FIG. 9b illustrates an example in which a visibility analysis unit (390) determines a relatively high visibility metric for a region of interest (916) (e.g., 802 or 808). In FIG. 9b, the camera area (912) represents an area of the 3D scene depicted by the image data, for example, when the image data is not obscured. The camera area (912) is similar to the camera area (902) and is determined in a similar manner. The camera area (912) includes an obscured area (914) (exemplified by horizontal hatching). The obscured area (914) represents a portion of the 3D scene that is obscured in the image data. In FIG. 9b, similar to FIG. 9a, the visibility analysis unit (390) projects the camera area (912) onto the region of interest (916) to identify an overlapping area (918) (as exemplified by vertical hatching) that represents the overlap between the camera area (912) and the region of interest (916) (e.g., how much of the region of interest (916) the camera area (912) depicts when the image data is not obscured). The visibility analysis unit (390) identifies the portion of the overlapping area (918) that is not obscured by the visibility area (920) (the portion of the overlapping area (918) that has vertical hatching and does not have horizontal hatching). Subsequently, the visibility analysis unit (390) determines a visibility metric based on the pixels of the image data representing the visibility area (920). In FIG. 9b, the visibility analysis unit (390) determines a relatively high visibility metric because the number of pixels representing the visibility area (920) is relatively large and some pixels representing the visibility area (920) have relatively high weights in relation to determining the visibility metric (e.g., because the pixels represent the central part of the area of interest (916). The example in FIG. 9b results in a relatively high visibility metric due to a relatively small amount of occlusion of the camera area (912) (as exemplified by the size of the occlusion area (914) within the camera area (912)) and a relatively large amount of overlap between the camera area (912) and the area of interest (916). The large amount of overlap between the camera area (912) and the area of interest (916) may be because the orientation(s) of the camera correspond to the orientation of the user's head in a forward-facing pose. For example, the large amount of overlap between the camera area (912) and the area of interest (916) is because the user is looking straight ahead and the camera(s) are also oriented to face approximately straight ahead (e.g., relative to the user's head). As described below in connection with FIGS. 10a through 10g, the visibility analysis unit (390) is configured to cause the device (1000) to perform various actions based on whether a determined visibility metric satisfies a condition (e.g., a threshold). In some examples, the visibility metric satisfies the condition if the visibility metric is above the threshold, and the visibility metric satisfies the condition if the visibility metric is below the threshold. FIGS. 10a through 10g illustrate that a device (1000) performs various actions in response to receiving natural language input and according to a determined visibility metric, according to some examples. The device (1000) implements at least some of the components of the computer system (101). In some examples, the device (1000) includes one or more sensors configured to detect audio data (e.g., natural language user requests), one or more cameras configured to detect image data for which visibility metrics are determined, and one or more audio output devices (e.g., speakers) configured to provide audio output. In the examples of FIGS. 10a through 10g, the device (1000) is worn on the head of a user (1010). For example, the device (1000) is an XR headset, a pair of smart glasses, headphones, or a set of earbuds. In other examples, the device (1000) is another type of device such as a smart watch, a smartphone, a tablet device, a laptop computer, or a projection-based device. In the examples of FIGS. 10a through 10g, the device (1000) includes a camera (1002) (e.g., a set of one or more cameras) and a camera (1004) (e.g., a set of one or more cameras). The camera (1002) is positioned near a first side of the user (1010)'s head (e.g., worn), and the camera (1004) is positioned near a second side opposite the user (1010)'s head (e.g., worn). For example, a camera (1002) is physically housed in a first device (e.g., a first earbud) worn on a first side of the user's (1010) head (e.g., worn on the user's (1010) first ear), and a camera (1004) is physically housed in a second device (e.g., a second earbud) worn on a second side of the user's (1010) head opposite (e.g., worn on the user's (1010) opposite second ear). In FIGS. 10a to 10g, the left portion is a lateral dimension ( x ), height dimension( y ), and depth dimensions ( zA coordinate system defining ) is exemplified. FIGS. 10a through 10g exemplify a coordinate system defining directions relative to the user (1010) in the context of FIGS. 10a through 10g, namely, up, down, right, left, forward, and rear. Specifically, an object or area is at the height coordinates ( of the user (1010) y′ Height coordinates greater than ) y In the case where it has ), the object or area is above (e.g., above) the user (1010), and the object or area is at the height coordinates (of the user (1010) y′ Height coordinates smaller than ) y In the case where it has ), the object or area is downward (e.g., below) with respect to the user (1010), and the object or area is at the user's (1010) lateral coordinates ( x′ Lateral coordinates larger than ) x In the case where it has ), the object or area is to the right of the user (1010), and the object or area is at the lateral coordinates (of the user (1010) x′ Lateral coordinates smaller than ) x In the case where it has ), the object or region is to the left of the user (1010), and the object or region is at the depth coordinates (of the user (1010) z′ Depth coordinates greater than ) z In the case where it has ), the object or area is in front of the user (1010) (e.g., in front of the user (1010)), and the object or area is at the depth coordinates of the user (1010) ( z′ Depth coordinates smaller than ) z In the case of having ), the object or area is behind the user (1010) (e.g., behind the user (1010)). In some examples, the height, lateral, and depth coordinates of the user (1010) ( x′ , y′ , z′) are the height, lateral, and depth coordinates of the user's (1010) head. In some examples, the height, lateral, and depth coordinates of the user's (1010) head are the height, lateral, and depth coordinates of other parts of the user's (1010) head, e.g., the user's (1010) face or chest. In FIGS. 10a through 10g, the right portion of the drawings illustrates an image (e.g., 1018, 1026, 1046, 1076, or 1093) captured by the device (1000) or a user interface (e.g., 1068) displayed by an external device (1062). In FIG. 10a, the camera (1002) has a forward-facing orientation (1002-1) (e.g., orientation of pose (706)), the camera (1004) has a forward-facing orientation (1004-1) (e.g., orientation of pose (704)), and the head of the user (1010) (e.g., head pose) has a forward-facing orientation (1006) (e.g., orientation of forward-facing head pose (702)). The forward-facing orientation (1006) corresponds to a region of interest (1012) (e.g., 802). In FIG. 10a, since the orientations (1002-1, 1004-1) of each of the cameras (1002, 1004) roughly coincide with the orientation (1006), if there is no camera occlusion, the image (1018) collectively captured by the cameras (1002, 1004) will depict a relatively large amount of the region of interest (1012). In FIG. 10a, a 3D scene includes an object (1014) located in front of a user (1010) and within a region of interest (1012). In FIG. 10a, the device (1000) receives a natural language request (1016) "What is the price of this?" uttered by the user (1010) because the user (1010) wants to know how much the object (1014) costs. Cameras (1002, 1004) capture an image (1018) associated with the request (1016). For example, the camera (1002) captures a first image, the camera (1004) simultaneously captures a second image, and the device (1000) constructs an image (1018) based on combining the first and second images. Each of the first and second images is captured simultaneously with and / or in response to receiving the request (1016). The image (1018) includes a relatively small occluded area (1020) representing a portion of the 3D scene that is obscured in the image (1018) (e.g., because the camera (1002 and / or 1004) is obscured by the user's (1010) face, the user's (1010) hair, the user's (1010) clothing, stains on the camera (1002 and / or 1004), etc.). The image (1018) further includes an area (1022) (inside the dashed lines) corresponding to a portion of the region of interest (1012) (e.g., corresponding to the overlapping area (908 or 918)). In FIGS. 10a through 10g, the dashed lines within the images (e.g., 1018, 1026, 1046, 1076, or 1093) are for exemplary purposes only and are not included in each image. Due to a relatively small amount of occlusion and because the area (1022) corresponds to a large amount of the area of interest (1012), the image (1018) contains a relatively complete depiction of the object (1014) that the user is asking about. In FIG. 10a, because the image (1018) has a relatively small amount of occlusion and because the image (1018) corresponds to a large portion of the region of interest (1012), in response to receiving a user request (1016), the device (1000) determines a high visibility metric for the region of interest (1012) (e.g., according to the techniques described above in relation to FIG. 9a and 9b). Because the visibility metric is high (e.g., exceeding a threshold), the device (1000) attempts to perform a task to satisfy the request (1016), and the device (1000) provides an audio output (1024). Specifically, a digital assistant (provided, e.g., by a DA unit (350)) processes a request (1016) along with an image (1018) to determine the price of an object (1014), and the device (1000) provides an audio output (1024) “The price of this is $100.” In FIG. 10b, as in FIG. 10a, the camera (1002) has a forward-facing orientation (1002-1) (e.g., the orientation of pose (706)), the camera (1004) has a forward-facing orientation (1004-1) (e.g., the orientation of pose (704)), and the head of the user (1010) (e.g., head pose) has a forward-facing orientation (1006) (e.g., the orientation of forward-facing head pose (702)). The forward-facing orientation (1006) corresponds to a region of interest (1012) (e.g., 802). In FIG. 10b, since the orientations (1002-1, 1004-1) of each of the cameras (1002, 1004) roughly coincide with the orientation (1006), if there is no camera occlusion, the image (1026) collectively captured by the cameras (1002, 1004) will depict a relatively large amount of the region of interest (1012). In FIG. 10b, the 3D scene includes objects (1028, 1030) located in front of the user (1010) and within a region of interest (1012). In FIG. 10b, the device (1000) receives a natural language request (1032) "What is the price of this?" uttered by the user (1010) because the user (1010) wants to know how much the object (1028) costs. The cameras (1002, 1004) capture an image (1026) associated with the request (1032), similar to how the cameras (1002, 1004) capture an image (1018) in FIG. 10a, for example. The image (1026) includes a relatively small occluded area (1034) representing a portion of the 3D scene that is obscured in the image (1026). The image (1026) further includes an area (1036) (inside the dashed lines) corresponding to a portion of the region of interest (1012) (e.g., corresponding to an overlapping area (908 or 918)). Due to the relatively small amount of occlusion and because the area (1036) corresponds to a large portion of the region of interest (1012), the image (1026) includes a relatively complete depiction of objects (1028, 1030) that are of potential user interest. In FIG. 10b, because the image (1026) has a small amount of occlusion and the image (1026) corresponds to a large amount of the region of interest (1012), in response to receiving the request (1032), the device (1000) determines a high visibility metric for the region of interest (1012) (e.g., according to the techniques described above in relation to FIG. 9a and FIG. 9b). Because the visibility metric is high (e.g., exceeding a threshold), the device (1000) attempts to perform a task to satisfy the request (1032). Specifically, the digital assistant attempts to determine a response to the request (1032) by processing the image (1026) along with the request (1032) “What is the price of this?” The digital assistant determines that the image (1026) contains multiple objects (1028, 1030) (e.g., multiple objects are detected within the region of interest (1012)). Accordingly, the device (1000) provides an audio output (1038) “Which object do you mean?” requesting the user (1010) to resolve the ambiguity between the objects (1028, 1030). After the device (1010) provides the audio output (1038), the device (1010) receives a response (1040) “The object on the left” spoken by the user (1010) and resolving the ambiguity between the objects (1028, 1030). In response to receiving the response (1040), the digital assistant processes the request (1032), the response (1040), and the image (1026) together to determine that the price of the object (1028) is $150, and the device (1000) provides the audio output (1042) "The price of the object on the left is $150". In some examples, before the device (1000) receives the request (1032), the device (1000) captures one or more images of a 3D scene, and the device (1000) detects an object (1044) based on the captured image(s). In some examples, in response to receiving the request (1032), the device (1000) provides one or more audio outputs based on the detected object (1044). For example, the one or more audio outputs refer to each position of the objects (1028 and / or 1030) relative to the position of the detected object (1044), e.g., each position determined based on the detection of the objects (1028, 1030, 1044), the head pose of the user (1010), and a map of the 3D scene. As a specific example, if the object (1044) is a green ball, the audio output (1038) is instead "Do you mean the object closer to the green ball or the object further from the green ball?" and / or the audio output (1042) is instead "The price of the object closer to the green ball is $150." In this way, the device (1000) uses the position and / or identity of the previously detected object when providing audio outputs, which can help the user (1010) provide the device (1000) an improved response (e.g., which the device (1000) can interpret more accurately) and help the device (1000) provide the user (1010) improved (more informative and / or unambiguous) audio outputs. In FIG. 10c, as in FIG. 10b, the camera (1002) has a forward-facing orientation (1002-1) (e.g., the orientation of pose (706)), the camera (1004) has a forward-facing orientation (1004-1) (e.g., the orientation of pose (704)), and the head of the user (1010) (e.g., head pose) has a forward-facing orientation (1006) (e.g., the orientation of forward-facing head pose (702)). The forward-facing orientation (1006) corresponds to a region of interest (1012) (e.g., 802). In FIG. 10c, since the orientations (1002-1, 1004-1) of each of the cameras (1002, 1004) roughly coincide with the orientation (1006), if there is no camera occlusion, the image (1046) collectively captured by the cameras (1002, 1004) will depict a relatively large amount of the region of interest (1012). In FIG. 10c, the 3D scene includes an object (1048) located in front of the user (1010) and within a region of interest (1012). In FIG. 10c, the device (1000) receives a natural language request (1050) “Please add this to my shopping list” spoken by the user (1010) because the user (1010) wants to add the object (1048) to their shopping list. The cameras (1002, 1004) capture an image (1046) associated with the request (1050), similar to how the cameras (1002, 1004) capture an image (1018) in FIG. 10a, for example. The image (1046) includes a relatively large occluded area (1052) representing a portion of the 3D scene that is obscured in the image (1046). In FIG. 10c, the image (1046) includes an occluded area (1052) caused by obscuring of the cameras (1002 and / or 1004) by the user's (1010) face, the user's (1010) hair, and / or the user's (1010) clothing. The image (1046) further includes an area (1054) (inside the dashed lines) corresponding to a portion of the region of interest (1012) (e.g., corresponding to an overlapping area (908 or 918)). Even though the image (1046) corresponds to a large portion of the region of interest (1012), due to the large amount of occlusion, the image (1046) does not depict the object (1048) that the user (1010) is burying. In FIG. 10c, even though the image (1046) corresponds to a large portion of the region of interest (1012), due to a large amount of occlusion, in response to receiving a request (1050), the device (1000) determines a low visibility metric for the region of interest (1012) (e.g., according to the techniques described above in relation to FIG. 9a and 9b). Because the visibility metric is low (e.g., below a threshold), the device (1000) provides an audio output (1056) “Which object do you mean?” asking the user (1010) to specify the object (1048) corresponding to the request (1050) (e.g., to specify which object the user (1010) is referring to). After the device (1000) provides an audio output (1056), the device (1000) receives a response (1058) “object in front of me” spoken by the user (1010). In response to receiving the response (1058), the digital assistant processes the response (1058) and the image (1046) to determine whether the image (1046) depicts the object (1048). For example, based on the user (1010)’s head pose, the orientations (1002-1, 1004-1) of the cameras (1002, 1004), and the image (1048), the digital assistant determines whether the object in front of the user (1010) can be detected (e.g., identified) from the image (1048) with sufficient confidence. In FIG. 10c, the digital assistant determines that the image (1046) does not depict the object (1048) (due to image obscuration), and therefore, the device (1000) provides an audio output (1060) “I cannot see the object, please capture an image of the object (1048) using an external device (1062) (Figs. 10d and 10e).” In FIG. 10d, after the device (1000) provides an audio output (1060), the user (1010) holds an external device (1062) and captures an image of a desired object (1048). The external device (1062) implements at least some of the components of the computer system (101), and the external device (1062) includes one or more cameras. FIG. 10d and FIG. 10e illustrate that the external device (1062) is a smartphone, but in other examples, the external device (1062) is a different type of device, such as a laptop computer, a tablet device, or a smart watch. In FIG. 10d, after providing the audio output (1060) (or simultaneously), the external device (1062) displays a camera icon (1066) prompting the user (1010) to activate one or more cameras of the external device (1010). In FIG. 10d, the external device (1062) receives a user input (1064) (e.g., touch input, voice input, gesture input, motion input, gaze input, and / or input received through a peripheral device) that selects the camera icon (1066). In FIG. 10e, in response to receiving user input (1066), the external device (1062) displays a camera user interface (1068). The camera user interface (1068) includes a live view of a 3D scene captured by one or more cameras of the external device (1062) and includes a shutter button (1070) selectable to capture an image. The live view depicts the object (1048) because the user (1010) has captured an image of the object (1048) by directing one or more cameras of the external device (1062) toward the object (1048). In FIG. 10e, the external device (1062) receives user input (1072) (e.g., touch input, voice input, gesture input, motion input, gaze input, and / or input received through a peripheral device) to select the shutter button (1070). In response to receiving user input (1072), the external device (1062) captures an image of the object (1048). Subsequently, the digital assistant processes the image of the object (1048) along with the request (1050) "Please add this to my shopping list" to perform the requested task of adding the object (1048) (e.g., canned sardines) to the user's (1010) shopping list, and the device (1000) provides the audio output (1074) "Okay. I have added the canned sardines to your shopping list." In this way, if the initial image (1046) does not provide sufficient information for the device (1000) to satisfy the user request (1050) (e.g., due to orientations (1002-1 and / or 1004-2) of each of the cameras (1002 and / or 1004) and / or image obscuration), the device (1000) and the external device (1062) can satisfy the user request (1050) without requiring the user (1010) to repeat the user request (1050). Examples of FIGS. 10d and FIGS. 10e illustrate an external device (1062) displaying a camera user interface (1068) in response to receiving a user input (1064) selecting a camera icon (1066), but in other examples, the external device (1062) displays the camera user interface (1068) in response to a different triggering event. For example, the external device (1062) automatically displays the camera user interface (1068) simultaneously with the device (1000) providing the audio output (1060) without user input, or the external device (1062) automatically displays the camera user interface (1068) after the device (1000) provides the audio output (1060) without user input (e.g., thus Fig. 10c proceeds directly to Fig. 10e), or the external device (1062) displays the camera user interface in response to the detection of a motion input corresponding to a lifting motion of the external device (1062) (e.g., a motion associated with taking the external device (1062) out of the user's (1010) pocket or bag) after the device (1000) provides the audio output (1060), or the external device (1062) displays the camera user interface after the device (1000) provides the audio output (1060) and the user of an icon for running a camera application Displays the camera user interface (1068) in response to the selection. In some examples, if the device (1000) determines that the visibility metric does not satisfy the condition (e.g., is below a threshold), the device (1000) provides an audio output (1060) that requests the user (1010) to capture an image of the object (1048) using an external device (1062) without providing an audio output (1056) that requests the user (1010) to specify the object (1048) (e.g., to specify which object the user (1010) is referring to) (and / or without determining whether the image (1046) depicts the object (1048). After the device (1000) has provided the audio output (1060) (or simultaneously with the device (1000) providing the audio output (1060), the external device (1062) displays a camera user interface (1068) according to the techniques discussed above. Accordingly, in some examples, when the device (1000) determines a low visibility metric in response to receiving a user request, the device (1000) directly prompts the user to capture an image of the relevant object without asking the user (1010) to specify the relevant object. In FIG. 10f, the camera (1002) has a downward facing orientation (1002-2) (e.g., the orientation of pose (706)), the camera (1004) has a downward facing orientation (1004-2) (e.g., the orientation of pose (704)), and the head of the user (1010) (e.g., head pose) has a forward facing orientation (1006) (e.g., the orientation of forward facing head pose (702)). The forward facing orientation (1006) corresponds to the region of interest (1012) (e.g., 802). In FIG. 10f, because the orientations (1002-2, 1004-2) of each of the cameras (1002, 1004) do not coincide with the orientation (1006), the image (1076) collectively captured by the cameras (1002, 1004) depicts a relatively small amount of the region of interest (1012) (or does not depict the region of interest (1012) at all). In FIG. 10f, the 3D scene includes an object (1078) that is downward to the user (1010) (e.g., on the floor) and is not within the region of interest (1012). In FIG. 10f, the device (1000) receives a natural language request (1080) "What is the price of this?" uttered by the user (1010) because the user (1010) wants to know how much the object (1078) costs. FIG. 10f illustrates an example where the user (1010) requests to perform a task based on the object (1078) that is not within the region of interest (1012). Specifically, while the user (1010) is looking forward, the user [is on the floor] (e.g., x - z Issue a request (1080) to ask about an object (1078) on the plane (e.g., the user (1010)'s head is facing forward, but the user is looking downward to ask about an object (1078) on the floor). In FIG. 10f, cameras (1002, 1004) capture an image (1076) associated with a request (1080), similar to how cameras (1002, 1004) in FIG. 10a capture an image (1018). The image (1076) includes a relatively small occlusion area (1082) representing a portion of the 3D scene that is obscured in the image (1076). The image (1076) further includes an area (1084) (inside the dashed lines) corresponding to a portion of the region of interest (1012) (e.g., corresponding to an overlapping area (908 or 918)). The area (1084) is relatively small due to the discrepancy between orientations (1002-2, 1006) and between orientations (1004-2, 1006). For example, because the cameras (1002, 1004) are facing downward toward the user (1010), only the upper part of the image (1076) depicts the region of interest (1012) facing forward toward the user (1010). The image (1076) depicts the object (1078) that the user (1010) is asking about. In FIG. 10f, even though the image (1076) contains a small amount of occlusion, because the image (1076) depicts a small portion of the region of interest (1012), in response to receiving a request (1080), the device (1000) determines a low visibility metric for the region of interest (1012) (e.g., according to the techniques described above in relation to FIG. 9a and 9b). Because the visibility metric is low (e.g., below a threshold), the device (1000) provides an audio output (1086) “Which object do you mean?” asking the user (1010) to specify the object (1078) corresponding to the request (1080) (e.g., to specify which object the user (1010) is referring to). After the device (1000) provides an audio output (1086), the device (1000) receives a response (1088) "an object on the floor below me" spoken by the user (1010). In response to receiving the response (1088), the digital assistant processes the response (1088) and the image (1076) to determine whether the image (1076) depicts the object (1078). For example, based on the user (1010)'s head pose, the orientations (1002-2, 1004-2) of the cameras (1002, 1004), and the image (1048), the digital assistant determines whether the object on the floor below the user (1010) can be detected (e.g., identified) from the image (1076) with sufficient confidence. In FIG. 10f, the digital assistant determines that the image (1076) depicts an object (1078) (e.g., with a sufficient amount of confidence). Because the image (1076) depicts the object (1078), the digital assistant processes the image (1076) with the request (1080) “What is the price of this?” to retrieve the price of the object (1078), and the device (1000) provides the audio output (1090) “The price of this object is $68.” In some examples, if the device (1000) determines that the visibility metric does not satisfy the condition (e.g., is below the threshold), the device (1000) provides an audio output (e.g., 1060 in FIG. 10c) requesting the user (1010) to capture an image of the object (1078) using an external device (1062) without providing an audio output (1086) requesting the user (1010) to identify the object (1078) (and / or without determining whether the image (1076) depicts the object (1078). In some examples, the user (1010) then uses the external device (1062) to capture an image of the object (1078) so that the device (1000) provides an audio output (1090) that satisfies the request (1080), similar to what is described in connection with FIG. 10d and FIG. 10e. Accordingly, in some examples, in contrast to that of FIG. 10f, when the visibility metric does not satisfy the condition, the device (1000) prompts the user (1010) to capture an image of the relevant object (1078) using an external device (1062), even though the initial image (1076) can already depict the relevant object (1078). In some examples, before the device (1000) receives a request (1080), the device (1000) captures one or more images of a 3D scene, and the device (1000) detects an object (1091) based on the captured image(s). In some examples, in response to receiving the request (1080), the device (1000) provides one or more audio outputs based on the detected object (1091). For example, one or more audio outputs indicate the position of the object (1078) relative to the position of the detected object (1091). As a specific example, if the object (1091) is a red ball, for instance, similar to how the device (1000) uses the position of the previously detected object (1044) in FIG. 10b, the audio output (1086) is instead "Do you mean the object under the red ball?" and / or the audio output (1090) is instead "The price of the object under the red ball is $68". In FIG. 10g, the camera (1002) has a forward-facing orientation (1002-3) (e.g., orientation of pose (704)), the camera (1004) has a forward-facing orientation (1004-3) (e.g., orientation of pose (706)), and the head of the user (1010) (e.g., head pose) has a forward-facing orientation (1006) (e.g., orientation of forward-facing head pose (702)). The forward-facing orientation (1006) corresponds to a region of interest (1092) (e.g., 808). In FIG. 10g, since the orientations (1002-3, 1004-3) of each of the cameras (1002, 1004) roughly coincide with the orientation (1006), if there is no camera occlusion, the image (1093) collectively captured by the cameras (1002, 1004) will depict a relatively large amount of the region of interest (1092). In FIG. 10g, the 3D scene includes an object (1094) held in the hand of the user (1010) and located within a region of interest (1092) (e.g., a region of interest for handheld objects as described above in relation to FIG. 8b). In FIG. 10g, the device (1000) receives a natural language request (1095) "What is the price of this object in my hand?" uttered by the user (1010) because the user (1010) wants to know how much the object (1094) (held in the user's hand) costs. In FIG. 10g, cameras (1002, 1004) capture an image (1093) associated with a request (1095), similar to how cameras (1002, 1004) in FIG. 10a capture an image (1018). The image (1093) includes a relatively small occluded area (1097) representing a portion of the 3D scene that is obscured in the image (1093). The image (1093) further includes an area (1096) (inside the dashed lines) corresponding to the region of interest (1092). Because of the relatively small amount of occlusion and because the area (1096) corresponds to a large amount of the area of interest (1092) (e.g., all of the area of interest (1092) when the area (1096) is not occluded), the image (1093) contains a relatively complete depiction of the object (1094) in the user's (1010) hand. In FIG. 10g, the device (1000) determines that the request (1095) corresponds to an object held in the user's (1010) hand. Because the request (1095) corresponds to an object held in the user's (1010) hand, the device (1000) determines a visibility metric for a region of interest (1092) (e.g., a region of interest for handheld objects) (e.g., a region of interest (808) as described in relation to FIG. 8b). In contrast, in FIGS. 10a through 10f, since the device (1000) has not determined that each user request (e.g., 1016, 1032, 1050, 1080) corresponds to an object held in the hand of the user (1010), the device (1000) instead determines a visibility metric for a different region of interest (1012) (e.g., a region of interest (802) as described in relation to FIG. 8a). In FIG. 10g, because the image (1093) has a relatively small amount of occlusion and because the image (1093) corresponds to a large portion of the region of interest (1092), in response to receiving a user request (1095), the device (1000) determines a high visibility metric for the region of interest (1092) (e.g., according to the techniques described above in relation to FIG. 9a and 9b). Because the visibility metric is high (e.g., exceeding a threshold), the device (1000) attempts to perform a task to satisfy the request (1095), and the device (1000) provides an audio output (1098). Specifically, the digital assistant processes the request (1095) along with the image (1093) to determine the price of the object (1094), and the device (1000) provides the audio output (1098) "The price of this is $1,000." Further descriptions regarding FIGS. 7, FIGS. 8a and 8b, FIGS. 9a and 9b, and FIGS. 10a through 10g are provided below with reference to the method (1100) described below in relation to FIG. 11. FIG. 11 is a flowchart of a method (1100) for providing audio outputs in response to natural language input, according to some examples. In some examples, the method (1100) is performed in a computer system (e.g., device (1000)) that communicates with one or more visual imaging sensors (e.g., cameras, e.g., RGB cameras, infrared cameras, and / or depth cameras) and one or more audio output devices (e.g., speakers). In some examples, the method (1100) is controlled by instructions stored in a non-transient (or transient) computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processing units (302) of a computer system (101) (e.g., controller (110) of FIG. 1). In some examples, the operations of the method (1100) are distributed across a number of computer systems, such as a computer system and a separate server system. Some of the operations of the method (1100) are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted. The method (1100) includes the step (1102) of receiving a natural language input (e.g., 1016, 1032, 1050, 1080, or 1095) corresponding to a first object (e.g., 1014, 1028, 1048, 1078, or 1094) in a 3D scene while the head of a user of a computer system (e.g., 1010) has a head pose (e.g., 702) corresponding to a forward-facing area (e.g., 802, 808, 906, 916, 1012, or 1092) of a 3D scene (e.g., a forward-facing area determined by an area tracking unit (380)). The method (1100) includes the step (1104) of capturing image data (e.g., 1018, 1026, 1046, 1076, or 1093) associated with a natural language input corresponding to a first object in a 3D scene through one or more visual imaging sensors. The method (1100) comprises, in response to receiving natural language input corresponding to a first object in a 3D scene (1106): providing a first audio output (e.g., 1024, 1038, or 1098) corresponding to the first object through one or more audio output devices according to a determination that a visibility metric (e.g., determined by a visibility analysis unit (390)) representing the amount of a forward-facing area of the 3D scene described by image data satisfies a condition (e.g., exceeds a threshold); and providing a second audio output (e.g., 1056 or 1086) corresponding to the first object (e.g., different from the first audio output) through one or more audio output devices according to a determination that a visibility metric (e.g., is less than a threshold) representing the amount of a forward-facing area of the 3D scene described by image data does not satisfy a condition (e.g., is less than a threshold) (1110). In some examples, depending on the determination that the head pose is the first head pose, the forward-facing area of the 3D scene is the first area of the 3D scene; depending on the determination that the head pose is the second head pose different from the first head pose, the forward-facing area of the 3D scene is the second area of the 3D scene different from the first area of the 3D scene. In some examples, the step of capturing image data associated with natural language input through one or more visual imaging sensors includes the step of capturing image data associated with natural language input through one or more visual imaging sensors while receiving natural language input. In some examples, the step of capturing image data associated with natural language input through one or more visual imaging sensors includes the step of capturing image data associated with natural language input through one or more visual imaging sensors in response to receiving natural language input. In some examples, a first audio output (e.g., 1024 or 1098) corresponding to a first object displays the result of a first task (e.g., a result satisfying a user request included in the natural language input) performed (e.g., by a digital assistant) based on the natural language input and the first object. In some examples, the step of providing a first audio output corresponding to a first object through one or more audio output devices includes providing an audio output (e.g., 1038) that requests a user to resolve ambiguity between a plurality of detected objects based on a determination that a forward-facing area of a 3D scene includes a plurality of detected objects (e.g., 1028, 1030) including the first object (e.g., 1028). In some examples, a second audio output (e.g., 1056) corresponding to a first object (e.g., 1048) includes a request to the user to specify the first object (e.g., to specify the attributes of the first object (e.g., identity, shape, size, color, orientation, and location)). In some examples, the second audio output corresponds to a request (e.g., 1060) to the user to capture an image of the first object using an external device (e.g., 1062). In some examples, after the external device has captured an image of the first object, the computer system provides an audio output (e.g., 1074) through one or more audio output devices that indicates a task to be performed based on natural language input and the captured image of the first object. In some examples, since the computer system provides audio output without receiving any additional natural language input after receiving natural language input, the user does not need to repeat their initial natural language request for the requested task related to the object to be performed and for the results of the requested task to be output through one or more audio output devices. In some examples, the computer system comprises a first device (e.g., 1002) and a second device (e.g., 1004), wherein the first device is different from the second device; natural language input is received while the first device is worn by a user and while the second device is worn by a user; and the forward-facing area of the 3D scene is determined based on the position of the first device while the first device is worn by a user, the orientation of the first device while the first device is worn by a user (e.g., 1002-1, 1002-2, or 1002-3), the position of the second device while the second device is worn by a user, and the orientation of the second device while the second device is worn by a user (e.g., 1004-1, 1004-2, or 1004-3). In some examples, one or more visual imaging sensors include a first visual imaging sensor (e.g., 1002) and a second visual imaging sensor (e.g., 1004) different from the first visual imaging sensor; and the step of capturing image data (e.g., 1018, 1026, 1046, 1076, or 1093) associated with natural language input through one or more visual imaging sensors includes: capturing first image data through the first visual imaging sensor; and capturing second image data different from the first image data through the second visual imaging sensor (e.g., capturing the first image data and the second image data simultaneously). In some examples, the first image data is captured while the first visual imaging sensor is worn on a first side of the user's head (e.g., left or right) (e.g., while a device including the first visual imaging sensor is worn (e.g., while the device is inserted into the ear)); and the second image data is captured while the second visual imaging sensor is worn on a second side of the user's head (e.g., left or right) (e.g., while a device including the second visual imaging sensor is worn (e.g., while the device is inserted into the ear)), wherein the first side of the user's head is opposite to the second side of the user's head. In some examples, a visibility metric representing the amount of a forward-facing area (e.g., 1012) of a 3D scene depicted by image data (e.g., 1018) has a first value (as described in relation to FIG. 10a), depending on the determination that the first image data is captured while the first visual imaging sensor has a first orientation (e.g., 1002-1 in FIG. 10a) (e.g., with respect to the device including the second visual imaging sensor and / or with respect to the user's head) and the second image data is captured while the second visual imaging sensor has a second orientation (e.g., 1004-1 in FIG. 10a); According to the determination that the first image data is captured while the first visual imaging sensor has a third orientation (e.g., 1002-2 in FIG. 10f) that is different from the first orientation (e.g., with respect to a device including the first visual imaging sensor and / or with respect to the user's head) and the determination that the second image data is captured while the second visual imaging sensor has a fourth orientation (e.g., 1004-2 in FIG. 10f) that is different from the second orientation (e.g., with respect to a device including the second visual imaging sensor and / or with respect to the user's head), a visibility metric representing the amount of a forward-facing area (e.g., 1012) of a 3D scene depicted by the image data (e.g., 1076) has a second value that is different from the first value (e.g., as described in relation to FIG. 10f) (e.g., the visibility metric is for each orientation of the first visual imaging sensor when the first image data is captured and when the second image data is captured (depends on the orientation of each of the second visual imaging sensors). In some examples, depending on the determination that the first image data and the second image data describe (e.g., collectively) a first amount of the forward-facing area of the 3D scene (e.g., 1012), the visibility metric has a first value (e.g., as described in connection with FIG. 10f); depending on the determination that the first image data and the second image data describe (e.g., collectively) a second amount of the forward-facing area of the 3D scene that is larger than the first amount of the forward-facing area of the 3D scene, the visibility metric has a second value that is larger than the first value (e.g., as described in connection with FIG. 10a). In some examples, image data includes image regions (e.g., pixels) (e.g., 904, 914, 1020, 1034, 1052, 1082, or 1097) representing the occlusion of the foreground area of a 3D scene, wherein a visibility metric representing the amount of the foreground area of a 3D scene depicted by the image data is based on the image regions representing the occlusion of the foreground area of a 3D scene. In some examples, depending on the determination that an image area (e.g., 1020) representing the occlusion of the foreground area of a 3D scene has a first size, a visibility metric representing the amount of the foreground area of a 3D scene depicted by the image data has a third value (e.g., as described in connection with FIG. 10a); depending on the determination that an image area (e.g., 1052) representing the occlusion of the foreground area of a 3D scene has a second size larger than the first size, a visibility metric representing the amount of the foreground area of a 3D scene depicted by the image data has a fourth value smaller than the third value (e.g., as described in connection with FIG. 10c). In some examples, the visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data (e.g., 802, 808, 906, 916, 1012, or 1092) is based on the amount of overlap between the unobstructed area of the 3D scene depicted by the image data and the forward-facing area of the 3D scene (e.g., as represented by the visibility area (910 or 920)) (e.g., so that more overlap results in a higher visibility metric and less overlap results in a lower visibility metric). In some examples, the foreground area of a 3D scene has predefined (e.g., fixed) dimensions (e.g., length, width, depth, area, and / or volume) (e.g., dimensions that do not depend on the user's head pose) (e.g., dimensions determined before natural language input is received and image data is captured). In some examples, the forward-facing area of the 3D scene is at least a predefined (e.g., fixed) distance (e.g., 804) from the user's head (e.g., so that the part of the forward-facing area of the 3D scene closest to the user's head is at least a predefined non-zero distance from the user's head). In some examples, a handheld object area is determined; the handheld object area is the area where each user holds each object in their hand while each user issues a query regarding each object; and the foreground area of the 3D scene (e.g., 808 or 1092) is determined based on the handheld object area. In some examples, depending on the determination that the natural language input corresponding to the first object is a natural language input of type 1 (e.g., 1016, 1032, 1050, or 1080), the forward-facing area of the 3D scene is a third area of the 3D scene (e.g., 802 or 1012); depending on the determination that the natural language input corresponding to the first object (e.g., 1095) is a natural language input of type 2 different from the natural language input of type 1, the forward-facing area of the 3D scene is a fourth area of the 3D scene different from the third area of the 3D scene (e.g., 808 or 1092). In some examples, the method (1100) provides, after providing a second audio output (e.g., 1056 or 1086) corresponding to a first object through one or more audio output devices: receiving a user input (e.g., 1058 or 1088) corresponding to a first object (e.g., voice input, gaze input, and / or gesture input) (e.g., user input responding to the second audio output and identifying the first object and / or identifying the location of the first object); In response to receiving user input corresponding to a first object: providing a third audio output (e.g., 1090) that displays the result of a second task performed (e.g., by a digital assistant) based on natural language input and the first object through one or more audio output devices, according to a determination that image data (e.g., 1076) satisfies a predetermined condition in relation to the first object (e.g., 1078) (e.g., that the image data is determined to describe the first object with at least a threshold amount of confidence); The method further includes the step of providing a fourth audio output (e.g., 1060) that requests a user to use an external device (e.g., 1062) (e.g., to capture an image of the first object using an external device) through one or more audio output devices, based on a user input (e.g., 1058) corresponding to the first object, and a determination that the image data (e.g., 1046) does not satisfy a predetermined condition in relation to the first object (e.g., that the image data is not determined to describe the first object with at least a threshold amount of confidence). In some examples, after providing a fourth audio output requesting a user to use an external device through one or more audio output devices, the external device displays a camera user interface (e.g., 1068), and the external device captures an image of a first object (e.g., 1048) while displaying the camera user interface. In some examples, the method (1100) includes the step of, after the external device has captured an image of the first object while displaying the camera user interface (e.g., in response to receiving user input (1072), providing a fifth audio output (e.g., 1074) through one or more audio output devices that displays the result of a third task performed (e.g., by a digital assistant) based on the image of the first object and natural language input (e.g., 1050). In some examples, the external device displays a camera user interface in response to a selection (e.g., 1064) of a user interface element (e.g., 1066) displayed by the external device. In some examples, the method (1100) includes the step of capturing third image data representing a 3D scene through one or more visual imaging sensors before receiving natural language input (e.g., 1032 or 1080), wherein the first audio output is based on a second object (e.g., 1044) detected based on the third image data representing the 3D scene (e.g., as described in connection with FIG. 10b); and / or the second audio output is based on a second object (e.g., 1091) detected based on the third image data representing the 3D scene (e.g., as described in connection with FIG. 10f). The foregoing description has been described with reference to specific embodiments for the purpose of explanation. However, the exemplary discussions above are not intended to limit the invention to the exact forms disclosed or to be complete. Many modifications and variations are possible in light of the teachings above. The embodiments have been selected and described to best illustrate the principles of the invention and its practical applications, thereby enabling those skilled in the art to best use the invention and the various described embodiments with various modifications suitable for the specific use being considered. As described above, one aspect of the present technology is the collection and use of data available from various sources to facilitate user interactions with a three-dimensional scene. The present disclosure considers that, in some cases, such collected data may include personal data that can be used to uniquely identify a specific individual or to contact him / her or to determine his / her location. Such personal data may include demographic data, location-based data, telephone numbers, email addresses, Twitter IDs, home addresses, data or records regarding a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information. The present disclosure recognizes that the use of such personal data in the present technology may be used to benefit users. For example, personal data may be used to output spoken responses to assist users. Additionally, other uses of personal data that benefit users are also considered by the present disclosure. For example, health and fitness data may be used to provide insights into a user's general wellness, or may be used as positive feedback to individuals using technology to pursue wellness goals. The present disclosure considers that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal data will comply with well-established privacy policies and / or privacy practices. In particular, such entities must implement and consistently use privacy policies and practices that are recognized as meeting or exceeding industrial or administrative requirements for keeping personal data private and secure. Such policies must be readily accessible to users and updated as the collection and / or use of data changes. Personal data from users must be collected for the entity's lawful and proper use and must not be shared or sold outside of these lawful uses. Additionally, such collection / sharing must occur after the users' notified consent has been received. Furthermore, such entities should consider taking any necessary steps to protect and secure access to such personal data and to ensure that others with access to the personal data adhere to their privacy policies and procedures. Additionally, such entities may be evaluated by third parties to demonstrate their adherence to widely recognized privacy policies and practices. Furthermore, policies and practices must be adapted to the specific types of personal data being collected and / or accessed, and to applicable laws and standards that include jurisdiction-specific considerations. For example, in the United States, the collection or access to certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); whereas in other countries, health data may be subject to and must be handled by other statutes and policies.Therefore, different privacy practices must be maintained for different types of personal data in each country. Notwithstanding the foregoing, the present disclosure also considers embodiments in which users selectively block the use of or access to personal information data. That is, the present disclosure considers that hardware and / or software elements may be provided to prevent or block access to such personal information data. For example, in the case of outputting spoken responses to a user, the present technology may be configured to allow users to select "Agree" or "Disagree" regarding participation in the collection of personal information data during or at any time thereafter registration for services. In another example, users may choose not to provide personal information data that causes spoken responses to be generated. In yet another example, users may choose to limit the length of time such data is retained. In addition to providing "Agree" and "Disagree" options, the present disclosure considers providing notices regarding the access or use of personal information. For example, a user may be notified when downloading an app in which their personal information data will be accessed, and then reminded again immediately before the personal information data is accessed by the app. Furthermore, it is the intent of the present disclosure that personal data should be managed and processed in a manner that minimizes the risks of unintended or unauthorized access or use. Risks can be minimized by limiting the collection of data and by deleting data when it is no longer needed. Additionally, and where applicable, including those applicable to certain health-related applications, data deidentification may be used to protect user privacy. Where appropriate, deidentification may be facilitated by removing specific identifiers (e.g., date of birth, etc.), by controlling the amount or specificity of stored data (e.g., by collecting location data at the city level rather than the address level), by controlling the way data is stored (e.g., by aggregating data across users), and / or by other methods. Accordingly, while the present disclosure extensively covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also takes into account that the various embodiments may also be implemented without the need to access such personal information data. That is, the various embodiments of the present technology are not rendered inoperable by the absence of all or part of such personal information data. For example, spoken responses may be generated based on non-personal information data, such as content requested by a device associated with a user, other non-personal information available to the service, or publicly available information, or the minimum amount of personal information required.
Claims
Claim 1 As a method, in a first computer system communicating with one or more image sensors, A step of acquiring a first image using the above one or more image sensors; A step of receiving a user request related to the first image above; and In response to acquiring the first image and receiving the user request: A step of causing the second computer system to provide a prompt to capture the second image to the second computer system based on a determination that the quality of the first image does not satisfy a quality standard; and Based on the determination that the quality of the first image satisfies the quality standard: A step of generating a response to the user request based on the first image above; and A method comprising the step of providing an output including the response to the user request based on the first image. Claim 2 A method according to claim 1, wherein the first computer system is a head-mounted electronic device and the second computer system is a smartphone. Claim 3 A method according to claim 1 or 2, wherein the image sensor of the first computer system has a lower quality metric than the quality metric of the image sensor of the second computer system. Claim 4 In any one of paragraphs 1 to 3, in response to detecting the first image and receiving the user request: A step of selecting a first quality standard as a quality standard based on a determination that the above user request includes a request of the first type; and A method further comprising the step of selecting a second quality standard different from the first quality standard as the quality standard, based on a determination that the user request includes a second type of request different from the first type. Claim 5 In any one of paragraphs 1 through 4, after the second computer system has provided a prompt to capture the second image to the second computer system: A step of detecting user input to capture the second image using the second computer system; A step of generating a response to the user request based on the second image above; and A method further comprising the step of providing an output including the response to the user request based on the second image. Claim 6 A method according to any one of claims 1 to 5, wherein the determination of whether the quality of the first image satisfies the quality standard comprises providing a prompt to a large language model (LLM), and the prompt comprises a request for whether the first image is of sufficient quality to complete a task determined from the user request. Claim 7 A method according to any one of claims 1 to 5, wherein the determination of whether the quality of the first image satisfies the quality standard comprises: generating an embedding of the first image; and comparing the embedding of the first image with a learned set of embeddings representing a high-quality image or a low-quality image. Claim 8 A method according to claim 7, further comprising: a step of selecting a first learned embedding set as the learned embedding set based on a determination that the user request is a first type of request; and a step of selecting a second learned embedding set as the learned embedding set based on a determination that the user request is a second type of request. Claim 9 In any one of claims 1 to 8, in response to detecting the first image and receiving the user request: Based on the determination that the context of the first computer system indicates that the first image does not satisfy the quality criteria: A step of abandoning the determination of whether the first image satisfies the quality criteria; and A method further comprising the step of causing the second computer system to provide the prompt to capture the second image to the second computer system. Claim 10 In claim 9, the determination that the context of the first computer system indicates that the first image does not satisfy the quality standard includes the determination that the first computer system is moving. Claim 11 A method according to claim 9 or 10, wherein the determination that the context of the first computer system indicates that the first image does not satisfy the quality criteria includes a determination that the lighting level of the environment of the first computer system is below a lighting threshold. Claim 12 A method according to any one of claims 9 to 11, wherein the determination that the context of the first computer system indicates that the first image does not satisfy the quality standard includes the determination that the one or more image sensors are obscured. Claim 13 A method according to any one of claims 9 through 12, wherein the determination that the context of the first computer system indicates that the first image does not satisfy the quality standard includes the determination that the field of view of the one or more image sensors includes text. Claim 14 A method according to any one of claims 1 to 13, further comprising the step of displaying a camera user interface as a display generating component communicating with a second computer system, based on a determination that the quality of the first image does not satisfy the quality standard. Claim 15 In paragraph 14, after the camera user interface is displayed as a display generating component communicating with the second computer system: A step of detecting user input to capture the second image; and A method further comprising the step of stopping the display of the camera user interface to the display generating component communicating with the second computer system in response to detecting the user input to capture the second image. Claim 16 In claim 14, the method wherein the camera user interface is displayed on the lock screen by the display generating component communicating with the second computer system. Claim 17 A method according to any one of claims 1 to 16, wherein the second computer system is not physically connected to the first computer system. Claim 18 A method according to any one of claims 1 to 16, wherein the second computer system is physically connected to the first computer system by a wire, and the second computer system and the first computer system are not located within the same housing. Claim 19 A non-transient computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a first computer system communicating with one or more image sensors, wherein the one or more programs include instructions for performing the method of any one of claims 1 to 16. Claim 20 A first computer system configured to communicate with one or more image sensors, comprising: one or more processors; and one or more memories for storing one or more programs configured to be executed by the one or more processors, wherein the one or more programs comprise instructions for performing the method of any one of claims 1 to 16. Claim 21 A first computer system configured to communicate with one or more image sensors, comprising means for performing the method of any one of claims 1 to 16. Claim 22 A computer program product comprising one or more programs configured to be executed by one or more processors of a first computer system communicating with one or more image sensors, wherein the one or more programs comprise instructions for performing the method of any one of claims 1 to 16. Claim 23 A non-transient computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a first computer system communicating with one or more image sensors, wherein the one or more programs use the one or more image sensors to acquire a first image; receive a user request related to the first image; and in response to acquiring the first image and receiving the user request: Based on the determination that the quality of the first image does not satisfy the quality standard, the second computer system is made to provide a prompt to the second computer system to capture the second image; Based on the determination that the quality of the first image satisfies the quality standard: Generate a response to the user request based on the first image above; A non-transient computer-readable storage medium comprising instructions for providing an output including the response to the user request based on the first image. Claim 24 A computer system configured to communicate with one or more image sensors, wherein the one or more computer systems comprise: one or more processors; and one or more memories for storing one or more programs configured to be executed by the one or more processors, and the one or more programs are Using the above one or more image sensors, a first image is acquired; Receiving a user request related to the first image above; In response to acquiring the first image and receiving the user request: Based on the determination that the quality of the first image does not satisfy the quality standard, the second computer system is made to provide a prompt to the second computer system to capture the second image; Based on the determination that the quality of the first image satisfies the quality standard: Generate a response to the user request based on the first image above; A computer system comprising instructions for providing an output including the response to the user request based on the first image. Claim 25 A computer system configured to communicate with one or more image sensors, wherein the computer system comprises: means for acquiring a first image using the one or more image sensors; means for receiving a user request related to the first image; and in response to acquiring the first image and receiving the user request: Based on the determination that the quality of the first image does not satisfy the quality standard, the second computer system is made to provide a prompt to the second computer system to capture the second image; Based on the determination that the quality of the first image satisfies the quality standard: Generate a response to the user request based on the first image above; A computer system comprising means for providing an output including the response to the user request based on the first image. Claim 26 A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more image sensors, wherein the one or more programs use the one or more image sensors to acquire a first image; receive a user request related to the first image; and in response to acquiring the first image and receiving the user request: Based on the determination that the quality of the first image does not satisfy the quality standard, the second computer system is made to provide a prompt to the second computer system to capture the second image; Based on the determination that the quality of the first image satisfies the quality standard: Generate a response to the user request based on the first image above; A computer program product comprising instructions for providing an output including the response to the user request based on the first image. Claim 27 As a method, in a computer system communicating with one or more visual imaging sensors and one or more audio output devices: receiving a natural language input corresponding to a first object in a three-dimensional (3D) scene while the head of a user of the computer system has a head pose corresponding to a forward-facing region of a three-dimensional (3D) scene; capturing image data associated with the natural language input corresponding to the first object in the 3D scene through the one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: A step of providing a first audio output corresponding to the first object through one or more audio output devices, in accordance with a determination that a visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data satisfies a condition; and A method comprising the step of providing a second audio output corresponding to the first object through one or more audio output devices, in accordance with a determination that the visibility metric representing the amount of the forward-facing area of the 3D scene described by the image data does not satisfy the condition. Claim 28 A method according to claim 27, wherein, in accordance with the determination that the head pose is a first head pose, the forward-facing area of the 3D scene is a first area of the 3D scene; and in accordance with the determination that the head pose is a second head pose different from the first head pose, the forward-facing area of the 3D scene is a second area of the 3D scene different from the first area of the 3D scene. Claim 29 A method according to claim 27 or 28, wherein the step of capturing the image data associated with the natural language input through the one or more visual imaging sensors comprises the step of capturing the image data associated with the natural language input through the one or more visual imaging sensors while receiving the natural language input. Claim 30 A method according to any one of claims 27 to 29, wherein the step of capturing the image data associated with the natural language input through the one or more visual imaging sensors comprises the step of capturing the image data associated with the natural language input through the one or more visual imaging sensors in response to receiving the natural language input. Claim 31 A method according to any one of claims 27 to 30, wherein the first audio output corresponding to the first object displays the result of a first task performed based on the natural language input and the first object. Claim 32 A method according to claim 31, wherein the step of providing the first audio output corresponding to the first object through the one or more audio output devices comprises the step of providing an audio output that requests the user to resolve ambiguity between the plurality of detected objects, based on a determination that the forward-facing area of the 3D scene includes a plurality of detected objects including the first object. Claim 33 A method according to any one of claims 27 to 32, wherein the second audio output corresponding to the first object includes a request to the user to specify the first object. Claim 34 A method according to any one of claims 27 to 33, wherein the computer system comprises a first device and a second device, the first device being different from the second device; the natural language input being received while the first device is worn by the user and while the second device is worn by the user; and the forward-facing area of the 3D scene being determined based on the position of the first device while the first device is worn by the user, the orientation of the first device while the first device is worn by the user, the position of the second device while the second device is worn by the user, and the orientation of the second device while the second device is worn by the user. Claim 35 In any one of claims 27 to 34, the one or more visual imaging sensors include a first visual imaging sensor and a second visual imaging sensor different from the first visual imaging sensor; and the step of capturing the image data associated with the natural language input through the one or more visual imaging sensors is A step of capturing first image data through the first visual imaging sensor; and A method comprising the step of capturing second image data different from the first image data through the second visual imaging sensor. Claim 36 In paragraph 35, the first image data is captured while the first visual imaging sensor is worn on the first side of the user's head; the second image data is captured while the second visual imaging sensor is worn on the second side of the user's head, and the first side of the user's head is opposite to the second side of the user's head, method. Claim 37 A method according to claim 35 or 36, wherein, in accordance with a determination that the first image data is captured while the first visual imaging sensor has a first orientation and a determination that the second image data is captured while the second visual imaging sensor has a second orientation, the visibility metric expressing the amount of the forward-facing area of the 3D scene depicted by the image data has a first value; and in accordance with a determination that the first image data is captured while the first visual imaging sensor has a third orientation different from the first orientation and a determination that the second image data is captured while the second visual imaging sensor has a fourth orientation different from the second orientation, the visibility metric expressing the amount of the forward-facing area of the 3D scene depicted by the image data has a second value different from the first value. Claim 38 A method according to any one of claims 35 to 37, wherein, in accordance with the determination that the first image data and the second image data describe a first amount of the forward-facing area of the 3D scene, the visibility metric has a first value; and in accordance with the determination that the first image data and the second image data describe a second amount of the forward-facing area of the 3D scene that is greater than the first amount of the forward-facing area of the 3D scene, the visibility metric has a second value greater than the first value. Claim 39 A method according to any one of claims 27 to 38, wherein the image data includes an image region representing the occlusion of the forward-facing region of the 3D scene, and the visibility metric representing the amount of the forward-facing region of the 3D scene depicted by the image data is based on the image region representing the occlusion of the forward-facing region of the 3D scene. Claim 40 A method according to claim 39, wherein, in accordance with a determination that the image area representing the occlusion of the forward-facing area of the 3D scene has a first size, the visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data has a third value; and in accordance with a determination that the image area representing the occlusion of the forward-facing area of the 3D scene has a second size larger than the first size, the visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data has a fourth value smaller than the third value. Claim 41 A method according to any one of claims 27 to 40, wherein the visibility metric expressing the amount of the forward-facing region of the 3D scene depicted by the image data is based on the amount of overlap between the non-occluded region of the 3D scene depicted by the image data and the forward-facing region of the 3D scene. Claim 42 A method according to any one of claims 27 to 41, wherein the forward-facing area of the 3D scene has predefined dimensions. Claim 43 A method according to any one of claims 27 to 42, wherein the forward-facing area of the 3D scene is located at least a predetermined distance from the head of the user. Claim 44 A method according to any one of claims 27 to 43, wherein a handheld object area is determined; said handheld object area is an area in which said user holds said object in their hand while saying user issues a query regarding said object; and said forward-facing area of said 3D scene is determined based on said handheld object area. Claim 45 A method according to any one of claims 27 to 44, wherein, in accordance with a determination that the natural language input corresponding to the first object is a natural language input of the first type, the forward-facing area of the 3D scene is a third area of the 3D scene; and in accordance with a determination that the natural language input corresponding to the first object is a natural language input of the second type different from the natural language input of the first type, the forward-facing area of the 3D scene is a fourth area of the 3D scene different from the third area of the 3D scene. Claim 46 In any one of claims 27 to 45, after providing the second audio output corresponding to the first object through the one or more audio output devices: A step of receiving user input corresponding to the first object; and In response to receiving the user input corresponding to the first object: A step of providing a third audio output that displays the result of a second task performed based on the natural language input and the first object through the one or more audio output devices, based on a determination that the image data satisfies a predetermined condition in relation to the first object based on the user input corresponding to the first object; and A method further comprising the step of providing a fourth audio output requesting the user to use an external device through one or more audio output devices, based on the user input corresponding to the first object, and according to the determination that the image data does not satisfy the predetermined condition in relation to the first object. Claim 47 In claim 46, after providing the fourth audio output requesting the user to use the external device through the one or more audio output devices, the external device displays a camera user interface, and while the external device displays the camera user interface, the external device captures an image of the first object, and the method further comprises the step of providing a fifth audio output through the one or more audio output devices that displays the result of a third task performed based on the image of the first object and the natural language input, after the external device captures the image of the first object while the external device displays the camera user interface. Claim 48 In claim 47, the method wherein the external device displays the camera user interface in response to the selection of a user interface element displayed by the external device. Claim 49 In any one of claims 27 to 48, the method further comprises the step of capturing third image data representing the 3D scene through the one or more visual imaging sensors before receiving the natural language input, and The first audio output is based on a second object detected based on the third image data representing the 3D scene; and / or A method in which the second audio output is based on the second object detected based on the third image data representing the 3D scene. Claim 50 A non-transient computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs include instructions for performing the method of any one of claims 27 to 49. Claim 51 A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, comprising: one or more processors; and one or more memories for storing one or more programs configured to be executed by the one or more processors, wherein the one or more programs comprise instructions for performing the method of any one of claims 27 to 49. Claim 52 A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, comprising means for performing the method of any one of claims 27 to 49. Claim 53 A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs comprise instructions for performing the method of any one of claims 27 to 49. Claim 54 A non-transient computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs receive natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing area of a 3D scene; capture image data associated with the natural language input corresponding to the first object in the 3D scene through the one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: Based on the determination that a visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data satisfies the condition, a first audio output corresponding to the first object is provided through the one or more audio output devices; A non-transient computer-readable storage medium comprising instructions for providing a second audio output corresponding to the first object through one or more audio output devices, in accordance with a determination that the visibility metric representing the amount of the forward-facing area of the 3D scene described by the image data does not satisfy the condition. Claim 55 A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, comprising: one or more processors; and one or more memories for storing one or more programs configured to be executed by the one or more processors, wherein the one or more programs are, While the head of the user of the above computer system has a head pose corresponding to a forward-facing area of a three-dimensional (3D) scene, natural language input corresponding to a first object in the 3D scene is received; Through the one or more visual imaging sensors, image data associated with the natural language input corresponding to the first object in the 3D scene is captured; In response to receiving the natural language input corresponding to the first object within the 3D scene: Based on the determination that a visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data satisfies the condition, a first audio output corresponding to the first object is provided through the one or more audio output devices; A computer system comprising instructions for providing a second audio output corresponding to the first object through one or more audio output devices, in accordance with a determination that the visibility metric representing the amount of the forward-facing area of the 3D scene described by the image data does not satisfy the condition. Claim 56 A computer system configured to communicate with one or more visual imaging sensors and one or more audio output devices, wherein means for receiving natural language input corresponding to a first object in a 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing area of a 3D scene; means for capturing image data associated with the natural language input corresponding to the first object in the 3D scene through the one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: Based on the determination that a visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data satisfies the condition, a first audio output corresponding to the first object is provided through the one or more audio output devices; A computer system comprising means for providing a second audio output corresponding to the first object through one or more audio output devices, in accordance with a determination that the visibility metric representing the amount of the forward-facing area of the 3D scene described by the image data does not satisfy the condition. Claim 57 A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more visual imaging sensors and one or more audio output devices, wherein the one or more programs receive natural language input corresponding to a first object in the 3D scene while the head of a user of the computer system has a head pose corresponding to a forward-facing area of the 3D scene; capture image data associated with the natural language input corresponding to the first object in the 3D scene through the one or more visual imaging sensors; and in response to receiving the natural language input corresponding to the first object in the 3D scene: Based on the determination that a visibility metric representing the amount of the forward-facing area of the 3D scene depicted by the image data satisfies the condition, a first audio output corresponding to the first object is provided through the one or more audio output devices; A computer program product comprising instructions for providing a second audio output corresponding to the first object through one or more audio output devices, in accordance with a determination that the visibility metric representing the amount of the forward-facing area of the 3D scene described by the image data does not satisfy the condition.