Digital assistant object placement

The placement of a digital assistant object outside the user's field of view and animating it into view effectively reduces interaction time and power usage.

JP2025178252APending Publication Date: 2025-12-05APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025141089
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-23
Filing Date
2025-08-27
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Interacting with a digital assistant within a computer-generated reality (CGR) environment can be difficult, confusing, or distracting from the immersion of the CGR environment, due to the complexity of the interface.

Method used

A digital assistant object is placed outside the user's field of view and indicated with visual, auditory, or tactile outputs to attract attention, reducing the time and input required to access functions, thus improving battery life.

Benefits of technology

The placement of a digital assistant object outside the user's field of view and animated into view effectively reduces interaction time and power usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025178252000001_ABST
    Figure 2025178252000001_ABST
Patent Text Reader

Abstract

To provide technology relating to placing an object representing a digital assistant in a computer-generated reality (CGR) environment.SOLUTION: A system and a process are provided for operating an intelligent automated assistant in a computer-generated reality (CGR) environment. For example, a user input invoking a digital assistant session is received, and in response thereto, the digital assistant session is initiated. Initiating the digital assistant session includes placing a digital assistant object at a first location within the CGR environment but outside a currently displayed portion of the CGR environment at a first time, and providing a first output indicating the location of the digital assistant object.SELECTED DRAWING: Figure 2B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This relates generally to digital assistants, and more specifically to placing objects representing digital assistants within a computer-generated reality (CGR) environment. [Background technology]

[0002] A digital assistant can act as a useful interface between a human user and their electronic device, for example, using spoken or typed natural language, gestures, or other convenient or intuitive input modes. For example, a user can utter a natural language request to a digital assistant on an electronic device. The digital assistant can interpret the user's intent from the spoken input and act on the user's intent into a task. The task can then be performed by executing one or more services on the electronic device and can return an associated output response to the user request to the user.

[0003] Unlike the physical world, which a person can interact with and perceive without using electronic devices, electronic devices are used to interact with and / or perceive a wholly or partially simulated computer-generated reality (CGR) environment. For example, a CGR environment may include mixed reality (MR) content, augmented reality (AR) content, virtual reality (VR) content, etc. One way to interact with a CGR system is by tracking some of the person's physical movements and adjusting the properties of simulated elements in the CGR environment accordingly in a manner that appears to obey at least one law of physics. For example, as a user moves a device presenting the CGR environment and / or the user's head, the CGR system can detect the movement and adjust graphical and auditory content according to the user's viewpoint to create a spatial sound effect. In some situations, the CGR system can adjust the properties of the CGR content in response to user input, such as button inputs or voice commands.

[0004] Many different electronic devices and / or systems can be used to interact with and / or perceive a CGR environment, including head-up displays (HUDs), head-mountable systems, projection-based systems, headphones / earphones, speaker arrays, smartphones, tablets, and desktop / laptop computers. For example, a head-mountable system may include one or more speakers (e.g., a speaker array), an integrated or external opaque, translucent, or transparent display, an image sensor that captures video of the physical environment, and / or a microphone that captures audio of the physical environment. The display may be implemented using a variety of display technologies, including uLED, OLED, LED, liquid crystal on silicon, laser scanning light source, digital light projection, etc., and may implement optical waveguides, optical reflectors, holographic media, optical combiners, combinations thereof, or similar technologies as the medium for directing light to the user's eye. In implementations using transparent or translucent displays, the transparent or translucent displays may also be controlled to be opaque. The display can implement a projection-based system that projects images onto the user's retina and / or projects virtual CGR elements into the physical environment (e.g., as holograms or as projections mapped onto physical surfaces or objects).

[0005] An electronic device can be used to implement the use of a digital assistant in a CGR environment. Implementing a digital assistant in a CGR environment can assist a user of the electronic device in interacting with the CGR environment and can allow the user to access digital assistant functionality without having to stop interacting with the CGR environment. However, because the interface of a CGR environment can be large and complex (e.g., the CGR environment can fill and extend beyond the user's field of view), invoking and interacting with a digital assistant within the CGR environment can be difficult, confusing, or distracting from the immersion of the CGR environment. Summary of the Invention

[0006] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors, memory, a display, and one or more sensors, detecting a first user input using one or more sensors while displaying a portion of a computer-generated reality (CGR) environment representing a current field of view of a user of the electronic device, and initiating a first digital assistant session in accordance with a determination that the first user input satisfies at least one criterion for initiating a digital assistant session, wherein initiating the first digital assistant session includes placing a digital assistant object at a first location within the CGR environment at a first time and outside the displayed portion of the CGR environment, and providing a first output indicating the first location of the digital assistant within the CGR environment.

[0007] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. The one or more programs, when executed by one or more processors of the electronic device, cause the electronic device to detect a first user input using one or more sensors, and, in accordance with a determination that the first user input meets at least one criterion for starting a digital assistant session, start a first digital assistant session, wherein starting the first digital assistant session includes placing a digital assistant object in a first location within the CGR environment at a first time and outside a displayed portion of the CGR environment, and providing a first output indicating the first location of the digital assistant within the CGR environment.

[0008] An exemplary electronic device is disclosed herein. The exemplary electronic device includes one or more processors, a memory, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs including: detecting a first user input using one or more sensors; and, in accordance with a determination that the first user input satisfies at least one criterion for initiating a digital assistant session, initiating a first digital assistant session, wherein the initiation of the first digital assistant session includes placing a digital assistant object in a first location within the CGR environment at a first time and outside a displayed portion of the CGR environment; and providing a first output indicating the first location of the digital assistant within the CGR environment.

[0009] An exemplary electronic device includes means for detecting a first user input using one or more sensors; means for initiating a first digital assistant session in accordance with a determination that the first user input satisfies at least one criterion for initiating a digital assistant session, wherein initiating the first digital assistant session includes placing a digital assistant object at a first location within the CGR environment at a first time and outside a displayed portion of the CGR environment; and means for providing a first output indicating the first location of the digital assistant within the CGR environment.

[0010] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors, memory, a display, and one or more sensors, detecting a user input using one or more sensors, and initiating a first digital assistant session according to a determination that the first user input satisfies criteria for initiating a digital assistant session, wherein starting the first digital assistant session includes displaying a first portion of a computer-generated reality (CGR) environment on the display. The method includes placing a digital assistant object at a first location within the CGR environment and outside the first portion of the CGR environment, and providing a first output indicating the first location of the digital assistant object within the CGR environment.

[0011] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. When executed by one or more processors of the electronic device, the one or more programs cause the electronic device to detect a user input using one or more sensors, and, in accordance with a determination that the first user input meets the criteria for starting a digital assistant session, start a first digital assistant session, wherein starting the first digital assistant session includes displaying a first portion of the computer-generated reality (CGR) environment on the display. While displaying a first portion of the CGR environment, placing a digital assistant object at a first location outside the first portion of the CGR environment, and providing a first output indicating the first location of the digital assistant object within the CGR environment.

[0012] An exemplary electronic device is disclosed herein. The exemplary electronic device includes one or more processors, a memory, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs detecting a user input using one or more sensors, and initiating a first digital assistant session according to a determination that the first user input satisfies criteria for initiating a digital assistant session, wherein the initiation of the first digital assistant session includes displaying a first portion of the computer-generated reality (CGR) environment on the display. The initiation includes placing a digital assistant object at a first location within the CGR environment and outside the first portion of the CGR environment, and providing a first output indicating the first location of the digital assistant object within the CGR environment.

[0013] An exemplary electronic device includes means for detecting user input using one or more sensors; and means for initiating a first digital assistant session in accordance with a determination that the first user input meets criteria for initiating a digital assistant session, wherein initiating the first digital assistant session includes placing a digital assistant object at a first location within a computer-generated reality (CGR) environment but outside the first portion of the CGR environment while displaying the first portion of the CGR environment on a display; and means for providing a first output indicating the first location of the digital assistant object within the CGR environment.

[0014] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors, memory, a display, and one or more sensors, detecting a first user input using one or more sensors while displaying a portion of a computer-generated reality (CGR) environment on the display, and initiating a first digital assistant session according to a determination that the first user input meets criteria for initiating a digital assistant session, wherein starting the first digital assistant session includes starting a digital assistant object at a first location within the CGR environment at a first time and outside the portion of the CGR environment, and animating the digital assistant object to move to a second location within the portion of the CGR environment at the first time.

[0015] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. When executed by one or more processors of an electronic device, the one or more programs cause the electronic device to detect a first user input using one or more sensors while displaying a portion of a computer-generated reality (CGR) environment on a display, and to start a first digital assistant session according to a determination that the first user input meets the criteria for starting a digital assistant session, wherein starting the first digital assistant session includes starting a digital assistant object at a first location within the CGR environment at a first time and outside the portion of the CGR environment, and animating the digital assistant object to move to a second location within the portion of the CGR environment at the first time.

[0016] An exemplary electronic device is disclosed herein. The exemplary electronic device includes one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for detecting a first user input using one or more sensors while displaying a portion of a computer-generated reality (CGR) environment on a display, and instructions for starting a first digital assistant session according to a determination that the first user input meets the criteria for starting a digital assistant session, wherein starting the first digital assistant session includes starting a digital assistant object at a first location within the CGR environment at a first time and outside the portion of the CGR environment, and animating the digital assistant object to move to a second location within the portion of the CGR environment at the first time.

[0017] An exemplary electronic device includes means for detecting a first user input using one or more sensors while displaying a portion of a computer-generated reality (CGR) environment on a display; and means for initiating a first digital assistant session according to a determination that the first user input meets criteria for initiating a digital assistant session, wherein initiating the first digital assistant session includes initiating a digital assistant object at a first location within the CGR environment at the first time and outside the portion of the CGR environment, and animating the digital assistant object to move to a second location within the portion of the CGR environment at the first time.

[0018] Placing a representation of a digital assistant in a computer-generated reality (CGR) environment as described herein provides an intuitive and efficient user interface for interacting with the digital assistant in the CGR environment. For example, initially placing a digital assistant object outside the user's field of view and providing an indication of the location of the digital assistant object effectively attracts the user's attention to the digital assistant, reducing the time and user input required for the user to access a desired function, and thus reducing power usage and improving the device's battery life. As another example, initializing a digital assistant object outside the user's field of view and animating the digital assistant as it moves into the user's field of view also effectively attracts the user's attention to the digital assistant, reducing the time and user input required for the user to access a desired function, and thus reducing power usage and improving the device's battery life. [Brief explanation of the drawings]

[0019] [Figure 1A] 1 illustrates an exemplary system for use with various augmented reality technologies, according to various embodiments. [Figure 1B] 1 illustrates an exemplary system for use with various augmented reality technologies, according to various embodiments.

[0020] [Figure 2A] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 2B] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 2C] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 2D] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 2E]1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments.

[0021] [Figure 3A] 1 illustrates a flow diagram of a method for placing a representation of a digital assistant in a CGR environment, according to various embodiments. [Figure 3B] 1 illustrates a flow diagram of a method for placing a representation of a digital assistant in a CGR environment, according to various embodiments.

[0022] [Figure 4A] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 4B] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 4C] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 4D] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 4E] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments. [Figure 4F] 1 illustrates a process for placing a representation of a digital assistant within a CGR environment, according to various embodiments.

[0023] [Figure 5A] 1 illustrates a flow diagram of a method for placing a representation of a digital assistant in a CGR environment, according to various embodiments. [Figure 5B] 1 illustrates a flow diagram of a method for placing a representation of a digital assistant in a CGR environment, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0024] In the following description of the embodiments, reference is made to the accompanying drawings, which show, by way of illustration, specific embodiments that may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the various embodiments.

[0025] The digital assistant can be used in a CGR environment. In some embodiments, upon invocation, a digital assistant object representing the digital assistant can be placed in a first location within the CGR environment but outside the user's current field of view, and an indication of the location of the digital assistant object can be provided. In some embodiments, upon invocation, a digital assistant object can be placed in a first location within the CGR environment but outside the user's current field of view, and then animated to move from the first location to a second visible location.

[0026] In the following description, terms such as "first" and "second" are used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first input can be referred to as a second input, and similarly, a second input can be referred to as a first input, without departing from the scope of various embodiments described. The first input and the second input are both inputs, and in some cases, are separate and distinct inputs.

[0027] The terminology used in the description of the various embodiments set forth herein is for the purpose of describing particular embodiments only and is not intended to be limiting. When used in the description of the various embodiments set forth and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" should be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0028] The term "if" can be interpreted to mean "when" or "upon," or "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "upon determining," or "in response to determining," or "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]," depending on the context. 1. The process of placing a representation of a digital assistant within a CGR environment

[0029] 1A and 1B illustrate an exemplary system 800 for use in various computer-generated reality technologies.

[0030] 1A , system 800 includes device 800a. Device 800a includes various components, such as processor(s) 802, RF circuit(s) 804, memory(s) 806, image sensor(s) 808, orientation sensor(s) 810, microphone(s) 812, location sensor(s) 816, speaker(s) 818, display(s) 820, and touch-sensitive surface 822. These components optionally communicate via communication bus(es) 850 of device 800a.

[0031] In some embodiments, elements of system 800 are implemented within a base station device (e.g., a computing device such as a remote server, a mobile device, or a laptop), and other elements of system 800 are implemented within a head-mounted display (HMD) device designed to be worn by a user, and the HMD device communicates with the base station device. In some embodiments, device 800a is implemented within the base station device or the HMD device.

[0032] 1B , in some embodiments, system 800 includes two (or more) devices in communication, such as via a wired or wireless connection. A first device 800b (e.g., a base station device) includes processor(s) 802, RF circuit(s) 804, and memory(s) 806. These components optionally communicate via communication bus(es) 850 of device 800b. A second device 800c (e.g., a head-mounted device) includes various components, such as processor 802, RF circuit(s) 804, memory 806, image sensor 808, orientation sensor 810, microphone 812, location sensor 816, speaker 818, display 820, and touch-sensitive surface 822. These components optionally communicate via communication bus(es) 850 of device 800c.

[0033] System 800 includes processor(s) 802 and memory(s) 806. Processor(s) 802 include one or more general-purpose processors, one or more graphics processors, and / or one or more digital signal processors. In some embodiments, memory(s) 806 are one or more non-transitory computer-readable storage media (e.g., flash memory, random access memory) that store computer-readable instructions configured to be executed by processor(s) 802 to perform the techniques described below.

[0034] System 800 includes RF circuit(s) 804. RF circuit(s) 804 optionally include circuitry for communicating with electronic devices, networks such as the Internet, an intranet, and / or wireless networks such as cellular networks and wireless local area networks (LANs). RF circuit(s) 804 optionally include circuitry for communicating using near-field and / or short-range communications such as Bluetooth®.

[0035] System 800 includes display(s) 820. In some embodiments, display(s) 820 include a first display (e.g., a left-eye display panel) and a second display (e.g., a right-eye display panel), each display displaying an image to a respective eye of the user. Corresponding images are simultaneously displayed on the first and second displays. Optionally, the corresponding images include representations of the same virtual object and / or the same physical object viewed from different viewpoints, resulting in a parallax effect on the displays that provides the user with the illusion of depth of the objects. In some embodiments, display(s) 820 include a single display. Corresponding images are simultaneously displayed on first and second regions of the single display, for each eye of the user. Optionally, the corresponding images include representations of the same virtual object and / or the same physical object viewed from different viewpoints, resulting in a parallax effect on the single display that provides the user with the illusion of depth of the objects.

[0036] In some embodiments, system 800 includes touch-sensitive surface(s) 822 that receive user input, such as tap input or swipe input. In some embodiments, display(s) 820 and touch-sensitive surface(s) 822 form touch-sensitive display(s).

[0037] System 800 includes image sensor(s) 808. Image sensor(s) 808 optionally include one or more visible light image sensors, such as charge-coupled device (CCD) sensors, and / or complementary metal-oxide semiconductor (CMOS) sensors, operable to obtain images of physical objects from the real environment. Image sensor(s) also optionally include one or more infrared (IR) sensor(s), such as passive IR sensors or active IR sensors, for detecting infrared light from the real environment. For example, an active IR sensor includes an IR emitter, such as an IR dot emitter, for emitting infrared light into the real environment. Image sensor(s) 808 also optionally include one or more event cameras configured to capture movement of physical objects in the real environment. Image sensor(s) 808 also optionally include one or more depth sensors configured to detect the distance of physical objects from system 800. In some embodiments, system 800 uses a combination of a CCD sensor, an event camera, and a depth sensor to detect the physical environment around system 800. In some embodiments, image sensor(s) 808 include a first image sensor and a second image sensor. The first image sensor and the second image sensor are configured to capture images of physical objects in the real environment, optionally from two different perspectives. In some embodiments, system 800 uses image sensor(s) 808 to receive user input, such as hand gestures. In some embodiments, system 800 uses image sensor(s) 808 to detect the position and orientation of system 800 and / or display(s) 820 in the real environment. For example, system 800 uses image sensor(s) 808 to track the position and orientation of display(s) 820 relative to one or more fixed objects in the real environment.

[0038] In some examples, system 800 includes microphone(s) 812. System 800 uses microphone(s) 812 to detect sounds from the user and / or the user's physical environment. In some embodiments, microphone(s) 812 optionally include an array of microphones (including multiple microphones) operating in concert, such as to identify ambient noise or to identify the location of a sound source within the space of a real environment.

[0039] System 800 includes orientation sensor(s) 810 that detect orientation and / or movement of system 800 and / or display(s) 820. For example, system 800 uses orientation sensor(s) 810 to track changes in position and / or orientation of system 800 and / or display(s) 820 relative to physical objects in a real-world environment, etc. Orientation sensor(s) 810 optionally include one or more gyroscopes and / or one or more accelerometers.

[0040] 2A-2E illustrate a process (e.g., method 1000) for deploying a representation of a digital assistant in a CGR environment according to various embodiments. The process is performed, for example, using one or more electronic devices implementing the digital assistant. In some examples, the process is performed using a client-server system, and the process steps are divided in any manner between a server and a client device. In other examples, the process steps are divided between a server and multiple client devices (e.g., a head-wearable system (e.g., a headset) and a smartwatch). Thus, while portions of the process are described herein as being performed by a particular device in a client-server system, it will be understood that the process is not so limited. In other examples, the process is performed using only a client device (e.g., device 906) or multiple client devices. In this process, some steps are optionally combined, the order of some steps is optionally changed, and some steps are optionally omitted. In some examples, additional steps may be performed in combination with the illustrated process.

[0041] 1A-1B, e.g., device 800a or 800c. In some embodiments, device 906 communicates (e.g., using 5G, WiFi, a wired connection, etc.), directly or indirectly (e.g., via a hub device, a server, etc.) with one or more other electronic devices, such as a computer, a mobile device, a smart home device, etc. For example, as shown in FIGS. 2A-2E, device 906 may be directly or indirectly connected to a smart watch device 910, a smart speaker device 912, a television 914, and / or a stereo system 916. In some embodiments, as described with respect to FIGS. 1A-1B, device 906 has (and / or communicates directly or indirectly with other devices that have) one or more sensors, such as an image sensor (e.g., for capturing visual content of the physical environment, gaze detection, etc.), an orientation sensor, a microphone, a location sensor, a touch-sensitive surface, an accelerometer, etc.

[0042] 2A-2E, a user 902 is shown immersed in a computer-generated reality (CGR) environment 904 using a device 906 at various steps of a process 900. The right panels of FIGS. 2A-2E each show a corresponding currently displayed portion 908 of the CGR environment 904 (e.g., the current field of view of the user 902) at each step of the process 900, as displayed on one or more displays of the device 906.

[0043] In some embodiments, the CGR environment 904 may include mixed reality (MR) content, augmented reality (AR) content, virtual reality (VR) content, etc. For example, as shown in Figures 2A-2E, the CGR environment includes MR content and allows the user to view both physical objects and the environment (e.g., physical devices 910, 912, 914, or 916, furniture or walls in a room in which the user 902 is located, etc.) and virtual objects 918A-C (e.g., virtual furniture items including a chair, a picture frame, and a vase).

[0044] 2A, device 906 detects user input (e.g., using one or more sensors). In currently displayed portion 908A (e.g., the current field of view of user 902 while providing user input), physical devices television 914 and stereo system 916 are visible along with virtual object 918B.

[0045] In some embodiments, the user input may include an auditory input, such as a voice input including a trigger phrase, an eye gaze input, such as a user directing their gaze at a particular location for at least a threshold period, a gesture input, a button press, a tap, a controller, a touchscreen, or a device input, etc. For example, as shown in FIG. 2A, the device 906 detects that the user 902 says "Hey, Assistant" and / or that the user 902 raises their wrist in a "raise and speak" gesture. While both the verbal input "Hey, Assistant" and the "raise and speak" gesture are provided, just one of these inputs may be sufficient to trigger a digital assistant session (e.g., as described with respect to FIG. 2B).

[0046] 2B, a digital assistant session is initiated according to a determination that the user input satisfies one or more criteria for initiating a digital assistant session. For example, as shown in FIG. 2B, the criteria for initiating a digital assistant session may include criteria for matching a predefined voice trigger (e.g., "Hey, Assistant") and / or criteria for matching a predefined gesture trigger (e.g., a "raise and speak" gesture). Thus, when user 902 raises their wrist with a "raise and speak" gesture as illustrated (as in FIG. 2A) and utters the trigger phrase "Hey, Assistant," a digital assistant session is initiated.

[0047] Starting a digital assistant session includes placing a digital assistant object 920 at a first location (e.g., a first position) within the CGR environment 904. As shown in FIG. 2B, the digital assistant object 920 is a virtual (e.g., VR) sphere located at a first location near the physical smart speaker device 912. However, because the first location of the digital assistant object 920 is to the right of the user 902 and the user 902 is looking ahead, the first position of the digital assistant object 920 is outside (e.g., not visible within) the currently displayed portion 908B at the time the digital assistant session is initiated.

[0048] In some embodiments, the first location of the digital assistant object 920 may be determined based on one or more environmental factors, such as characteristics of the physical environment, characteristics of the CGR environment 904, the location, position, or posture of the user 902, or the locations, positions, or postures of other users that may be present.

[0049] For example, the first location of the digital assistant object 920 may be selected to be near the location of the physical smart speaker device 912 and to avoid collisions (e.g., visual intersections) with physical objects (such as a table on which the physical smart speaker device 912 is placed) and / or virtual objects 918A-C (such as virtual object 918C). The location of the smart speaker device 912 may be determined based on a pre-specified location (e.g., a user of the smart speaker device 912 manually identifies and tags the device location), visual analysis of image sensor data (e.g., analyzing image sensor data such as data from image sensor(s) 808 to recognize the smart speaker device 912), analysis of other sensor data, etc. (e.g., using a Bluetooth connection when in general proximity).

[0050] In some embodiments, device 906 may provide an output indicating the state of the digital assistant session. The state output may be selected from two or more different outputs representing a selected state from two or more different states. For example, the two or more states may include a listening state, which may be further subdivided into active and passive listening states, a responsive state, a thinking (e.g., processing) state, an attention acquisition state, etc., which may be indicated by visual, auditory, and / or tactile output.

[0051] For example, as shown in FIG. 2B, the appearance of the digital assistant object 920 (e.g., a virtual sphere) includes a cloud shape in the center of the sphere, indicating that the digital assistant object 920 is in an attention acquisition state. Other state outputs include a change in the size of the digital assistant object 920, animation of movement (e.g., animation of the digital assistant object 920 hopping, hovering, etc.), auditory output (e.g., directional audio output "emanating" from the current location of the digital assistant object 920), changes in lighting effects (e.g., changes in the lighting or light emitted by the digital assistant object 920, or changes in pixels or light sources "pinned" to the edge of the display in the direction of the digital assistant object 920), haptic output, etc.

[0052] Initiating the digital assistant session includes generating an output indicating the first location of the digital assistant object 920 within the CGR environment 904. In some embodiments, the output may include an auditory output, such as an auditory output indicating the location using spatial sound, a tactile output, or a visual indication. For example, as shown in FIG. 2B, the output includes the verbal auditory output "Huh?" output from the smart speaker device 912 and the stereo system 916. The verbal auditory output "Huh?" may be provided using spatial sound technology, generating components of the audio signal coming from the smart speaker device 912 at a louder volume than components of the audio signal coming from the stereo system 916, so that the output sounds as if it is coming from the first location of the digital assistant object 920 (e.g., from the right hand side of the user as shown in FIG. 2B).

[0053] As another example, as shown in FIG. 2B, the output further includes a skin tap haptic output from smartwatch device 910 on the right wrist of user 902 indicating that the first location is to the right of the user.

[0054] As a further example, the output may include a visual indication of the first location (i.e., a visual indication other than the display of a digital assistant object currently outside the field of view). Visual output may be provided using the display of device 906, such as a change in the lighting of the CGR environment to indicate light emanating from the first location (e.g., rendering light and shadow in the 3D space visible in the currently displayed portion 908B, "pinning" illuminated pixels or light sources to the edge of the currently displayed portion 908B in the direction of the first location, and / or displaying a head-up display or 2D lighting overlay). Visual output may also be provided using non-display hardware, such as edge lighting (e.g., LEDs, etc.) illuminated in the direction of the first location. The lighting or light may vary in intensity to further draw attention to the first location.

[0055] Referring now to FIG. 2C, in some embodiments, after starting a digital assistant session, in accordance with a determination that the first location is within the currently displayed portion 908C (e.g., when user 902 is positioned so that the first location is within the current field of view of user 902), digital assistant object 920 is displayed (e.g., on the display of device 906).

[0056] As the digital assistant session progresses, device 906 can provide updated output indicating the updated state of the digital assistant session. For example, as shown in FIG. 2C, when the digital assistant session enters an active listening state, digital assistant object 920 includes a spiral shape in the center of the sphere, indicating that digital assistant object 920 is in an active listening state.

[0057] In some embodiments, device 906 detects a second user input (e.g., using one or more sensors). Device 906 then determines the intent of the second user input, for example, using natural language processing methods. For example, as shown in FIG. 2C , device 906 detects that user 902 is uttering the request "Play some music on my stereo," which corresponds to an intent to play audio.

[0058] In some embodiments, device 906 determines the intention according to a determination that the first location is within currently displayed portion 908C (eg, the current field of view of user 902 at the time user 902 provides the second user input). That is, in some embodiments, the digital assistant responds solely according to a determination that user 902 intends to pay attention to digital assistant object 920 (eg, look at digital assistant object 920) and respond to digital assistant object 920.

[0059] 2D , in some embodiments, device 906 determines the intent of the second user input and then provides a response output based on the determined intent. For example, as shown in FIG. 2D , in response to the user input "play some music on the stereo," device 906 causes music to be played from stereo system 906 based on the determined intent to play audio.

[0060] In some embodiments, providing a responsive output in accordance with a determination that the determined intent relates to an object (e.g., either a physical object or a virtual object) located at an object location within the CGR environment 904 includes positioning a digital assistant object 920 near the object location. For example, as shown in FIG. 2D, in addition to playing music from the stereo system 916, the device 906 repositions the digital assistant object 920 near the location of the stereo system 906 to indicate that the digital assistant is performing a task. Additionally, as shown in FIG. 2D, the appearance of the digital assistant object 920 may be updated to indicate that the digital assistant object 920 is in a responsive state, such as by including a star shape in the center of a sphere as shown, or some other appropriate status output (e.g., as described above with respect to FIG. 2B) may be provided.

[0061] 2E, in some embodiments, the digital assistant session may end, for example, when the user 902 explicitly dismisses the digital assistant (e.g., using voice input, gaze input, gesture input, button press, tap, controller, touch screen, or device input, etc.), or automatically (e.g., after a predetermined threshold period without interaction). When the digital assistant session ends, the device 906 dismisses the digital assistant object 920. As shown in FIG. 2E, if the current location of the digital assistant object 920 is within the currently displayed portion 908E (e.g., the current field of view of the user 902 at the time the digital assistant session ends), dismissing the digital assistant object 920 includes ceasing to display the digital assistant object 920.

[0062] In some embodiments, dismissing the digital assistant may also include providing a further output indicating the dismissal. The dismissal output may include an indication such as an audible output (e.g., a chime, a verbal output, etc.) or a visual output (e.g., a displayed object, a change in the lighting of the CGR environment 904, etc.). For example, as shown in FIG. 2E, the dismissal output includes a digital assistant indicator 922 positioned at a first location (e.g., the initial location of the digital assistant object 920), thereby indicating where the digital assistant object 920 will reappear if another digital assistant session is started. Because the first location is within the currently displayed portion 908E (e.g., the current field of view of the user 902 at the time the digital assistant session ends), the device 906 displays the digital assistant indicator 922.

[0063] The processes described above with reference to Figures 2A-2E are optionally performed by the components shown in Figures 1A-1B. For example, the operations of the illustrated processes may be performed by electronic devices (e.g., 800a, 800b, 800c, or 906). It will be apparent to one skilled in the art how other processes may be performed based on the components shown in Figures 1A-1B.

[0064] 3A-3B show a flow diagram of a method 1000 for placing a representation of a digital assistant in a computer-generated reality (CGR) environment according to some embodiments. Method 1000 can be performed using one or more electronic devices (e.g., devices 800a, 80b, 800c, 906) having one or more processors and memory. In some embodiments, method 1000 is performed using a client-server system, and the operations of method 1000 are divided between client devices (e.g., 800c, 906) and the server in any manner. Some operations of method 1000 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0065] Method 1000 is performed while displaying at least a portion of a CGR environment. That is, at a particular time, the particular portion of the CGR environment being displayed represents the current field of view of a user (e.g., a user of a client device(s)), and other portions of the CGR environment (e.g., behind the user or outside the user's peripheral vision) are not displayed. Thus, while method 1000, for example, refers to placing virtual objects and generating "visual" output, the actual visibility of the virtual objects and output to the user may vary depending on the particular portion of the CGR environment currently being displayed. Terms such as "first time," "second time," "first portion," and "second portion" are used to distinguish displayed virtual content from non-displayed virtual content and are not intended to indicate a fixed order or predetermined portions of the CGR environment.

[0066] In some embodiments, the CGR environment of method 1000 may include virtual and / or physical content (e.g., physical devices 910, 912, 914, or 916, furniture or walls in the room in which user 902 is located, and virtual objects 918A-C shown in FIGS. 2A-2E). The content of the CGR environment may be static or dynamic. For example, static physical objects in the CGR environment may include physical furniture, walls, ceilings, etc., while static virtual objects in the CGR environment may include virtual objects that are located at a fixed location in the CGR environment (e.g., virtual furniture such as virtual objects 918A-C) or virtual objects that are located at a fixed location relative to a display (e.g., a heads-up display overlaid on a displayed portion of the CGR environment). Dynamic physical objects in the CGR environment may include the user, other users, pets, etc., and dynamic virtual items in the CGR environment include objects that move (e.g., virtual characters, avatars, or pets) or objects that change size, shape, or form.

[0067] 3A, at block 1002, a first user input is detected by one or more sensors of the device(s) performing method 1000. For example, the one or more sensors may include an audio sensor (e.g., a microphone), a vibration sensor, a motion sensor (e.g., an accelerometer, a camera, etc.), a visual sensor (e.g., a light sensor, a camera, etc.), a touch sensor, etc.

[0068] In some embodiments, the first user input includes an auditory input. For example, the first user input may include a voice input including a trigger phrase (e.g., "Hey, Assistant"). In some embodiments, the first user input includes a gaze input. For example, a user may direct their gaze at a particular location (e.g., a default digital assistant location, a location of a smart speaker or device, a location of an object that the digital assistant can assist with, etc.) for at least a threshold period of time. In some embodiments, the first user input includes a gesture (e.g., a user's body movement) input. For example, a user may raise their wrist in a "raise and speak" gesture. In some embodiments, the first user inputs a button press, a tap, a controller, a touchscreen, or a device input. For example, a user may press and hold the touchscreen of the smartwatch device 910.

[0069] In block 1004, a digital assistant session is initiated in accordance with a determination that the first user input meets the criteria for initiating a digital assistant session.

[0070] For example, if the first user input includes an auditory input, the criteria may include matching a predetermined voice trigger (e.g., "Hey, assistant") with sufficient confidence. As another example, if the first user input includes a gaze input, the criteria may include the user directing their gaze at a particular location (e.g., a predetermined digital assistant location, the location of an object with which the digital assistant can interact, etc.) for at least a threshold period. As another example, if the first user input includes a gesture (e.g., a user's body movement) input, the criteria may include matching a predetermined gesture trigger (e.g., a "raise wrist to speak" movement, etc.) with sufficient confidence. To determine whether the user has invoked a digital assistant session, one or more possible trigger inputs may be considered together or individually.

[0071] Starting the digital assistant session in block 1004 includes placing the digital assistant object in a first (eg, initial) location in the CGR environment at a first time and outside a first (eg, currently displayed) portion of the CGR environment. That is, the electronic device performing block 1004 places the digital assistant object in the CGR environment so that the digital assistant object is not visible to the user when the digital assistant session is initiated (eg, not displayed or "off-screen").

[0072] In some embodiments, the first (e.g., initial) location in the CGR environment is a predetermined location in the CGR environment. For example, the predetermined location may be a predetermined set of coordinates in the coordinate system of the CGR environment. The predetermined location may be determined (e.g., selected) by the user in a previous digital assistant session (e.g., as described below with respect to the second digital assistant session).

[0073] In some embodiments, in block 1006, a first (e.g., initial) location within the CGR environment is determined based on one or more environmental factors, such as the physical environment in which the user is working, the user's position, or the positions of multiple users.

[0074] The one or more environmental factors may include characteristics of the CGR environment. For example, the first location may be determined based on the physical location of the electronic device (e.g., smart speaker device 912 in FIGS. 2A-2E), the static or dynamic location of another physical object (e.g., furniture or a pet running into the room), the static or dynamic location of a virtual object (e.g., virtual objects 918A-C in FIGS. 2A-2E), etc. The location of the electronic device may be determined based on a pre-identified location (e.g., a user manually identifying and tagging the device location), based on visual analysis of image sensor data (e.g., analyzing image sensor data to recognize or visually understand the device), based on analysis of other sensor or connection data (e.g., using a Bluetooth connection when in general proximity), etc.

[0075] The one or more environmental factors may also include the position (e.g., location and / or posture) of the user. For example, the first location may be determined to be a location behind the user based on the orientation of the user's body or head, or the position of the user's gaze. As another example, the first location may be determined to be a location on or near the user's body, such as placing a sphere on the user's wrist.

[0076] The one or more environmental factors may also include multiple positions (e.g., location and / or posture) of multiple users of the CGR environment. For example, in a shared CGR environment such as a virtual conference room or multiplayer game, the first location may be determined to be a location that minimizes (or maximizes) the visibility of the digital assistant object to the majority of shared users based on where each user is facing and / or looking.

[0077] A digital assistant object is a virtual object that represents a digital assistant (e.g., an avatar in a digital assistant session). For example, a digital assistant object may be a virtual sphere, a virtual character, a virtual light ball, an avatar, etc. The digital assistant object can change appearance and / or form throughout the digital assistant session, for example, from a sphere to a virtual light ball, from a translucent sphere to an opaque sphere, and / or similar.

[0078] In block 1008, a first output indicating a first location of the digital assistant object within the CGR environment is provided. That is, although the first (eg, initial) location is outside the first (eg, currently displayed) portion of the CGR environment, the first output indicating the first location helps the user locate the digital assistant object within the CGR environment (eg, find the digital assistant object), for example, by quickly and intuitively attracting the user's attention to the digital assistant session, thereby increasing the efficiency and effectiveness of the digital assistant session.

[0079] In some embodiments, providing a first output indicative of the first location includes generating a first auditory output. That is, the device performing method 1000 may generate the first auditory output itself (e.g., using a built-in speaker or a headset) and / or cause one or more other suitable auditory devices to generate the first auditory output. For example, the first auditory output may include a verbal output (e.g., "Huh?", "Yes?", "Can I help you?", etc.), another auditory output (e.g., a chime, a hum, etc.), and / or a mixed auditory / tactile output (e.g., a hum resulting from a vibration that can also be felt by the user).

[0080] In some embodiments, the first auditory output may be provided using spatial sound technology, such as using multiple speakers (e.g., a speaker array or a surround sound system) to emit multiple audio components (e.g., channels) at different volumes, making the overall auditory output appear to be emanating specifically from a particular location. For example, the first auditory output may include a first audio component (e.g., a channel) generated by a first speaker of the multiple speakers and a second audio component generated by a second speaker of the multiple speakers. A determination is made as to whether the first location of the digital assistant object is closer to the location of the first speaker or the location of the second speaker. In accordance with a determination that the first location is closer to the location of the first speaker, the first audio component is generated at a louder volume than the second audio component. Similarly, in accordance with a determination that the first location is closer to the location of the second speaker, the second audio component is generated at a louder volume than the first audio component.

[0081] In some embodiments, providing a first output indicating the first location includes generating a first haptic output. That is, the device(s) performing method 1000 may generate the first haptic output itself and / or cause one or more other suitable haptic devices to generate the first haptic output. The haptic output includes vibrations, taps, and / or other tactile outputs felt by a user of the device(s) performing method 1000. For example, as shown in FIG. 2B, device 910 may generate a vibration haptic output on the right wrist of user 902, indicating that digital assistant object 920 is on the right side of user 902.

[0082] In some embodiments, providing a first output indicating the first location includes displaying a visual indication of the first location. For example, the visual indication may include emitting a virtual light from the first location, changing pass-through filtering of the physical environment lighting, and / or changing the lighting of the physical environment using an appropriate home automation device to directionally illuminate the CGR environment. As another example, the visual indication may include displaying an indicator other than the digital assistant object to guide the user to the first location.

[0083] In some embodiments, in block 1010, in accordance with a determination that the first location is within a second portion of the CGR environment (eg, the portion currently displayed at the second time), the digital assistant object is displayed at the first location (eg, on one or more displays of the device implementing method 1000). That is, after starting the first digital assistant session with the digital assistant object positioned off-screen (eg, outside the user's field of view at the start), when the user changes their viewpoint (eg, by looking in a different direction or providing another input) and looks at the location of or near the digital assistant object, the digital assistant object becomes visible to the user.

[0084] In some embodiments, in block 1012, a second user input is detected (e.g., at a third time after the start of the first digital assistant session). For example, the second user input may include a verbal input such as a verbal command, a question, a request, a shortcut, etc. As another example, the second user input may also include a gesture input such as a gesture command, a question, a request, a gesture representing an interaction with the CGR environment (e.g., "grabbing" and "dropping" a virtual object), etc. As another example, the second user input may include a gaze input.

[0085] In some embodiments, at block 1014, the intent of the second user input is determined. The intent may correspond to one or more tasks that may be performed using one or more parameters. For example, if the second user input includes a verbal or gestural command, question, or request, the intent may be determined using natural language processing techniques, such as determining the intent to play audio from the verbal user input "Play some music on the stereo" shown in FIG. 2C. As another example, if the second user input includes a gesture, the intent may be determined based on the gesture type(s), the gesture location(s), and / or the locations of various objects within the CGR environment. For example, a grab-and-drop type gesture may correspond to an intent to move a virtual object located at or near the location of the "grab" gesture to the location of the "drop" gesture.

[0086] In some embodiments, the determination of the intent of the second user input is made solely according to a determination that the current location of the digital assistant object is within the currently displayed portion of the CGR environment. That is, some detected user inputs are not intended for the digital assistant session, such as when the user talks to another person in a physical room or another player in a multiplayer game. Processing and responding only to user inputs received while the user is looking at (or near) the digital assistant object reduces the likelihood of unintended interactions, for example, and improves the efficiency of the digital assistant session.

[0087] In some embodiments, at block 1016, a second output is provided based on the determined intent. Providing the second output may include causing one or more tasks corresponding to the determined intent to be performed. For example, as shown in FIGS. 2A-2E, in response to a verbal user input "Play some music on the stereo," the second output may include causing the stereo system 916 to play some music. As another example, in response to a gestural input "What's the weather like today?" the second output may include displaying a widget showing a thunderstorm icon and a temperature.

[0088] In some embodiments, providing the second output based on the determined intent includes determining whether the determined intent is related to relocating the digital assistant object. For example, the second user input may include an explicit request to relocate the digital assistant, such as a verbal input such as "Please pass by the TV" or a gesture input such as grabbing and dropping. As another example, the second user input may include a gaze input originating from the current location of the digital assistant object combined with a "pinch" or "grab" to initiate movement of the digital assistant object.

[0089] According to a determination that the determined intention is related to rearranging the digital assistant object, a second location is determined from the second user input. For example, the second location "by the TV" can be determined from the verbal input "pass by the TV", a second location located at or near the "drop" gesture can be determined from the grab and drop gesture input, or a second location located at or near the user's gaze input can be determined.

[0090] Further, in accordance with a determination that the determined intention is related to relocating the digital assistant object, the digital assistant object is placed at a second location. In accordance with a determination that the second location is within a currently displayed (e.g., second) portion of the CGR environment (e.g., at a fourth time), the digital assistant object is displayed at a second location within the CGR environment. That is, while relocating the digital assistant object from its location when the second user input was detected to the second location determined from the second user input, the digital assistant object is displayed as long as its location remains within the user's current field of view (e.g., including animating the movement of the digital assistant object from one location to another).

[0091] In some embodiments, providing a second output based on the determined intent includes determining whether the determined intent is related to an object (e.g., a physical or virtual object) located at an object location within the CGR environment. In accordance with a determination that the determined intent is related to an object located at the object location within the CGR environment, the digital assistant is placed at a third location near the object location. That is, the digital assistant object moves near the associated object to indicate an interaction with the object and / or draw attention to the object or interaction. For example, as shown in FIGS. 2A-2E, in response to the verbal user input "Play some music on the stereo," the digital assistant object 920 moves near the stereo system 916, as if the digital assistant object itself had turned on the stereo system 916 to indicate to the user that the requested task involving the stereo is being performed.

[0092] In some embodiments, in block 1018, a third output is provided based on one or more characteristics of the second output. The one or more characteristics of the second output may include the type of second output (e.g., visual, auditory, or tactile), the location of the second output within the CGR environment (e.g., for a task performed within the CGR environment), etc. For example, as shown in FIGS. 2A-2E, in response to a verbal user input "Play some music on the stereo," the third output may include the verbal output "OK, music is playing." As another example, in response to a gestural input "What's the weather like today?", the third output may include animating a digital assistant object bouncing or wiggling near a displayed weather widget (e.g., the second output). Providing a third output as described improves the efficiency of a digital assistant session, for example, by drawing the user's attention to the performance and / or completion of a requested task when that performance and / or completion is not immediately apparent to the user.

[0093] In some embodiments, in block 1020, a fifth output selected from two or more different outputs is provided to indicate a state of the first digital assistant session selected from two or more different states. The two or more different states of the first digital assistant session may include one or more listening states (e.g., active or passive listening), one or more response states, one or more processing (e.g., thinking) states, one or more failure states, one or more attention acquisition states, and / or one or more transition (e.g., moving, appearing, or disappearing) states. There may be a one-to-one correspondence between different outputs and different states, and one or more states may be represented by the same output, or one or more outputs may represent the same state (or variations of the same state).

[0094] For example, as shown in Figures 2A-2E, the digital assistant object 920 initially attracts the user's attention (e.g., shows an initial location), listens to user input, and responds to the user input / attracts the user's attention to the response. Other fifth (e.g., state) outputs may include changes in the size of the digital assistant object, animation of movement, auditory output (e.g., directional audio output "emanating" from the current location of the digital assistant object), changes in lighting effects, other visual output, tactile output, etc.

[0095] At any time during the first digital assistant session, if the currently displayed portion of the CGR environment is updated (e.g., in response to a user changing their viewpoint) and the current location of the digital assistant object is no longer included in the currently displayed portion (e.g., no longer visible to the user), an additional output indicating the current location of the digital assistant may be provided. For example, an additional output indicating the current position may be provided as described with respect to block 1008 (e.g., a spatial auditory output, a visual output, a tactile output, etc.).

[0096] In some embodiments, the first digital assistant session ends, for example, upon explicit exit by the user or after a threshold period has elapsed without interaction. When the first digital assistant session ends, the digital assistant object is dismissed (e.g., removed from the CGR environment) in block 1022. If the digital assistant object is located within a displayed portion of the CGR environment at the time the digital assistant session ends, dismissing the digital assistant object includes ceasing to display the digital assistant object.

[0097] In some embodiments, when the digital assistant object is dismissed, a fourth output indicating the dismissal of the digital assistant object is provided. The fourth output may include one or more visual outputs (e.g., a dim light, returning the lighting in the CGR environment to the state before the digital assistant session, etc.), one or more auditory outputs (e.g., a verbal output such as "bye," a chime, etc.), one or more tactile outputs, etc. For example, as shown in FIG. 2E, providing the fourth output may include placing a digital assistant indicator 922 at the first location. The fourth output may indicate to the user that the first digital assistant session has ended and may help the user more quickly locate the digital assistant object (e.g., find the digital assistant object) in subsequent calls.

[0098] In some embodiments, after the digital assistant object is dismissed, a third user input is detected. The third user input may be an auditory input, a gaze input, a gesture input, or a device input (e.g., a button press, a tap, a swipe, etc.) as described with respect to the first user input (e.g., the user input invoking the first digital assistant session). In accordance with a determination that the third user input satisfies at least one criterion for initiating a digital assistant session (e.g., as described above with respect to block 1004), a second digital assistant session is initiated.

[0099] In some embodiments in which a second user input related to relocating the digital assistant object (e.g., as described above with respect to block 1016) was received during the first digital assistant session, starting the second digital assistant session involves placing the digital assistant object in a second location (e.g., the location requested by the second user input). That is, in some embodiments, after the user explicitly moves the digital assistant object, the location selected by the user becomes the new "default" location where the digital assistant object will appear on subsequent calls.

[0100] The methods described above with reference to Figures 3A-3B are optionally performed by the components shown in Figures 1A-1B and 2A-2E. For example, the operations of the illustrated methods may be performed by electronic devices (e.g., 800a, 800b, 800c, or 906). It will be apparent to one skilled in the art how other processes may be implemented based on the components shown in Figures 1A-1B and 2A-2E.

[0101] 4A-4E illustrate a process (e.g., method 1200) for deploying a representation of a digital assistant in a CGR environment according to various embodiments. This process is performed, for example, using one or more electronic devices implementing the digital assistant. In some embodiments, the process is performed using a client-server system, and the process steps are divided in any manner between a server and a client device. In other embodiments, the process steps are divided between a server and multiple client devices (e.g., a head-wearable system (e.g., a headset) and a smartwatch). Thus, while portions of the process are described herein as being performed by a particular device in a client-server system, it will be understood that the process is not so limited. In other embodiments, the process is performed using only a client device (e.g., device 1106) or only multiple client devices. In this process, some steps are optionally combined, the order of some steps is optionally changed, and some steps are optionally omitted. In some embodiments, additional steps may be performed in combination with the illustrated process.

[0102] 4A-4E, a user 1102 is shown immersed in a computer-generated reality (CGR) environment 1104 using a device 1106 at various steps in the process. The right panels of FIGS. 4A-4E show a corresponding currently displayed portion 1108 of the CGR environment 1104 (e.g., the current field of view of the user 1102) at each step in the process, as displayed on one or more displays of the device 1106. The device 1106 may be implemented as described above with respect to the device 906 of FIGS. 1A-1B and 2A-2E.

[0103] In some embodiments, the CGR environment 1104 may include mixed reality (MR) content, augmented reality (AR) content, virtual reality (VR) content, etc. For example, as shown in Figures 4A-4E, the CGR environment 1104 includes MR content and allows the user to view both physical objects and the environment (e.g., physical devices 1110, 1112, 1114, or 1116, furniture or walls in a room in which the user 1102 is located, etc.) and virtual objects 1118A-C (e.g., virtual furniture items including a chair, a picture frame, and a vase).

[0104] 4A, device 1106 detects user input (e.g., using one or more sensors). In currently displayed portion 1108A (e.g., the current field of view of user 902 while providing user input), physical devices television 1114 and stereo system 1116 are visible, along with virtual object 1118B.

[0105] In some embodiments, the user input may include an auditory input, such as a voice input including a trigger phrase, an eye gaze input, such as a user directing their gaze at a particular location for at least a threshold period, a gesture input, a button press, a tap, a controller, a touchscreen, or a device input. For example, as shown in FIG. 4A, the device 1106 detects that the user 1102 says "Hey, assistant." To determine whether the user has invoked a digital assistant session, one or more possible trigger inputs may be considered, either together or individually.

[0106] 4B, a first digital assistant session is initiated according to a determination that the user input satisfies one or more criteria for initiating a digital assistant session. For example, as shown in FIG. 4B, the criteria for initiating a digital assistant session may include a criterion that matches a predefined voice trigger (e.g., "Hey, Assistant"). Thus, when user 1102 (as in FIG. 4A) utters the trigger phrase "Hey, Assistant" as illustrated, a digital assistant session is initiated.

[0107] Starting the first digital assistant session includes instantiating a digital assistant object 1120 at a first location (e.g., a first position) within the CGR environment 1104. As shown in FIG. 4B, the digital assistant object 1120 is a virtual (e.g., VR) sphere located at a first location behind the user 1102. Thus, the first position of the digital assistant object 1120 is outside (e.g., not visible within) the currently displayed portion 1108B at the time the digital assistant session is initiated.

[0108] In some embodiments, the first location of the digital assistant object 1120 may be determined based on one or more environmental factors, such as characteristics of the physical environment, characteristics of the CGR environment 1104, the location, position, or posture of the user 1102, and the locations, positions, or postures of other users that may be present. For example, as shown in FIG. 4B, the first location of the digital assistant object 1120 is selected to be a location that the user 1102 cannot currently see (e.g., a location outside the currently displayed portion 1108B) based on where the user 1102 is standing in the CGR environment and the direction the user 1102 is looking.

[0109] In some embodiments, device 1106 may provide an output indicating the state of the digital assistant session. The state output may be selected from two or more different outputs representing a selected state from two or more different states. For example, the two or more states may include a listening state, a responsive state, a thinking (e.g., processing) state, an attention acquisition state, etc., which may be further subdivided into active and passive listening states, and these may be indicated by visual, auditory, and / or tactile output. For example, as shown in FIG. 4B, the appearance of digital assistant object 1120 (e.g., a virtual sphere) includes a spiral shape in the center of the sphere, indicating that digital assistant object 1120 is in a listening state. As described with respect to FIGS. 2A-2E, updated state outputs representing different states may be provided as the conversation progresses.

[0110] 4B-4C, starting the first digital assistant session further includes animating the second moving digital assistant object 1120 within the currently displayed portion 1108C. That is, the digital assistant object 1120 moves within the CGR environment 1104 and becomes visible to the user 1102.

[0111] In some embodiments, animating the repositioning of the digital assistant object 1120 includes determining a movement path corresponding to a portion of the path between the first location and the second location within the currently displayed portion 1108C. That is, as shown in FIG. 4C, the digital assistant object 1120 is animated to "fly into view" from the initial (e.g., first) location and stop at the final (e.g., second) location visible to the user 1102.

[0112] In some embodiments, the travel path is determined not to pass through one or more objects (e.g., physical or virtual objects) located in the currently displayed portion(s) 1108B and 1108C of the CGR environment. For example, as shown in FIG. 4C, the travel path does not intersect with (physical) furniture in the room in which the user 1102 is located or with virtual objects (e.g., virtual object 1118C) in the CGR environment 1104, as if the digital assistant object 1120 were a "dodge" object in the CGR environment.

[0113] In some embodiments, the second location of the digital assistant object 1120 may be determined based on one or more environmental factors, such as characteristics of the physical environment, characteristics of the CGR environment 1104, the location, position, or posture of the user 1102, or the location, position, or posture of other users that may be present. For example, as shown in FIG. 4C, the second location of the digital assistant object 1120 is selected to be visible to the user 1102 (e.g., based on the position of the user 1102) and to avoid collisions (e.g., visual interference) with physical objects (such as the television 1114) and / or virtual objects 1118 (such as virtual object 1118C).

[0114] 4C, in some embodiments, the device 1106 detects a second user input (for example, using one or more sensors). The device 1106 then determines the intent of the second user input, for example, using natural language processing. A response output is provided based on the determined intent. For example, as shown in FIG. 4C, the device 1106 detects that the user 1102 is uttering the request "sit on the shelf," which corresponds to the intent to rearrange the digital assistant object 1120.

[0115] 4D, in an embodiment in which the determined intent is to relocate the digital intent object 1120, providing the response output includes determining a third location from the second user input. The digital assistant object 120 is then placed in the third location, and in accordance with a determination that the third location is within the currently displayed portion 1108D, the digital assistant object 1120 is displayed in the third location within the CGR environment 1104.

[0116] For example, based on the user input shown in FIG. 4C, in FIG. 4D, the determined third location is a location on a (physical) shelf in the CGR environment 1104, and therefore the digital assistant object 1120 moves from the second location (directly in front of the user 1102) to the third location on the shelf. Because the shelf is currently within the displayed portion 1108, the device 1106 displays the digital assistant object 1120 "sitting" on the shelf.

[0117] 4E, in some embodiments, the digital assistant session may end, for example, when user 1102 explicitly dismisses the digital assistant (e.g., using voice input, gaze input, gesture input, button press, tap, controller, touch screen, or device input, etc.), or automatically (e.g., after a predetermined threshold period without interaction). When the digital assistant session ends, device 1106 dismisses digital assistant object 1120. As shown in FIG. 4E, if the third location of digital assistant object 1120 is within currently displayed portion 1108E (e.g., the current field of view of user 1102 at the time the digital assistant session ends), dismissing digital assistant object 1120 includes ceasing to display digital assistant object 1120.

[0118] In some embodiments, dismissing the digital assistant may also include providing a further output indicating the dismissal. The dismissal output may include an indication such as an audible output (e.g., a chime, a verbal output, etc.) or a visual output (e.g., a displayed object, a change in the lighting of the CGR environment 1104, etc.). For example, as shown in FIG. 4E, the dismissal output includes a digital assistant indicator 1122 positioned at the first location (e.g., the initial location of the digital assistant object 1120), thereby indicating where the digital assistant object 920 will reappear if another digital assistant session is started. Because the third location is still within the currently displayed portion 1108E (e.g., the current field of view of the user 1102 at a time after the first digital assistant session has ended), the device 1106 displays the digital assistant indicator 1122.

[0119] In some embodiments, after dismissing the digital assistant object 1120, the device 1106 detects a third user input. For example, as shown in FIG. 4F, the device 1106 detects that the user 1102 is uttering the user input "Hey, assistant." As described above with respect to FIG. 4A, if the third input meets the criteria for starting a digital assistant session (e.g., matches a default voice trigger, gaze trigger, or gesture trigger), a second digital assistant session is started.

[0120] However, as shown in FIG. 4F, because user 1102 had previously rearranged digital assistant object 1120 on the shelf, starting the second digital assistant session involves placing the digital assistant object in a third location. That is, in some embodiments, after manually rearranging digital assistant object 1120 during a digital assistant session, the default (e.g., initial) location of digital assistant object 1120 changes according to user 1102's request.

[0121] The processes described above with reference to Figures 4A-4F are optionally performed by the components shown in Figures 1A-1B. For example, the operations of the illustrated processes may be performed by electronic devices (e.g., 800a, 800b, 800c, or 906). It will be apparent to one skilled in the art how other processes may be performed based on the components shown in Figures 1A-1B.

[0122] 5A-5B are flow diagrams illustrating a method 1200 for placing a representation of a digital assistant in a computer-generated reality (CGR) environment according to some embodiments. Method 1200 can be performed using one or more electronic devices (e.g., devices 800a, 800b, 800c, and / or 1106) having one or more processors and memory. In some embodiments, method 1200 is performed using a client-server system, and the operations of method 1200 are divided in any manner between a client device (e.g., 800c, 1106) and a server. Some operations of method 1200 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0123] Method 1200 is performed while displaying at least a portion of a CGR environment. That is, at a particular time, the particular portion of the CGR environment being displayed represents the current field of view of a user (e.g., a user of a client device(s)), and other portions of the CGR environment (e.g., behind the user or outside the user's peripheral vision) are not displayed. Thus, while method 1200, for example, refers to placing virtual objects and generating "visual" output, the actual visibility of the virtual objects and output to the user may vary depending on the particular portion of the CGR environment currently being displayed. Terms such as "first time," "second time," "first portion," and "second portion" are used to distinguish displayed virtual content from non-displayed virtual content and are not intended to indicate a fixed order or predetermined portions of the CGR environment.

[0124] In some embodiments, the CGR environment of method 1200 may include virtual and / or physical content (e.g., physical devices 1110, 1112, 1114, or 1116, furniture or walls in the room in which user 1102 is located, and virtual objects 118A-C shown in FIGS. 4A-4F). The content of the CGR environment may be static or dynamic. For example, static physical objects in the CGR environment may include physical furniture, walls, ceilings, etc., while static virtual objects in the CGR environment may include virtual objects that are located at a fixed location in the CGR environment (e.g., virtual furniture such as virtual objects 1118A-C) or virtual objects that are located at a fixed location relative to a display (e.g., a heads-up display overlaid on a displayed portion of the CGR environment). Dynamic physical objects in the CGR environment may include the user, other users, pets, etc., and dynamic virtual items in the CGR environment include objects that move (e.g., virtual characters, avatars, or pets) or objects that change size, shape, or form.

[0125] 5A, at block 1202, a first user input is detected by one or more sensors of the device(s) performing method 1200. For example, the one or more sensors may include an audio sensor (e.g., a microphone), a vibration sensor, a motion sensor (e.g., an accelerometer, a camera, etc.), a visual sensor (e.g., a light sensor, a camera, etc.), a touch sensor, etc.

[0126] In some embodiments, the first user input includes an auditory input. For example, the first user input may include a voice input including a trigger phrase (e.g., "Hey, Assistant"). In some embodiments, the first user input includes a gaze input. For example, a user may direct their gaze at a particular location (e.g., a default digital assistant location, the location of an object the digital assistant can assist with, etc.) for at least a threshold period of time. In some embodiments, the first user input includes a gesture (e.g., a user's body movement) input. For example, a user may raise their wrist in a "raise and speak" gesture. In some embodiments, the first user inputs a button press, a tap, a controller, a touchscreen, or a device input. For example, a user may press and hold the touchscreen of the smartwatch device 910.

[0127] In block 1204, a first digital assistant session is initiated in accordance with a determination that the first user input meets the criteria for initiating a digital assistant session.

[0128] For example, if the first user input includes an auditory input, the criteria may include matching a predefined voice trigger (e.g., "Hey, assistant") with sufficient confidence. As another example, if the first user input includes a gaze input, the criteria may include the user directing their gaze at a particular location (e.g., a predefined digital assistant location, the location of an object with which the digital assistant can interact, etc.) for at least a threshold period. As another example, if the first user input includes a gesture (e.g., a user's body movement) input, the criteria may include matching a predefined gesture trigger (e.g., a "raise wrist to speak" gesture, etc.) with sufficient confidence.

[0129] Starting the digital assistant session in block 1204 includes instantiating a digital assistant object in a first (e.g., initial) location in the CGR environment at a first time and outside a first (e.g., currently displayed) portion of the CGR environment in block 1206. That is, the electronic device performing block 1204 initially places the digital assistant object in the CGR environment so that the digital assistant object is not visible to the user when the digital assistant session is initiated (e.g., not displayed or "off-screen").

[0130] In some embodiments, the first (e.g., initial) location in the CGR environment is a predetermined location in the CGR environment. For example, the predetermined location may be a predetermined set of coordinates in the coordinate system of the CGR environment. The predetermined location may be determined (e.g., selected) by the user in a previous digital assistant session (e.g., as described below with respect to the second digital assistant session).

[0131] In some embodiments, block 1208 determines a first (e.g., initial) location within the CGR environment based on one or more first environmental factors, such as the physical environment in which the user is working, the user's position, or the positions of multiple users. The one or more first environmental factors may include characteristics of the CGR environment, for example, as described above with respect to block 1006 of FIG. 3A. The one or more first environmental factors may also include the user's position (e.g., location and / or posture). For example, the first location may be determined to be a location behind the user based on the user's body or head orientation or the user's gaze position.

[0132] A digital assistant object is a virtual object that represents a digital assistant (e.g., an avatar in a digital assistant session). For example, a digital assistant object may be a virtual sphere, a virtual character, a virtual light ball, etc. The digital assistant object can change appearance and / or form throughout the digital assistant session, for example, from a sphere to a virtual light ball, from a translucent sphere to an opaque sphere, and / or similar.

[0133] In block 1210, the digital assistant object is animated to move to a second location within a first (e.g., currently displayed) portion of the CGR environment at a first time. That is, the first (e.g., initial) location is outside the first (e.g., currently displayed) portion of the CGR environment, but the digital assistant object quickly changes position and becomes visible to the user, for example, by quickly and intuitively attracting the user's attention to the digital assistant session without reducing immersion in the CGR environment, thereby increasing the efficiency and effectiveness of the digital assistant session.

[0134] In some embodiments, at block 1212, a second location is determined based on one or more second environmental factors, similar to the determination of the first location at block 1208.

[0135] The one or more second environmental factors may include characteristics of the CGR environment. For example, the second location may be determined based on the static or dynamic location of physical or virtual objects in the CGR environment, such as ensuring that the digital assistant object does not collide (e.g., visually intersect) with those other objects. For example, the second location may be determined to be at or near the location of the electronic device. The device location may be determined based on a pre-identified location (e.g., a user manually identifying and tagging the device location), based on visual analysis of image sensor data (e.g., analyzing image sensor data to recognize or visually understand the device), based on analysis of other sensor or connection data (e.g., using a Bluetooth connection when in general proximity), etc.

[0136] The one or more second environmental factors may also include the position (e.g., location and / or posture) of the user. For example, the first location may be determined to be a location that is in front of and / or visible to the user based on the orientation of the user's body or head, or the position of the user's gaze.

[0137] In some embodiments, animating the digital assistant object to move to a second location within the first portion of the CGR environment includes animating the digital assistant object to disappear at the first location and reappear at the second location (e.g., teleport from one location to another, either instantly or with a predetermined delay).

[0138] In some embodiments, in block 1214, a travel path is determined. In some embodiments, the travel path corresponds to a portion of a path between a first location and a second location that is within a first (e.g., currently displayed) portion of the CGR environment at a first time. That is, animating a digital assistant object moving to a second location includes determining the visible movement the digital assistant object should take to reach the second location. For example, the digital assistant object may simply take the shortest (e.g., most direct) path, or a longer path that achieves a particular visual effect, such as a smooth "flying" path, a path with unusual movements such as bouncing, bouncing, and winding.

[0139] In some embodiments, the movement path is determined so that it does not pass through at least one additional object located in a first (e.g., currently displayed) portion of the CGR environment at the first time. For example, the movement path of the digital assistant object may be context-aware and move to avoid or dodge other physical or virtual objects in the CGR environment. The digital assistant object may also be able to dodge dynamic objects, such as a (real, physical) pet running into the room, another user of the CGR environment, or a moving virtual object (such as virtual objects 1118A-C in FIGS. 4A-4F).

[0140] In some embodiments, in block 1216, a second user input is detected (e.g., at some point after the start of the first digital assistant session). For example, the second user input may include a verbal input such as a verbal command, a question, a request, a shortcut, etc. As another example, the second user input may also include a gesture input such as a gesture command, a question, a request, a gesture representing an interaction with the CGR environment (e.g., "grabbing" and "dropping" a virtual object), etc. As another example, the second user input may include a gaze input.

[0141] In some embodiments, at block 1218, the intent of the second user input is determined. The intent may correspond to one or more tasks that may be performed using one or more parameters. For example, if the second user input includes a verbal or gestural command, question, or request, the intent may be determined using natural language processing techniques, such as determining the intent to play audio from the verbal user input "sit on the shelf" shown in FIG. 4C. As another example, if the second user input includes a gesture, the intent may be determined based on the gesture type(s), the gesture location(s), and / or the locations of various objects within the CGR environment. For example, a grab-and-drop type gesture may correspond to the intent to move a virtual object located at or near the location of the "grab" gesture to the location of the "drop" gesture.

[0142] In some embodiments, the determination of the intent of the second user input is made solely according to a determination that the current location of the digital assistant object is within the currently displayed portion of the CGR environment at the time the second user input is received. That is, some detected user inputs are not intended for the digital assistant session, such as when the user talks to another person in a physical room or another player in a multiplayer game. Processing and responding only to user inputs received while the user is looking at (or near) the digital assistant object reduces the likelihood of unintended interactions, for example, improving the efficiency of the digital assistant session.

[0143] In some embodiments, in block 1220, a first output is provided based on the determined intent. Providing the first output may include performing one or more tasks corresponding to the determined intent. For example, as shown in FIGS. 4A-4F, for a verbal user input "sit on the shelf," the second output may include rearranging the digital assistant object 1120 on the shelf in the CGR environment. As another example, for a gesture input "What time is it?", the first output may include displaying a widget showing a clock face.

[0144] In some embodiments, providing the second output based on the determined intent includes determining whether the determined intent is related to relocating the digital assistant object. For example, the second user input may include an explicit request to relocate the digital assistant, such as the verbal output "Sit on the shelf" or the grab and drop gesture input shown in Figures 4A-4F. As another example, the second user input may include a specific gaze input that starts from the current location of the digital assistant object, moves across the displayed portion of the CGR environment, and stops at the new location.

[0145] According to a determination that the determined intention is related to rearranging the digital assistant object, a third location is determined from the second user input. For example, a third location "on the shelf" can be determined from the verbal input "sit on the shelf," a third location located at or near the "drop" gesture can be determined from the grab and drop gesture input, or a third location located at or near the user's gaze can be determined from the gaze input.

[0146] Further, in accordance with a determination that the determined intention is related to relocating the digital assistant object, the digital assistant object is placed at a second location. In accordance with a determination that a third location is within the currently displayed (e.g., second) portion of the CGR environment (e.g., at a second time), the digital assistant object is displayed at a third location within the CGR environment. That is, while relocating the digital assistant object from its location when the second user input was detected to the third location determined from the second user input, the digital assistant object is displayed as long as its location remains within the user's current field of view (e.g., including animating the movement of the digital assistant object from one location to another).

[0147] In some embodiments, providing a first output based on the determined intent includes determining whether the determined intent is related to an object (e.g., a physical or virtual object) located at an object location within the CGR environment. In accordance with a determination that the determined intent is related to an object located at the object location within the CGR environment, the digital assistant is placed at a fourth location near the object location. That is, the digital assistant object moves near the associated object to indicate an interaction with the object and / or draw attention to the object or interaction. For example, in the CGR environment 1104 shown in FIGS. 4A-4F, in response to a verbal user input such as "What's on the console table?", the digital assistant object 1120 can move near the virtual object 1118C on the console table and then provide more information about the virtual object.

[0148] In some embodiments, in block 1222, a second output is provided based on one or more characteristics of the first output. The one or more characteristics of the second output may include the type of second output (e.g., visual, auditory, or tactile), the location of the second output within the CGR environment (e.g., for a task performed within the CGR environment), etc. For example, as shown in FIGS. 4A-4F, in response to a verbal user input "Please sit on the shelf," the second output may include the verbal output "OK!" along with actual movement to the shelf. As another example, in response to a gestural input "What time is it?", the second output may include animating a digital assistant object bouncing or wiggling near a displayed clock widget (e.g., the first output). Providing a second output as described improves the efficiency of a digital assistant session, for example, by drawing the user's attention to the performance and / or completion of a requested task when that performance and / or completion is not immediately apparent to the user.

[0149] In some embodiments, in block 1224, a fourth output selected from two or more different outputs is provided to indicate a state of the first digital assistant session selected from two or more different states. The two or more different states of the first digital assistant session may include one or more listening states (e.g., active or passive listening), one or more response states, one or more processing (e.g., thinking) states, one or more failure states, one or more attention acquisition states, and / or one or more transition (e.g., moving, appearing, or disappearing) states. There may be a one-to-one correspondence between different outputs and different states, and one or more states may be represented by the same output, or one or more outputs may represent the same state (or variations of the same state).

[0150] For example, as shown in Figures 4A-4F, the digital assistant object 1120 initially attracts the user's attention (e.g., moves to a second location position) and responds to the user input / takes on a different appearance while attracting the user's attention to the response. Other fourth (e.g., state) outputs may include changes in the size of the digital assistant object, animation of movement, auditory output (e.g., directional audio output "emanating" from the current location of the digital assistant object), changes in lighting effects, other visual output, tactile output, etc.

[0151] At any time during the first digital assistant session, if the currently displayed portion of the CGR environment is updated (e.g., as the user changes their viewpoint) and the current location of the digital assistant object is no longer included in the currently displayed portion (e.g., no longer visible to the user), an output indicating the current position of the digital assistant may be provided. For example, additional outputs indicating the current position may be provided as described above with respect to FIGS. 3A-3B and block 1008 (e.g., spatial auditory output, visual output, tactile output, etc.).

[0152] In some embodiments, the first digital assistant session ends, for example, upon explicit exit by the user or after a threshold period has elapsed without interaction. When the first digital assistant session ends, the digital assistant object is dismissed (e.g., removed from the CGR environment) in block 1226. If the digital assistant object is located within a displayed portion of the CGR environment at the time the digital assistant session ends, dismissing the digital assistant object includes ceasing to display the digital assistant object.

[0153] In some embodiments, when the digital assistant object is dismissed, a third output indicating the dismissal of the digital assistant object is provided. The third output may include one or more visual outputs (e.g., a dim light, restoring the lighting in the CGR environment to the state before the digital assistant session, etc.), one or more auditory outputs (e.g., a verbal output such as "bye," a chime, etc.), one or more tactile outputs, etc. For example, as shown in FIG. 4E, providing the third output may include placing a digital assistant indicator 1122 in a third location. The third output may indicate to the user that the first digital assistant session has ended and may help the user more quickly locate the digital assistant object (e.g., find the digital assistant object) in subsequent calls.

[0154] In some embodiments, after the digital assistant object is dismissed, a third user input is detected. The third user input may be an auditory input, a gaze input, a gesture input, or a device input (e.g., a button press, a tap, a swipe, etc.) as described with respect to the first user input (e.g., the user input invoking the first digital assistant session). In accordance with a determination that the third user input satisfies at least one criterion for initiating a digital assistant session (e.g., as described above with respect to block 1204), a second digital assistant session is initiated.

[0155] In an embodiment in which a second user input related to relocating a digital assistant object is received during the first digital assistant session, starting the second digital assistant session involves placing the digital assistant object in a third location (eg, the location requested by the second user input) (eg, generating an instance of it). That is, as shown in FIGS. 4A-4F, after the user explicitly moves the digital assistant object, the location selected by the user becomes the new "default" location where the digital assistant object will appear on subsequent calls.

[0156] The processes described above with reference to Figures 5A-5B are optionally performed by the components shown in Figures 1A-1B. For example, the operations of the illustrated methods may be performed by an electronic device (e.g., 800a, 800b, 800c, or 906), such as one implementing system 700. It will be apparent to one skilled in the art how other processes may be performed based on the components shown in Figures 1A-1B.

[0157] According to some implementations, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided that stores one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for performing any of the methods or processes described herein.

[0158] According to some implementations, there is provided an electronic device (eg, a portable electronic device) comprising means for performing any of the methods or processes described herein.

[0159] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes a processing unit configured to perform any of the methods or processes described herein.

[0160] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes one or more processors and a memory that stores one or more programs for execution by the one or more processors, the one or more programs including instructions for performing any of the methods or processes described herein.

[0161] The foregoing has been described with reference to specific embodiments for purposes of explanation. However, the exemplary discussion above is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described to best explain the principles of the technology and its practical applications so that others skilled in the art can best utilize the technology and various embodiments with various modifications as suited to the particular applications intended.

[0162] Although the present disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art, and such changes and modifications are to be understood as being included within the scope of the present disclosure and examples, as defined by the claims.

Claims

1. An electronic device having one or more processors, a memory, a display, and one or more sensors, while displaying a portion of a computer-generated reality (CGR) environment representing a current field of view of a user of said electronic device; Detecting a first user input using the one or more sensors; Initiating a first digital assistant session according to a determination that the first user input meets at least one criterion for initiating a digital assistant session, wherein initiating the first digital assistant session includes placing a digital assistant object in the CGR environment at a first location outside the displayed portion of the CGR environment at a first time; Providing a first output indicating the first location of the digital assistant object within the CGR environment; A method comprising:

2. The method of claim 1 , wherein the first user input comprises an auditory input.

3. The method of claim 2 , wherein the auditory input comprises a speech input that includes a trigger phrase.

4. The method of any one of claims 1 to 3, wherein the first user input comprises an eye-gaze input.

5. The method of any one of claims 1 to 4, wherein the first user input comprises a gesture input.

6. determining the first location within the CGR environment based on one or more environmental factors. The method according to any one of claims 1 to 5.

7. The method of claim 6 , wherein the one or more environmental factors include characteristics of the CGR environment.

8. The method of any one of claims 6 to 7, wherein the one or more environmental factors include a position of the user of the electronic device.

9. The method of any one of claims 6 to 8, wherein the one or more environmental factors include multiple positions of multiple users in the CGR environment.

10. The method of any preceding claim, wherein providing the first output indicative of the first location comprises generating a first audible output.

11. the first aural output includes at least a first audio component produced by a first speaker of a plurality of speakers and a second audio component produced by a second speaker of the plurality of speakers, and providing the first output indicative of the first location includes: determining whether the first location is closer to a location of the first speaker of the plurality of speakers or closer to a location of the second speaker of the plurality of speakers; generating the first audio component at a louder volume than the second audio component in accordance with determining that the first location is closer to the location of the first speaker than the location of the second speaker; and generating the second audio component louder than the first audio component in accordance with a determination that the first location is closer to the location of the second speaker than to the location of the first speaker.

12. The method of any preceding claim, wherein providing the first output indicative of the first location comprises generating a first tactile output.

13. 13. The method of any one of claims 1 to 12, wherein providing the first output indicative of the first location comprises displaying a visual indication of the first location on the display.

14. At a second time, Further comprising: displaying the digital assistant object at the first location according to a determination that the first location is within the displayed portion of the CGR environment at the second time; The method according to any one of claims 1 to 13.

15. detecting a second user input at a third time with the one or more sensors; and determining the intent of the second user input; and providing a second output based on the determined intent. The method according to any one of claims 1 to 14.

16. Detecting the second user input is performed while the digital assistant object is located at the first location, and determining the intention and providing the second output based on the determined intention are performed in accordance with the determination that the first location is within the displayed portion of the CGR environment at the third time. The method of claim 15.

17. Providing the second output based on the determined intent comprises: According to a determination that the determined intention is related to rearranging the digital assistant object, determining a second location from the second user input; and Placing the digital assistant object at the second location; At a fourth time, the digital assistant object is displayed at the second location within the CGR environment in accordance with a determination that the second location is within the displayed portion of the CGR environment at the fourth time. A method according to any one of claims 15 to 16, comprising:

18. After placing the digital assistant object at the second location, leaving the digital assistant object; detecting a third user input with the one or more sensors; Initiating a second digital assistant session in accordance with a determination that the third user input satisfies at least one criterion for initiating a digital assistant session, wherein initiating the second digital assistant session includes placing the digital assistant object at the second location.

18. The method of claim 17.

19. Providing the second output based on the determined intent comprises: According to a determination that the determined intention is related to an object located at an object location in the CGR environment, the digital assistant object is placed at a third location near the object location. A method according to any one of claims 15 to 16.

20. while providing the second output, providing a third output based on one or more characteristics of the second output.

20. The method according to any one of claims 14 to 19.

21. Leaving the digital assistant object; Further comprising: providing a fourth output indicating the departure of the digital assistant object; 21. The method according to any one of claims 1 to 20.

22. Providing the fourth output comprises: The method of claim 21 , comprising placing a digital assistant indicator at the first location within the CGR environment.

23. Further comprising providing a fifth output selected from two or more different outputs indicating a state selected from two or more different states of the first digital assistant session; 23. The method according to any one of claims 1 to 22.

24. An electronic device having one or more processors, a memory, a display, and one or more sensors, Detecting a first user input using the one or more sensors; Initiating a first digital assistant session in accordance with a determination that the first user input meets criteria for initiating a digital assistant session, wherein starting the first digital assistant session While displaying a first portion of a computer-generated reality (CGR) environment on the display, placing a digital assistant object at a first location within the CGR environment and outside the first portion of the CGR environment; Starting includes providing a first output indicating the first location of the digital assistant object within the CGR environment; A method comprising:

25. The method of claim 24 , wherein the first user input comprises an auditory input.

26. The method of claim 25 , wherein the auditory input comprises a speech input that includes a trigger phrase.

27. The method of any one of claims 24 to 26, wherein the first user input comprises an eye gaze input.

28. The method of any one of claims 24 to 27, wherein the first user input comprises a gesture input.

29. determining the first location within the CGR environment based on one or more environmental factors. The method according to any one of claims 24 to 28.

30. 30. The method of claim 29, wherein the one or more environmental factors include characteristics of the CGR environment.

31. The method of any one of claims 29 to 30, wherein the one or more environmental factors include a position of a user of the electronic device.

32. The method of any one of claims 29 to 31, wherein the one or more environmental factors include multiple positions of multiple users in the CGR environment.

33. Providing the first output indicative of the first location comprises: and generating a first auditory output.

34. the first aural output includes at least a first audio component produced by a first speaker of a plurality of speakers and a second audio component produced by a second speaker of the plurality of speakers, and providing the first output indicative of the first location includes: determining whether the first location is closer to a location of the first speaker of the plurality of speakers or closer to a location of the second speaker of the plurality of speakers; generating the first audio component at a louder volume than the second audio component in accordance with determining that the first location is closer to the location of the first speaker than the location of the second speaker; 34. The method of claim 33, further comprising: producing the second audio component at a louder volume than the first audio component in accordance with a determination that the first location is closer to the location of the second speaker than to the location of the first speaker.

35. Providing the first output indicative of the first location comprises: and generating a first haptic output.

36. The method of any one of claims 24 to 35, wherein providing the first output indicative of the first location comprises providing a visual representation of the first location.

37. detecting a second user input using the one or more sensors; determining the intent of the second user input; and providing a second output based on the determined intent. The method according to any one of claims 24 to 36.

38. displaying a second portion of the CGR environment on the display; While the digital assistant object is located at the first location, the digital assistant object is displayed at the first location within the CGR environment in accordance with a determination that the first location is within the second portion of the CGR environment.

38. The method of claim 37.

39. Detecting the second user input is performed while the digital assistant object is located at the first location, and determining the intention and providing the second output based on the determined intention are performed in accordance with the determination that the first location is within the second portion of the CGR environment. The method of claim 38.

40. Providing the second output based on the determined intent comprises: According to a determination that the determined intention is related to rearranging the digital assistant object, determining a second location from the second user input; and Placing the digital assistant object at the second location; While displaying the second portion of the CGR environment, in accordance with a determination that the second location is within the second portion of the CGR environment, displaying the digital assistant object at the second location within the CGR environment. A method according to any one of claims 38 to 39, comprising:

41. After displaying the digital assistant object at the second location, the digital assistant object leaves; detecting a third user input with the one or more sensors; Initiating a second digital assistant session in accordance with a determination that the third user input meets the criteria for initiating a digital assistant session, wherein starting the second digital assistant session includes placing the digital assistant object at the second location.

41. The method of claim 40.

42. Providing the second output based on the determined intent comprises: and placing the digital assistant object at a third location near the object location in accordance with a determination that the determined intention is related to an object located at an object location in the CGR environment. A method according to any one of claims 37 to 39, comprising:

43. while providing the second output, providing a third output based on one or more characteristics of the second output.

43. The method according to any one of claims 37 to 42.

44. While the digital assistant object is located at the first location, Leaving the digital assistant object; Further comprising: providing a fourth output indicating the departure of the digital assistant object; 44. The method according to any one of claims 24 to 43.

45. Providing the fourth output comprises:

45. The method of claim 44, including placing a digital assistant indicator at the first location within the CGR environment.

46. Further comprising providing a fifth output selected from two or more different outputs indicating a state selected from two or more different states of the first digital assistant session; 46. ​​The method according to any one of claims 24 to 45.

47. An electronic device having one or more processors, a memory, a display, and one or more sensors, While displaying a portion of a computer generated reality (CGR) environment on said display, Detecting a first user input using the one or more sensors; Initiating a first digital assistant session in accordance with a determination that the first user input meets criteria for initiating a digital assistant session, wherein starting the first digital assistant session Creating an instance of a digital assistant object at a first location within the CGR environment at a first time and outside the portion of the CGR environment; Starting includes animating the digital assistant object to move to a second location within the portion of the CGR environment at the first time; A method comprising:

48. 48. The method of claim 47, wherein the first user input comprises a voice input.

49. 49. The method of claim 48, wherein the speech input includes a trigger phrase.

50. The method of any one of claims 47 to 49, wherein the first user input comprises an eye gaze input.

51. The method of any one of claims 47 to 50, wherein the first user input comprises a gesture input.

52. determining the first location within the CGR environment based on one or more first environmental factors; 52. The method of any one of claims 47 to 51.

53. 53. The method of claim 52, wherein the one or more first environmental factors include a first characteristic of the CGR environment.

54. The method of any one of claims 52 to 53, wherein the one or more first environmental factors include a first position of a user of the electronic device.

55. determining the second location within the CGR environment based on one or more second environmental factors; 55. The method of any one of claims 47 to 54.

56. 56. The method of claim 55, wherein the one or more second environmental factors include a second characteristic of the CGR environment.

57. 57. The method of any one of claims 55 to 56, wherein the one or more second environmental factors include a first position of a user of the electronic device.

58. The method according to any one of claims 47 to 57, wherein animating the digital assistant object to move to the second location includes determining a movement path.

59. 59. The method of claim 58, wherein the travel path corresponds to a portion of a path between the first location and the second location, the portion of the path being within the portion of the CGR environment at the first time.

60. A method according to any one of claims 58 to 59, wherein the travel path does not pass through any additional objects located within the portion of the CGR environment at the first time.

61. detecting a second user input using the one or more sensors; determining the intent of the second user input; providing a first output based on the determined intent.

61. The method of any one of claims 47 to 60.

62. Providing the first output based on the determined intent comprises: According to a determination that the determined intention is related to rearranging the digital assistant object, determining a third location from the second user input; and Placing the digital assistant object at the third location; At a second time, the digital assistant object is displayed at the third location within the CGR environment in accordance with a determination that the third location is within the portion of the CGR environment at the second time. The method of any one of claims 61 to 65, comprising:

63. After displaying the digital assistant object at the third location, the digital assistant object is dismissed; detecting a third user input with the one or more sensors; Initiating a second digital assistant session in accordance with a determination that the third user input meets the criteria for initiating a digital assistant session, wherein initiating the second digital assistant session includes placing the digital assistant object at the third location.

63. The method of claim 62.

64. Leaving the digital assistant object 64. The method of claim 63, including placing a digital assistant indicator at the third location within the CGR environment.

65. Providing the first output based on the determined intent comprises:

62. The method of claim 61, further comprising: placing the digital assistant object at a fourth location near the object location in accordance with a determination that the determined intention is associated with an object located at an object location within the CGR environment.

66. while providing the first output, providing a second output based on one or more characteristics of the first output.

66. The method of any one of claims 61 to 65.

67. At a third time while the digital assistant object is located at the second location, Leaving the digital assistant object; Further comprising: providing a third output indicating the departure of the digital assistant object; 67. The method of any one of claims 47 to 66.

68. Further comprising providing a fourth output selected from two or more different outputs indicating a state selected from two or more different states of the first digital assistant session; 22. The method according to any one of claims 1 to 21.

69. 1. An electronic device comprising: The display and one or more sensors; one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: while displaying a portion of a computer-generated reality (CGR) environment representing a current field of view of a user of said electronic device; instructions for detecting a first user input using the one or more sensors; In accordance with a determination that the first user input satisfies at least one criterion for starting a digital assistant session, instructions for starting a first digital assistant session, wherein starting the first digital assistant session includes placing a digital assistant object in the CGR environment at a first location outside the displayed portion of the CGR environment at a first time; and providing a first output indicating the first location of the digital assistant object within the CGR environment. An electronic device.

70. 1. An electronic device comprising: The display and one or more sensors; one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: instructions for detecting a first user input using the one or more sensors; An instruction to start a first digital assistant session in accordance with a determination that the first user input meets criteria for starting a digital assistant session, wherein starting the first digital assistant session While displaying a first portion of a computer-generated reality (CGR) environment on the display, placing a digital assistant object at a first location within the CGR environment and outside the first portion of the CGR environment; and providing a first output indicating the first location of the digital assistant object within the CGR environment. An electronic device comprising:

71. 1. An electronic device comprising: The display and one or more sensors; one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: While displaying a portion of a computer generated reality (CGR) environment on said display, instructions for detecting a first user input using the one or more sensors; An instruction to start a first digital assistant session in accordance with a determination that the first user input meets criteria for starting a digital assistant session, wherein starting the first digital assistant session Creating an instance of a digital assistant object at a first location within the CGR environment at a first time and outside the portion of the CGR environment; Animating the digital assistant object moving to a second location within the portion of the CGR environment at the first time. An electronic device comprising:

72. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and one or more sensors, the one or more programs comprising: while displaying a portion of a computer-generated reality (CGR) environment representing a current field of view of a user of said electronic device; instructions for detecting a first user input using the one or more sensors; In accordance with a determination that the first user input satisfies at least one criterion for starting a digital assistant session, instructions for starting a first digital assistant session, wherein starting the first digital assistant session includes placing a digital assistant object in the CGR environment at a first location outside the displayed portion of the CGR environment at a first time; and providing a first output indicating the first location of the digital assistant object within the CGR environment. A non-transitory computer-readable storage medium.

73. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and one or more sensors, the one or more programs comprising: instructions for detecting a first user input using the one or more sensors; An instruction to start a first digital assistant session in accordance with a determination that the first user input meets criteria for starting a digital assistant session, wherein starting the first digital assistant session While displaying a first portion of a computer-generated reality (CGR) environment on the display, placing a digital assistant object at a first location within the CGR environment and outside the first portion of the CGR environment; and providing a first output indicating the first location of the digital assistant object within the CGR environment.

74. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and one or more sensors, the one or more programs comprising: While displaying a portion of a computer generated reality (CGR) environment on said display, instructions for detecting a first user input using the one or more sensors; An instruction to start a first digital assistant session in accordance with a determination that the first user input meets criteria for starting a digital assistant session, wherein starting the first digital assistant session Creating an instance of a digital assistant object at a first location within the CGR environment at a first time and outside the portion of the CGR environment; and animating the digital assistant object to move to a second location within the portion of the CGR environment at the first time. A non-transitory computer-readable storage medium comprising:

75. 1. An electronic device comprising: while displaying a portion of a computer-generated reality (CGR) environment representing a current field of view of a user of said electronic device; means for detecting a first user input using the one or more sensors; A means for initiating a first digital assistant session according to a determination that the first user input satisfies at least one criterion for initiating a digital assistant session, wherein initiating the first digital assistant session includes placing a digital assistant object in the CGR environment at a first location outside the displayed portion of the CGR environment at a first time; and means for providing a first output indicating the first location of the digital assistant object within the CGR environment. An electronic device.

76. 1. An electronic device comprising: means for detecting a first user input using the one or more sensors; A means for starting a first digital assistant session in accordance with a determination that the first user input meets criteria for starting a digital assistant session, wherein starting the first digital assistant session includes: While displaying a first portion of a computer-generated reality (CGR) environment on the display, placing a digital assistant object at a first location within the CGR environment and outside the first portion of the CGR environment; and providing a first output indicating the first location of the digital assistant object within the CGR environment. An electronic device comprising:

77. 1. An electronic device comprising: While displaying a portion of a computer generated reality (CGR) environment on said display, means for detecting a first user input using the one or more sensors; A means for starting a first digital assistant session in accordance with a determination that the first user input meets criteria for starting a digital assistant session, wherein starting the first digital assistant session includes: Creating an instance of a digital assistant object at a first location within the CGR environment at a first time and outside the portion of the CGR environment; An electronic device comprising: an initiating means for animating the digital assistant object to move to a second location within the portion of the CGR environment at the first time.

78. 1. An electronic device comprising: The display and one or more sensors; one or more processors; and a memory storing one or more programs configured to be executed by said one or more processors, said one or more programs comprising instructions for performing the method of any one of claims 1 to 23.

79. 1. An electronic device comprising: The display and one or more sensors; one or more processors; and a memory storing one or more programs configured to be executed by said one or more processors, said one or more programs comprising instructions for performing the method of any one of claims 24 to 46.

80. 1. An electronic device comprising: The display and one or more sensors; one or more processors; and a memory storing one or more programs configured to be executed by said one or more processors, said one or more programs comprising instructions for performing the method of any one of claims 47 to 68.

81. 24. A non-transitory computer readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and one or more sensors, the one or more programs comprising instructions for performing the method of any one of claims 1 to 23.

82. 47. A non-transitory computer readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and one or more sensors, the one or more programs comprising instructions for performing the method of any one of claims 24 to 46.

83. 69. A non-transitory computer readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and one or more sensors, the one or more programs comprising instructions for performing the method of any one of claims 47 to 68.

84. comprising means for carrying out the method according to any one of claims 1 to 23, Electronic devices.

85. comprising means for carrying out the method according to any one of claims 25 to 46, Electronic devices.

86. comprising means for carrying out the method according to any one of claims 47 to 68, Electronic devices.