Devices, methods, and graphical user interfaces for navigating, inputting, or modifying content.

Improved navigation and interaction methods in virtual and augmented reality systems using gaze, gesture, and voice inputs address inefficiencies, enhancing user experience and reducing input complexity while conserving power.

JP2026062738APending Publication Date: 2026-04-10APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
APPLE INC
Filing Date
2025-12-18
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for interacting with virtual and augmented reality environments are cumbersome, inefficient, and impose a significant cognitive burden on users, leading to time-consuming and error-prone operations.

Method used

Implementing computer systems with improved navigation, editing, and interaction methods using gaze-based, gesture-based, and voice-based inputs, along with advanced graphical user interfaces, to enhance user interaction efficiency and reduce the number and complexity of inputs.

Benefits of technology

The improved methods and interfaces reduce user input complexity, conserve power in battery-operated devices, and enhance the overall user experience by providing intuitive and efficient interaction with virtual and augmented reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062738000001_ABST
    Figure 2026062738000001_ABST
Patent Text Reader

Abstract

To provide a computer system with improved methods and interfaces for scrolling, creating, editing, and navigating content that are more efficient and intuitive for the user. [Solution] The computer system scrolls scrollable content in response to various user inputs, inputs text into text entry fields in response to voice input, facilitates interaction with a soft keyboard, facilitates interaction with a cursor, facilitates text deletion, and facilitates interaction with hardware input devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 266,357, filed on January 3, 2022; U.S. Provisional Patent Application No. 63 / 337,539, filed on May 2, 2022; and U.S. Provisional Patent Application No. 63 / 377,025, filed on September 24, 2022, the contents of which are hereby incorporated by reference in their entirety for all purposes.

[0002] This disclosure generally relates to computer systems that provide computer - generated experiences, including but not limited to, electronic devices that provide virtual and mixed - reality experiences via display - generation components.

Background Art

[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary extended - reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch - sensing surfaces, and touch - screen displays for computer systems and other electronic computing devices are used to interact with virtual / extended - reality environments. Exemplary virtual elements include virtual objects such as digital images, videos, text, icons, and control elements such as buttons and other graphics.

Summary of the Invention

[0004] Some methods and interfaces for navigating and editing content are cumbersome, inefficient, and restrictive. For example, systems for scrolling through content, adding and editing text, and performing actions using a cursor are complex, tedious, error-prone, and impose a significant cognitive burden on the user, detracting from the virtual / augmented reality experience. In addition, these methods are unnecessarily time-consuming, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.

[0005] Therefore, there is a need for computer systems with improved methods and interfaces for scrolling, creating, editing, and navigating content that are more efficient and intuitive for the user. Such methods and interfaces may optionally complement or replace conventional methods for performing such operations. Such methods and interfaces reduce the number, degree, and / or type of user input by helping the user understand the connection between the inputs provided and the device's response to those inputs, thereby generating a more efficient human-machine interface.

[0006] The above-mentioned drawbacks and other problems associated with the user interface of a computer system are mitigated or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, tablet computer, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to display-generating components, the output devices include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI (and / or computer system) through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the user's body as captured by a camera and other motion sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gameplay, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing those functions optionally reside in a primary computer-readable storage medium and / or a non-primary computer-readable storage medium, or in other computer program products configured to be executed by one or more processors.

[0007] As described above, there is a need for electronic devices with improved methods and interfaces for interacting with content. Such methods and interfaces can complement or replace conventional methods for interacting with content. Such methods and interfaces reduce the number, degree, and / or type of user input, resulting in a more efficient human-machine interface. In the case of battery-operated computing devices, such methods and interfaces conserve power and extend the interval between battery charges.

[0008] In some embodiments, the computer system scrolls scrollable content in response to various user inputs. In some embodiments, the computer system inputs text into a text entry field in response to voice input. In some embodiments, the computer system facilitates interaction with a soft keyboard. In some embodiments, the computer system facilitates interaction with a cursor. In some embodiments, the computer system facilitates deletion of text from a text entry field. In some embodiments, the computer system facilitates interaction with a hardware input device.

[0009] It should be noted that the various embodiments described herein can be combined with any other embodiments described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, in particular, in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention. [Brief explanation of the drawing]

[0010] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts.

[0011] [Figure 1] This block diagram shows the operating environment of a computer system for providing an XR experience, according to several embodiments.

[0012] [Figure 2] Block diagram showing a controller for a computer system configured to manage and adjust the user's XR experience, according to several embodiments.

[0013] [Figure 3] This block diagram shows display generation components of a computer system configured to provide users with visual components of an XR experience, according to several embodiments.

[0014] [Figure 4] This is a block diagram showing a hand tracking unit for a computer system configured to capture user gesture input, according to several embodiments.

[0015] [Figure 5]A block diagram showing an eye-tracking unit of a computer system configured to capture a user's eye gaze input according to some embodiments.

[0016] [Figure 6] A flowchart showing a grint-assisted eye-tracking pipeline according to some embodiments.

[0017] [Figure 7A] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments. [Figure 7B] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments. [Figure 7C] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments. [Figure 7D] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments. [Figure 7E] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments. [Figure 7F] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments. [Figure 7G] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments. [Figure 7H] Illustrative techniques for scrolling scrollable content in response to various user inputs according to some embodiments.

[0018] [Figure 8A]A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8B] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8C] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8D] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8E] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8F] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8G] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8H] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8I] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8J] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8K] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments. [Figure 8L] A flowchart of a method for scrolling scrollable content in response to various user inputs according to various embodiments.

[0019] [Figure 9A]This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9B] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9C] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9D] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9E] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9F] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9G] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9H] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9I] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9J] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9K] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9L] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9M]This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments. [Figure 9N] This document illustrates exemplary techniques for entering text into a text entry field in response to voice input, based on several embodiments.

[0020] [Figure 10A] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10B] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10C] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10D] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10E] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10F] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10G] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10H] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10I] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10J] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10K] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10L]This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10M] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10N] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10O] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10P] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10Q] This is a flowchart illustrating various methods for entering text into a text entry field. [Figure 10R] This is a flowchart illustrating various methods for entering text into a text entry field.

[0021] [Figure 11A] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11B] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11C] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11D] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11E] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11F] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11G] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11H] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11I] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11J] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11K] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11L] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11M] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11N] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 11O] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments.

[0022] [Figure 12A] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12B] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12C] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12D] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12E] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12F]This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12G] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12H] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12I] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12J] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12K] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12L] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12M] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12N] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12O] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 12P] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments.

[0023] [Figure 13A] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 13B] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 13C] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 13D] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 13E] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments.

[0024] [Figure 14A] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14B] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14C] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14D] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14E] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14F] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14G] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14H] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14I] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 14J] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments.

[0025] [Figure 15A] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 15B]This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 15C] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 15D] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 15E] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 15F] This document presents exemplary techniques for facilitating interaction with a soft keyboard, based on several embodiments.

[0026] [Figure 16A] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16B] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16C] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16D] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16E] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16F] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16G] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16H] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16I] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16J] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments. [Figure 16K] This is a flowchart illustrating a method for facilitating interaction with a soft keyboard, based on several embodiments.

[0027] [Figure 17A] This document illustrates exemplary techniques for facilitating interaction with the cursor, based on several embodiments. [Figure 17B] This document illustrates exemplary techniques for facilitating interaction with the cursor, based on several embodiments. [Figure 17C] This document illustrates exemplary techniques for facilitating interaction with the cursor, based on several embodiments. [Figure 17D] This document illustrates exemplary techniques for facilitating interaction with the cursor, based on several embodiments. [Figure 17E] This document illustrates exemplary techniques for facilitating interaction with the cursor, based on several embodiments. [Figure 17F] This document illustrates exemplary techniques for facilitating interaction with the cursor, based on several embodiments.

[0028] [Figure 18A] This is a flowchart illustrating a method for facilitating interaction with the cursor, based on several embodiments. [Figure 18B] This is a flowchart illustrating a method for facilitating interaction with the cursor, based on several embodiments. [Figure 18C] This is a flowchart illustrating a method for facilitating interaction with the cursor, based on several embodiments. [Figure 18D] This is a flowchart illustrating a method for facilitating interaction with the cursor, based on several embodiments. [Figure 18E] This is a flowchart illustrating a method for facilitating interaction with the cursor, based on several embodiments.

[0029] [Figure 19A]This describes an exemplary technique, based on several embodiments, of entering text into a text entry field in response to receiving utterance input. [Figure 19B] This describes an exemplary technique, based on several embodiments, of entering text into a text entry field in response to receiving utterance input. [Figure 19C] This describes an exemplary technique, based on several embodiments, of entering text into a text entry field in response to receiving utterance input. [Figure 19D] This describes an exemplary technique, based on several embodiments, of entering text into a text entry field in response to receiving utterance input. [Figure 19E] This describes an exemplary technique, based on several embodiments, of entering text into a text entry field in response to receiving utterance input. [Figure 19F] This describes an exemplary technique, based on several embodiments, of entering text into a text entry field in response to receiving utterance input. [Figure 19G] This describes an exemplary technique, based on several embodiments, of entering text into a text entry field in response to receiving utterance input.

[0030] [Figure 20A] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20B] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20C] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20D] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20E]This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20F] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20G] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20H] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20I] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20J] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20K] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20L] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments. [Figure 20M] This is a flowchart illustrating a method for entering text into a text entry field in response to receiving utterance input, according to several embodiments.

[0031] [Figure 21A] This document illustrates exemplary techniques for modifying text contained in a text entry field, using several embodiments. [Figure 21B] This document illustrates exemplary techniques for modifying text contained in a text entry field, using several embodiments. [Figure 21C]This document illustrates exemplary techniques for modifying text contained in a text entry field, using several embodiments. [Figure 21D] This document illustrates exemplary techniques for modifying text contained in a text entry field, using several embodiments. [Figure 21E] This document illustrates exemplary techniques for modifying text contained in a text entry field, using several embodiments. [Figure 21F] This document illustrates exemplary techniques for modifying text contained in a text entry field, using several embodiments. [Figure 21G] This document illustrates exemplary techniques for modifying text contained in a text entry field, using several embodiments.

[0032] [Figure 22A] This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments. [Figure 22B] This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments. [Figure 22C] This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments. [Figure 22D] This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments. [Figure 22E] This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments. [Figure 22F] This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments. [Figure 22G] This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments. [Figure 22H]This is a flowchart illustrating a method for modifying text contained in a text entry field, according to several embodiments.

[0033] [Figure 23A] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23B] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23C] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23D] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23E] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23F] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23G] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23H] This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system. [Figure 23I]This document illustrates an exemplary technique, in several embodiments, for updating user interface elements according to the status of a hardware input device that communicates with a computer system.

[0034] [Figure 24A] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24B] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24C] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24D] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24E] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24F] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24G] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24H] This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Figure 24I]This is a flowchart illustrating a method for updating user interface elements according to the status of a hardware input device communicating with a computer system, according to several embodiments. [Modes for carrying out the invention]

[0035] This disclosure relates to user interfaces that provide users with Extended Reality (XR) experiences, in several embodiments.

[0036] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in multiple ways.

[0037] In some embodiments, the computer system scrolls content in response to various user inputs, such as gaze-based user input and gesture-based user input (e.g., air gesture input, which is described in more detail below). In some embodiments, the computer system presents scrollable content comprising a first region of scrollable content and a second region of scrollable content. In response to detecting the user's attention directed towards the second region of scrollable content, the computer system optionally scrolls the scrollable content to advance the content displayed in the second region toward the first region. In some embodiments, the computer system scrolls content in response to detecting air gesture input, including pinch-and-drag gestures, while the user's attention is directed towards the content.

[0038] In some embodiments, the computer system inputs text into a text entry field in response to voice input, according to some embodiments. Upon detecting user attention directed towards the text entry field, the computer system optionally initiates a process to accept dictation input directed towards the text entry field. The computer system optionally presents (e.g., visual, audio) feedback in response to utterance input directed towards the text entry field and displays a textual representation of the utterance input in the text entry field.

[0039] In some embodiments, the computer system facilitates interaction with a soft keyboard. The computer system optionally displays an object (e.g., a user interface, window, or another container) containing a text entry field located beyond a threshold distance from the user's viewpoint in a three-dimensional environment. In response to input directed at the text entry field, the computer system displays a soft keyboard. In some embodiments, the computer system displays the soft keyboard within the user's threshold distance.

[0040] In some embodiments, the computer system facilitates interaction with a soft keyboard. The computer system optionally displays a soft keyboard without displaying one or more cursors for interacting with the soft keyboard. In some embodiments, the computer system detects user input directed to one or more keys on the soft keyboard provided by a separate part of the user (e.g., the user's hand(s)). The computer system optionally displays the movement of one or more keys away from the separate part of the user toward the surface of the keyboard and performs one or more actions associated with one or more keys on the keyboard in response to the user input directed to one or more keys on the keyboard.

[0041] In some embodiments, the computer system facilitates interaction with a soft keyboard. The computer system optionally displays a soft keyboard along with one or more cursors for interaction with the soft keyboard. The computer system optionally moves the cursors in response to detecting movement of one or more parts of the user (e.g., hands). In some embodiments, in response to detecting input provided by one or more parts of the user corresponding to making a selection using the one or more cursors, the computer system activates one or more keys on the soft keyboard corresponding to the one or more cursors.

[0042] In some embodiments, the computer system facilitates interaction with the cursor. The computer system optionally displays the cursor in a specific area of ​​the three-dimensional environment. In some embodiments, the computer system updates the cursor's position according to the movement of a specific part of the user (e.g., the hand) and the user's attention. While the user's attention is directed to a specific area of ​​the three-dimensional environment and the cursor is displayed in that area, the computer system moves the cursor within the specific area in response to the movement of the specific part of the user. In some embodiments, upon detecting coordinated movement of the specific part of the user and a shift of the user's attention from the specific area to another location in the three-dimensional environment, the computer system displays the cursor in the new area in accordance with the attention and movement of the specific part of the user.

[0043] In some embodiments, the computer system facilitates text entry in response to utterance input. The computer system optionally displays a dictation user interface element overlaid on or at least partially on the text entry field to enable dictation of text into the text entry field. In some embodiments, the computer system enters text into the text entry field in response to confirmation input that the text in the dictation user interface element should be entered into the text entry field. In some embodiments, the computer system refrains from entering text into the text entry field until and after confirmation input has been received.

[0044] In some embodiments, the computer system facilitates the deletion of text from a text entry field. The computer system optionally displays a user interface element associated with a soft keyboard, which includes a text entry field containing a copy of text contained in a second text entry field within the user interface of an application that has the soft keyboard's current focus. In some embodiments, in response to detecting that the user's attention is directed to a portion of the text entry field contained in the user interface element, the computer system displays an option to delete one or more characters from the text entry field. In response to detecting the selection of an option and / or the selection of a portion of the text entry field contained in the user interface element, the computer system deletes one or more characters from the text.

[0045] In some embodiments, the computer system facilitates interaction with hardware input devices. The computer system optionally displays user interface elements that are within the computer system's field of view and have a predetermined spatial relationship to hardware input devices communicating with the computer system. In some embodiments, the user interface elements include a text entry field containing a representation of text contained in a second text entry field of the user interface of an application having the current focus of the hardware input device; an option for displaying a software input element; a dictation option; and an option for inserting suggested text into the text entry field.

[0046] Figures 1-6 illustrate exemplary computer systems for providing an XR experience to a user. Figures 7A-7H show exemplary techniques for scrolling scrollable content in response to various user inputs, according to several embodiments. Figures 8A-8L are flowcharts of methods for scrolling scrollable content in response to various user inputs, according to several embodiments. The user interfaces in Figures 7A-7H are used to illustrate the processes in Figures 8A-8L. Figures 9A-9N show exemplary techniques for entering text into a text entry field in response to voice input, according to several embodiments. Figures 10A-10R are flowcharts of methods for entering text into a text entry field, according to several embodiments. The user interfaces in Figures 9A-9N are used to illustrate the processes in Figures 10A-10R. Figures 11A-11O show exemplary techniques for facilitating interaction with a soft keyboard, according to several embodiments. Figures 12A-12P are flowcharts of methods for facilitating interaction with a soft keyboard, according to several embodiments. The user interfaces in Figures 11A to 11O are used to illustrate the processes in Figures 12A to 12P. Figures 13A to 13E illustrate exemplary techniques for facilitating interaction with a soft keyboard according to several embodiments. Figures 14A to 14J are flowcharts of methods for facilitating interaction with a soft keyboard according to several embodiments. The user interfaces in Figures 13A to 13E are used to illustrate the processes in Figures 14A to 14J. Figures 15A to 15F illustrate exemplary techniques for facilitating interaction with a soft keyboard according to several embodiments. Figures 16A to 16K are flowcharts of methods for facilitating interaction with a soft keyboard according to several embodiments. The user interfaces in Figures 15A to 15F are used to illustrate the processes in Figures 16A to 16K. Figures 17A to 17F illustrate exemplary techniques for facilitating interaction with a cursor according to several embodiments. Figures 18A to 18E are flowcharts of methods for facilitating interaction with a cursor according to several embodiments.The user interfaces in Figures 17A to 17F are used to illustrate the processes in Figures 18A to 18E. Figures 19A to 19G illustrate exemplary techniques, according to several embodiments, for entering text into a text entry field in response to receiving utterance input. Figures 20A to 20M are flowcharts, according to several embodiments, for a method of entering text into a text entry field in response to receiving utterance input. The user interfaces in Figures 19A to 19G are used to illustrate the processes in Figures 20A to 20M. Figures 21A to 21G illustrate exemplary techniques, according to several embodiments, for modifying text contained in a text entry field. Figures 22A to 22H are flowcharts, according to several embodiments, for a method of modifying text contained in a text entry field. The user interfaces in Figures 21A to 21G are used to illustrate the processes in Figures 22A to 22H. Figures 23A to 23I illustrate exemplary techniques, according to several embodiments, for updating user interface elements according to the status of a hardware input device communicating with a computer system. Figures 24A to 24I are flowcharts illustrating, in several embodiments, a method for updating user interface elements according to the status of a hardware input device communicating with a computer system. The user interfaces in Figures 23A to 23I are used to illustrate the processes in Figures 24A to 24I.

[0047] The processes described below enhance the usability of the device and make the user device interface more efficient (for example, by helping the user provide appropriate input and reducing user errors when operating / interacting with the device) through various technologies, including providing the user with improved visual feedback, reducing the number of inputs required to perform actions, providing additional control options without cluttering the user interface with additional displayed controls, performing actions without requiring further user input when a set of conditions is met, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving memory space, and / or additional technologies. These technologies also reduce power consumption and improve the battery life of the device by enabling the user to use the device more quickly and efficiently. Saving battery power, and therefore weight, improves the ergonomics of the device. These technologies also enable real-time communication, allow the use of fewer and / or less accurate sensors, resulting in more compact, lighter, and cheaper devices, and enabling the device to be used in a variety of lighting conditions. These technologies reduce energy consumption and thereby reduce the heat emitted by the device, which is especially important for wearable devices that can become uncomfortable for the user to wear if they generate excessive heat, even if the device is well within the operating parameters for its components.

[0048] Furthermore, in any method described herein that is conditional on one or more conditions being met in one or more steps, it should be understood that the method described can be repeated in multiple iterations such that all the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires that a first step be performed if a condition is met, and a second step be performed if the condition is not met, a person skilled in the art will understand that the steps described in the claim are repeated in a specific order until the conditions are met and then not met. Thus, a method described in one or more steps that depends on one or more conditions being met can be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this is not required for a claim of a system or computer-readable medium in which the system or computer-readable medium includes instructions that perform a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency has been met without explicitly repeating the steps of the method until all the conditions that the steps of the method are conditional on are met. Those skilled in the art will also understand that, as with a method having conditional steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.

[0049] In some embodiments, as shown in Figure 1, the XR experience is provided to the user via an operating environment 100 which includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, and / or a touchscreen), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, and other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, and / or a velocity sensor), and optionally one or more peripheral devices 195 (e.g., consumer electronics and / or wearable devices). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the display generation component 120 (for example, within a head-mounted device or handheld device).

[0050] When describing an XR experience, various terms are used to refer individually to several related but distinct environments that the user perceives and / or interacts with (for example, using inputs detected by the computer system 101, which causes the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the XR experience). The following is a subset of these terms.

[0051] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.

[0052] Extended reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through an electronic system. In XR, a subset of a person's bodily movements or their representations are tracked, and accordingly, one or more properties of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. For example, an XR system can detect a person's head rotation and, accordingly, adjust the graphic content and sound field presented to the person in a similar way to how such views and sounds would change in a physical environment. Depending on the circumstances (e.g., for reasons of accessibility), adjustments to the properties(s) of virtual objects(s) in the XR environment may be made in response to representations of bodily movements (e.g., voice commands). A person may perceive and / or interact with XR objects using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person can perceive and / or interact with audio objects that create a 3D or spatial audio environment, providing the perception of point audio sources in 3D space. In another example, audio objects may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may perceive and / or interact with only audio objects.

[0053] Examples of XR include virtual reality and mixed reality.

[0054] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of their presence within the computer-generated environment and / or through a simulation of a subset of their physical movement within the computer-generated environment.

[0055] Mixed Reality: A mixed reality (MR) environment is a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects), in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input. On a virtual continuum, a mixed reality environment is any location between, but not encompassing, the complete physical environment at one end and the virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track location and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may take movement into account so that a virtual tree appears stationary relative to the physical ground.

[0056] Examples of mixed reality include extended reality and augmented virtual reality.

[0057] Extended Reality: An Extended Reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system composites the image or video with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through the image or video of the physical environment. As used herein, a video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, so that a person can use the system to perceive the virtual objects superimposed on the physical environment. Extended reality environments also refer to simulated environments in which representations of the physical environment are transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion of it, so that the modified portion is a non-photorealistic altered version of the original captured image.As a further example, the representation of the physical environment may be altered by graphically removing or obscuring parts of it.

[0058] Augmented Virtuality (AV): An augmented virtuality (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, but people with faces might be realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.

[0059] Viewpoint-locked virtual objects: A virtual object is viewpoint-locked when the computer system displays the virtual object at the same location and / or position within the user's view, even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked in the forward direction of the user's head (e.g., the user's viewpoint is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's viewpoint remains fixed even if the user's gaze moves, without moving the user's head. In embodiments where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the extended reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object displayed in the upper-left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper-left corner of the user's viewpoint even if the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position in which a viewpoint-locked virtual object is displayed from the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so that the virtual object is also referred to as a "head-locked virtual object."

[0060] Environment-Locked Virtual Objects: A virtual object is environment-locked (or "world-locked") when a computer system displays it at a location and / or position in the user's viewpoint that is based on (e.g., selected by reference to and / or fixed to) a location and / or object in a three-dimensional environment (e.g., a physical or virtual environment). As the user's viewpoint shifts, the location and / or object in the environment relative to the user's viewpoint changes, and as a result, the environment-locked virtual object will appear at a different location and / or position in the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered in the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes left-leaning in the user's viewpoint (e.g., the tree's position in the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear left-leaning in the user's viewpoint. In other words, the location and / or position in which an environment-locked virtual object is displayed in the user's viewpoint depends on the location and / or object's position and / or orientation in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a fixed location in the physical environment and / or a coordinate system fixed to an object) to determine the position in which the environment-locked virtual object is displayed from the user's viewpoint. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object) or to a moving part of the environment (e.g., a vehicle, animal, person, or a representation of a part of the user's body that moves independently of the user's viewpoint, such as the user's hands, wrists, arms, or feet), so that the virtual object moves as the viewpoint or the part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.

[0061] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed tracking behavior, reducing or delaying the movement of the environment-locked or viewpoint-locked virtual object in response to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed tracking behavior, the computer system detects movement of the reference point that the virtual object is following (e.g., a part of the environment, a viewpoint, or a point fixed to the viewpoint, such as a point between 5 and 300 cm from the viewpoint) and intentionally delays the movement of the virtual object. For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first velocity, the virtual object is moved by the device so as to remain locked to the reference point, but at a second velocity slower than the first velocity (e.g., the virtual object begins to catch up to the reference point until the reference point stops or slows down). In some embodiments, when a virtual object exhibits delayed tracking behavior, the device ignores small amounts of movement of the reference point (e.g., ignoring movement of the reference point that is below a threshold amount of movement, such as movement between 0 and 5 degrees or movement between 0 and 50 cm). For example, when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a "delayed tracking" threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object that maintains a substantially fixed position with respect to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / behind the position of the reference point).

[0062] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing sounds of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface. In some embodiments, the controller 110 is configured to manage and adjust the XR experience for the user.In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server or a central server). In some embodiments, the controller 110 is communicably coupled to display generation components 120 (e.g., HMDs, displays, projectors, and / or touchscreens) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, and / or IEEE 802.3x). In another example, the controller 110 is contained within a housing (e.g., a physical housing) of one or more of the display generation components 120 (e.g., an HMD, or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.

[0063] In some embodiments, the display generation component 120 is configured to provide the user with an XR experience (e.g., at least the visual components of the XR experience). In some embodiments, the display generation component 120 includes a preferred combination of software, firmware, and / or hardware. The display generation component 120 is described in more detail below with reference to Figure 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.

[0064] According to some embodiments, the display generation component 120 provides the user with an XR experience while the user is virtually and / or physically present in the scene 105.

[0065] In some embodiments, the display generation component is mounted on a part of the user's body (e.g., the user's head or the user's hand). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing mounted on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content when the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with XR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the XR content response is displayed through the HMD. Similarly, a user interface showing interaction with XR content triggered based on the movement of a handheld or tripod-mounted device relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may be implemented similarly to an HMD where the movement is triggered by the movement of the HMD relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).

[0066] While relevant features of the operating environment 100 are shown in Figure 1, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the exemplary embodiments disclosed herein.

[0067] Figure 2 is a block diagram of an example of the controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0068] In some embodiments, one or more communication buses 204 include circuits for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0069] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and XR experience module 240.

[0070] The operating system 230 handles various basic system services and includes instructions for performing hardware-dependent tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for each group of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0071] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, and / or location data) from at least the display generation component 120 of Figure 1, and optionally from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0072] In some embodiments, the tracking unit 242 is configured to map scene 105 and track the position / location of at least the display generation component 120 relative to scene 105 in Figure 1, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes its instructions and / or logic, as well as heuristics and metadata therefor. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand, relative to scene 105 in Figure 1, relative to the display generation component 120, and / or relative to a coordinate system defined for the user's hand. The hand tracking unit 244 is described in more detail below with respect to Figure 4. In some embodiments, the eye-tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to XR content displayed via the display generation component 120. The eye-tracking unit 243 is described in more detail below with reference to Figure 5.

[0073] In some embodiments, the adjustment unit 246 is configured to manage and adjust the XR experience presented to the user by the display generation component 120 and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0074] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, and / or location data) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 248 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0075] While the data acquisition unit 241, tracking unit 242 (including, for example, eye-tracking unit 243 and hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (including, for example, eye-tracking unit 243 and hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.

[0076] Furthermore, Figure 2 is intended to illustrate the function of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules separately shown in Figure 2 can be realized in a single module, and the various functions of a single functional block can be realized by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0077] Figure 3 is a block diagram of an example of a display generation component 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting examples, the display generation component 120 (e.g., HMD) may include one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional in-facing and / or out-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0078] In some embodiments, one or more communication buses 304 include circuits for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, and / or a blood glucose sensor), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).

[0079] In some embodiments, one or more XR displays 312 are configured to provide the user with an XR experience. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal displays (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistors (OLET), organic light-emitting diodes (OLED), surface conduction electron emission displays (SED), field emission displays (FED), quantum dot light-emitting diodes (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffraction, reflection, polarization, holographic, and / or waveguide displays. For example, a display generation component 120 (e.g., HMD) includes a single XR display. In another embodiment, the display generation component 120 includes an XR display for each of the user's eyes. In some embodiments, one or more XR displays 312 can present MR or VR content.

[0080] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally, at least a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would view if a display generation component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.

[0081] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and XR presentation module 340.

[0082] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to the user via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.

[0083] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, and / or location data) from at least the controller 110 in Figure 1. To this end, in various embodiments, the data acquisition unit 342 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0084] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation unit 344 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0085] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (for example, a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. For this purpose, in various embodiments, the XR map generation unit 346 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0086] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, and / or location data) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 348 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0087] While the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 are shown as existing on a single device (e.g., the display generation component 120 in Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.

[0088] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular implementation, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be realized within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0089] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1) is controlled by the hand tracking unit 244 (Figure 2) to track the location / position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to the scene 105 of Figure 1 (e.g., relative to a part of the physical environment surrounding the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to the user's hand) in a defined coordinate system. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0090] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.

[0091] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), which drives the display generation components 120 accordingly. For example, a user can interact with the software running on the controller 110 by moving their hand 406 to change the orientation of their hand.

[0092] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines a set of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to a z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.

[0093] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the processor in the image sensor 404 and / or controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on a previous learning process, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D location of the user's wrist and fingertips.

[0094] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The pose estimation function described herein may be interleaved with the motion tracking function, so that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover changes in pose that occur over the remaining frames. Pose, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify the image presented on the display generation component 120, or perform other functions, depending on the pose and / or gesture information.

[0095] In some embodiments, the gestures include air gestures. Air gestures are gestures detected without (or independently of) the user touching an input element that is part of a device (e.g., a computer system 101, one or more input devices 125, and / or a hand tracking device 140), and are based on detected movements of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including the user's body movement relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), the user's body movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the user's other hand, and / or the movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movements of a part of the user's body (e.g., a tap gesture including the movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body).

[0096] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures, as in some embodiments, performed by moving one or more of the user's fingers relative to other fingers or parts of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device), and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving rotation of a part of the user's body by a predetermined speed or amount).

[0097] In some embodiments where the input gesture is an air gesture (i.e., without physical contact with an input device that provides the computer system with information about which user interface element is the target of user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor over a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of user input (e.g., in the case of direct input, as described below). Thus, in implementations involving air gestures, the input gesture is the detected attention (e.g., gaze) to the user interface element in combination (e.g., simultaneously) with the movement of the user's fingers (one or more) and / or hand to perform pinch and / or tap input, as described in more detail below.

[0098] In some embodiments, input gestures directed towards a user interface object are performed directly or indirectly by reference to the user interface object. For example, user input is performed directly towards the user interface object in response to the user performing an input gesture with their hand at a position corresponding to the user interface object's position in a three-dimensional environment (e.g., determined based on the user's current viewpoint). In some embodiments, the input gesture is performed indirectly towards the user interface object according to the user performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in a three-dimensional environment, while detecting the user's attention (e.g., gaze) to the user interface object. For example, in the case of a direct input gesture, the user can direct their input towards the user interface object by initiating the gesture at or near a position corresponding to the user interface object's display position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm from the optional outer edge or optional central portion). In the case of indirect input gestures, the user can direct their input towards the user interface object by paying attention to the user interface object (for example, by gazing at the user interface object), and while paying attention to the options, the user initiates the input gesture (for example, at any position detectable by the computer system) (for example, at a position that does not correspond to the display position of the user interface object).

[0099] In some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch and tap inputs for interacting with virtual or mixed reality environments, as in some embodiments. For example, the pinch and tap inputs described later are performed as air gestures.

[0100] In some embodiments, a pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture involves moving two or more fingers of one hand to touch each other, i.e., including an optional interruption (e.g., within 0 to 1 second) immediately after the fingers touch each other. A long pinch gesture that is an air gesture involves moving two or more fingers of one hand to touch each other for at least a threshold time (e.g., at least 1 second) before detecting an interruption of contact between the fingers. For example, a long pinch gesture involves the user holding a pinch gesture (e.g., if two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture, which is an air gesture, includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected directly from each other (e.g., within a predetermined period of time) in succession. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks contact between two or more fingers), and then performs a second pinch input within a predetermined period of time (e.g., within one or two seconds) after releasing the first pinch input.

[0101] In some embodiments, an air gesture, a pinch-and-drag gesture, includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in relation to (e.g., after) a drag input that changes the user's hand position from a first position (e.g., a drag initiation position) to a second position (e.g., a resistance termination position). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers) to terminate the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together and touches them to each other, and then moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by the user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of the user's hands. For example, an input gesture includes two (e.g., or more) pinch inputs performed in relation to each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using the user's first hand, and a second pinch input performed using the other hand (e.g., a second hand of the user's hands) in relation to performing a pinch input using the first hand. In some embodiments, movement between the user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).

[0102] In some embodiments, a tap input performed as an air gesture (e.g., directed towards a user interface element) includes the movement of one or more of the user's fingers toward the user interface element, the movement of the user's hand toward the user interface element with the user's fingers (one or more) optionally extended toward the user interface element, a downward movement of the user's fingers (e.g., mimicking a mouse click or a tap on a touchscreen), or other default movements of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand that performs the tap gesture movement away from the user's viewpoint and / or toward the object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand that performs the tap gesture (e.g., away from the user's viewpoint and / or the end of the movement toward the object that is the target of the tap input, a reversal of the direction of the finger or hand movement, and / or a reversal of the direction of acceleration of the finger or hand movement).

[0103] In some embodiments, the user's attention is determined to be directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards that part of the three-dimensional environment (optionally, without requiring any other conditions). In some embodiments, for the device to determine that the user's attention is directed towards a part of the three-dimensional environment, the device determines that the user's attention is directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards a part of the three-dimensional environment, with one or more additional conditions such as the gaze being directed towards the part of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the part of the three-dimensional environment, and / or the gaze being directed towards a part of the three-dimensional environment. If one of the additional conditions is not met, the device determines that the user's attention is not directed towards the part of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).

[0104] In some embodiments, the detection of a ready state configuration of the user or a part of the user is detected by the computer system. The detection of a ready state configuration of the hand is used by the computer system as an indication that the user is likely to be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap shape where one or more fingers are extended and the palm is facing away from the user), whether the hand is in a predetermined position relative to the user's line of sight (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular way (e.g., moved towards the area in front of the user above the user's waist, below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of the user interface is responsive to attention (e.g., gaze) input.

[0105] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although the controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or in other ways. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.

[0106] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from image sensor 404, with the gradation becoming richer as the depth increases. Controller 110 processes these depth values ​​to identify and segment image components (i.e., adjacent pixel groups) that have the characteristics of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame movement of the depth map sequence.

[0107] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the hand skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., knuckles, fingertips, palm center, and / or hand end connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.

[0108] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 243 (Figure 2) to track the position and movement of the user's gaze toward the scene 105 or toward the XR content displayed via the display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating XR content for user viewing and a component for tracking the user's gaze toward the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or XR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with a display generation component that is mounted on the head or a display generation component that is not mounted on the head. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally part of a non-head-mounted display generation component.

[0109] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames containing left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that the individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0110] As shown in Figure 5, in some embodiments, the eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) camera or a near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visible light to pass through. The eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60–120 frames per second (fps)), analyzes the images to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a separate eye-tracking camera and light source.

[0111] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR device to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. The user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil location, foveal location, optical axis, visual axis, and / or interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters of the eye-tracking device 130 are determined, images captured by the eye-tracking camera may be processed using a glint-assisted method to determine the user's current visual axis and gaze point relative to the display.

[0112] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and an eye-tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes(s) 592. The eye-tracking camera 540 is positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, and / or a projector) and may be directed towards a mirror 550 that transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of Figure 5).

[0113] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses gaze tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's viewpoint on the display 510 based on the gaze tracking input 542 obtained from the eye-tracking camera 540. The viewpoint estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.

[0114] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the XR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can use the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.

[0115] In some embodiments, the eye-tracking device is part of a head-mounted device mounted on a wearable housing, which includes a display (e.g., display 510), two eyepieces (e.g., one or more eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., a light source 530 (e.g., an IR LED or NIR LED)). The light source emits light (e.g., IR light or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in Figure 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.

[0116] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0117] Embodiments of eye-tracking systems, such as those shown in Figure 5, can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.

[0118] Figure 6 shows glint-assisted eye-tracking pipelines according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1 and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.

[0119] As shown in Figure 6, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which begins at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0120] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.

[0121] At 640, if the process proceeds from element 610, the current frame is analyzed to track the pupil and glint based in part on previous information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.

[0122] Figure 6 is intended to serve as an example of an eye-tracking technology that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking technologies that currently exist or may be developed in the future may be used in computer system 101 to provide users with XR experiences in various embodiments, either in place of or in combination with the Glint-assisted eye-tracking technology described herein.

[0123] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, for example, a mixed reality environment in which one or more virtual objects are superimposed on a representation of the real-world environment 602.

[0124] Accordingly, this description describes several embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and virtual objects. For example, a three-dimensional environment optionally includes a representation of a table existing in a physical environment that is captured and displayed within the three-dimensional environment (e.g., actively via a computer system's camera and display, or passively via a computer system's transparent or translucent display). As described above, a three-dimensional environment optionally is a mixed reality system based on a physical environment in which the three-dimensional environment is captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system may optionally selectively display parts and / or objects of the physical environment so that each part and / or object of the physical environment appears to exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system may optionally display virtual objects in a three-dimensional environment so that the virtual objects appear to exist in the real world (e.g., a physical environment) by placing virtual objects in each location within the three-dimensional environment that have corresponding locations in the real world. For example, a computer system may optionally display a vase in such a way that it appears as if a real vase were placed on a table in a physical environment. In some embodiments, individual locations in a three-dimensional environment have corresponding locations in the physical environment.Therefore, when a computer system is described as displaying virtual objects in separate locations relative to physical objects (for example, at or near the location of the user's hand, or on or near a physical table), the computer system displays the virtual objects in specific locations within a three-dimensional environment so that they appear to be at or near physical objects in the physical world (for example, if the virtual object is a real object at that specific location, the virtual object will be displayed in the location within the three-dimensional environment that corresponds to the location within the physical environment where the virtual object would have been displayed).

[0125] In some embodiments, real-world objects existing in a physical environment displayed within a three-dimensional environment (e.g., real-world objects visible via and / or display-generating components) can interact with virtual objects existing only within the three-dimensional environment. For example, the three-dimensional environment may include a table and a vase placed on the table, where the table is a view (or representation) of a physical table in the physical environment, and the vase is a virtual object.

[0126] Similarly, just as virtual objects are real objects in a physical environment, the user can optionally interact with virtual objects in a three-dimensional environment using one or more hands. For example, as described above, one or more sensors in the computer system can optionally capture one or more of the user's hands and display a representation of the user's hands in a three-dimensional environment (in a similar manner to, for example, displaying real-world objects in a three-dimensional environment as described above), or, in some embodiments, the user's hands are visible through the display-generating components by the ability to see the physical environment through the user interface, due to the transparency / transparency of some of the display-generating components displaying the user interface, or the projection of the user interface onto a transparent / translucent surface, or the projection of the user interface onto the user's eyes or the user's field of view. Thus, in some embodiments, the user's hands are displayed at separate locations in the three-dimensional environment and are treated as if they were objects in a three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, the computer system can update the display of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.

[0127] In some of the embodiments described below, the computer system can optionally determine the "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment for the purpose of determining, for example, whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, and / or holding the virtual object, or whether it is within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the fingers of a hand pressing a virtual button, a user's hand grasping a virtual vase, two fingers of a user's hand pinching / holding an application's user interface together, and other types of interactions described herein. For example, when determining whether a user is interacting with a virtual object and / or how a user is interacting with a virtual object, the computer system can optionally determine the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the virtual object of interest in the three-dimensional environment. For example, one or more of the user's hands are located in a specific position in the physical world, which the computer system optionally captures and displays at a specific corresponding position in a three-dimensional environment (e.g., the position in the three-dimensional environment where the hands are displayed, if the hands are virtual hands rather than physical hands). The position of the hands in the three-dimensional environment is optionally compared to the position of a target virtual object in the three-dimensional environment to determine the distance between the one or more of the user's hands and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing the position in the physical world (as opposed to comparing the position in the three-dimensional environment).For example, when determining the distance between one or more of the user's hands and a virtual object, the computer system optionally determines the corresponding location of the virtual object in the physical world (e.g., the position in the physical world where the virtual object would be located if it were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and one or more of the user's hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Thus, when determining whether a physical object is in contact with a virtual object, or whether a physical object is within a threshold distance of a virtual object, as described herein, the computer system optionally performs one of the techniques described above to map the location of the physical object to a three-dimensional environment and / or to map the location of the virtual object to a physical environment.

[0128] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed, and / or where and what the physical stylus held by the user is directed. For example, if the user's gaze is directed to a particular position in the physical environment, the computer system optionally determines the corresponding position in the three-dimensional environment (e.g., the virtual position of the gaze), and if a virtual object is located at that corresponding virtual position, the computer system optionally determines that the user's gaze is directed to that virtual object. Similarly, the computer system optionally determines, based on the orientation of the physical stylus, where in the physical environment the stylus is pointing. In some embodiments, based on this determination, the computer system optionally determines the corresponding virtual position in the three-dimensional environment corresponding to the location in the physical environment that the stylus is pointing to, and optionally determines that the stylus is pointing to the corresponding virtual position in the three-dimensional environment.

[0129] Similarly, embodiments described herein may refer to the location of a user (e.g., a user of a computer system) and / or the location of a computer system in a three-dimensional environment. In some embodiments, the user of a computer system is holding, wearing, or otherwise positioned near the computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the user's location. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to individual locations in the three-dimensional environment. For example, if a user is standing in a location facing an individual part of the physical environment that is visible by a display-generating component, the location of the computer system is the location in the physical environment (and its corresponding location in the three-dimensional environment) where the user can see objects in the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) or size as the objects are displayed by the display-generating component of the computer system in the three-dimensional environment. Similarly, if a virtual object displayed in a three-dimensional environment is a physical object in a physical environment (for example, the virtual object is located in the same physical environment location as it is in the three-dimensional environment, and has the same size and orientation as it does in the three-dimensional environment), then the computer system and / or user's location is the position from which the user will view the virtual object in the physical environment in the same position, orientation, and / or size (for example, absolutely, and / or relative to each other, and in relation to real-world objects) as it was displayed by the computer system's display generation components in the three-dimensional environment.

[0130] This disclosure describes various input methods for interaction with computer systems. Where one example is provided using one input device or method, and another example is provided using a different input device or method, each example may be compatible with the input device or method described in the other example, and their use should be considered optional. Similarly, various output methods for interaction with computer systems are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, each example may be compatible with the output device or method described in the other example, and their use should be considered optional. Similarly, various methods for interaction with virtual or mixed reality environments via computer systems are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, each example may be compatible with the method described in the other example, and their use should be considered optional. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User interface and related processes

[0131] Here, we focus on embodiments of user interfaces ("UI"), as well as related processes that can be implemented on a computer system such as a portable multifunction device or head-mounted device communicating with display generation components and one or more input devices.

[0132] Figures 7A to 7H illustrate exemplary techniques for scrolling scrollable content in response to various user inputs, according to several embodiments. The user interfaces in Figures 7A to 7H are used to illustrate the processes described below, including the processes in Figures 8A to 8L.

[0133] Figure 7A shows a computer system 101 that displays a three-dimensional environment 701 from the user's viewpoint via a display generation component (e.g., display generation component 120 in Figure 1). As described above with reference to Figures 1 to 6, the computer system 101 optionally includes a display generation component (e.g., a touchscreen) and a plurality of image sensors 314 (e.g., the image sensor 314 in Figure 3). The image sensors 314 optionally include one or more of the following: a visible light camera, an infrared camera, a depth sensor, or any other sensors that the computer system 101 can use to capture one or more images of the user or a part of the user (e.g., one or more of the user's hands) while the user is interacting with the computer system 101. In some embodiments, the user interfaces illustrated and described below may also be implemented on a head-mounted display, including a display generating component that displays the user interface or three-dimensional environment to the user, and sensors (e.g., external sensors facing outward from the user) and / or the user's hand movements, such as movement which is interpreted as a gesture by a computer system, such as an air gesture. In some embodiments, it should be understood that one or more of the techniques described herein may be applied to a two-dimensional environment without departing from the scope of this disclosure.

[0134] In Figure 7A, the computer system 101 presents scrollable content 702 via the display generation component 120. In some embodiments, the scrollable content 702 includes text content 707 and additional content 705. For example, the scrollable content 702 is an article, the text content 707 is the text of the article, and the additional content 705 is an embedded advertisement and / or one or more links to related articles. In some embodiments, the scrollable content includes a first scrollable area 704 and a second scrollable area 706. As will be described in more detail below, the computer system 101 scrolls the scrollable content 702 in response to detecting the user's gaze directed towards the first scrollable area 704 or the second scrollable area 706 without detecting the user's hand ready state. In some embodiments, detecting the user's hand ready state includes detecting a ready state associated with an air gesture, as will be described in more detail above. In some embodiments, upon detecting that the user's gaze is directed towards the area of ​​scrollable content 702 between scrollable areas 704 and 706, the computer system maintains the display of the scrollable content 702 without scrolling the scrollable content.

[0135] As shown in Figure 7A, in some embodiments, scroll regions 704 and 706 are close to the boundaries of the scrollable content 702. For example, the scrollable content 702 is vertically scrollable, and therefore the first scroll region 704 is at the top of the scrollable content 702, and the second scroll region 706 is at the bottom of the scrollable content 702. As shown in Figure 7A, the first scroll region 704 at the top of the scrollable content 702 is smaller than the second scroll region 706 at the bottom of the scrollable content 702. In some embodiments, if the scrollable content 702 is horizontally scrollable, the scrollable content 702 will include a left scroll region and a right scroll region (instead of, or in addition to, the upper scroll region such as the first scroll region 704 and the lower scroll region such as the second scroll region 706).

[0136] As shown in Figure 7A, the computer system 101 detects the user's gaze 713a directed towards the second scrollable area 706. In some embodiments, in response to detecting the user's gaze 713a directed towards the second scrollable area 706, the computer system 101 scrolls the scrollable content 702 downwards, as shown in Figure 7B.

[0137] Figure 7B illustrates how the computer system 101 scrolls the scrollable content 702 in response to detecting the user's gaze 713a directed towards the second scrollable area 706 in Figure 7A. As shown in Figure 7B, in response to detecting the user's gaze 713a in Figure 7A directed towards the second scrollable area 706 below the scrollable content 702, the computer system 101 scrolls the scrollable content 702 downwards (for example, by moving the scrollable content 702 upwards to reveal additional scrollable content 702 below the scrollable content 702). In some embodiments, if the user's gaze was directed towards the first scrollable area 704 above the scrollable content 702, the computer system 101 scrolls the scrollable content 702 upwards (for example, by moving the scrollable content 702 downwards to reveal additional scrollable content 702 above the scrollable content 702).

[0138] In some embodiments, the scrolling acceleration and / or speed differ when scrolling upwards (e.g., in response to detection of the user's gaze directed towards the first scrolling area 704) and when scrolling downwards (e.g., in response to detection of the user's gaze directed towards the second scrolling area 706). In some embodiments, the scrolling acceleration and / or speed is faster when scrolling upwards (e.g., in response to detection of the user's gaze directed towards the first scrolling area 704) than when scrolling downwards (e.g., in response to detection of the user's gaze directed towards the second scrolling area 706). In some embodiments, the scrolling acceleration and / or speed is slower when scrolling upwards (e.g., in response to detection of the user's gaze directed towards the first scrolling area 704) than when scrolling downwards (e.g., in response to detection of the user's gaze directed towards the second scrolling area 706).

[0139] In some embodiments, the computer system 101 detects a transition in the user's gaze 713a from a state where it is not directed towards one of the scroll regions 704 or 706 to a state where it is directed towards one of the scroll regions 704 or 706, and in response, gradually increases the scroll speed of the scrollable content 702 from a state where it is not scrolling to a state where it scrolls at an individual scroll speed. As described above, the individual scroll speed is based on which of the two scroll regions 704 or 706 the user's gaze is directed towards. In some embodiments, the individual scroll speed is based on the distance from the edge of the scrollable content 702 in the scroll region 704 or 706 where the user's gaze is detected. For example, in response to detecting the user's gaze 713a at the position shown in Figure 7A in the second scroll region 706, the computer system 101 scrolls the scrollable content 702 at a first speed. In Figure 7B, the computer system 101 detects the user's gaze 713b directed to a different location within the second scroll region 706, closer to the edge (e.g., the bottom edge) of the scrollable content 702, compared to the location of the user's gaze 713a shown in Figure 7A. In some embodiments, in response to detecting the user's gaze 713b at the position within the second scroll region 706 shown in Figure 7B, the computer system 101 scrolls the scrollable content 702 at a faster speed than the scrolling speed corresponding to the gaze 713a within the second scroll region 706 as shown in Figure 7A.

[0140] Figure 7C shows a computer system 101 scrolling scrollable content 702 in response to the user's line of sight 713b directed to a position within the second scrollable area 706 shown in Figure 7B. Because the user's line of sight 713b in Figure 7B is closer to the boundary (e.g., the lower edge) of the scrollable content 702 than the location of the user's line of sight 713a in Figure 7A, the amount of scrolling shown in Figure 7C is greater than the amount of scrolling shown in Figure 7B.

[0141] In some embodiments, the computer system 101 stops scrolling the scrollable content 702 in response to detecting the user's gaze directed to a portion of the scrollable content 702 outside of the scrollable area 704 or 706, or in response to detecting that the user's hand is in a ready state while the user's gaze is directed to one of the scrollable areas 704 or 706. For example, Figure 7C shows the user's gaze 713d directed to a portion of the scrollable content 702 that is not included in the first scrollable area 704 or the second scrollable area 706. Figure 7C also shows the user's hand 703a in a ready state (e.g., "hand state A") while the user's gaze 713c is directed to the second scrollable area 706 of the scrollable content 702. In response to detecting the user's gaze 713d shown in Figure 7C, or the ready state of the user's gaze 713c and hand 703a shown in Figure 7C, the computer system 101 stops scrolling the scrollable content as shown in Figure 7D.

[0142] Figure 7D shows a computer system 101 that maintains the display of scrollable content 702 without scrolling it, in response to one of the inputs described above with respect to Figure 7C. In some embodiments, when scrolling of scrollable content 702 is to be stopped, the computer system 101 gradually slows down the scrolling of scrollable content 702 until the scrolling is stopped.

[0143] Figure 7D also shows a computer system 101 that detects input for scrolling scrollable content 702 provided by the user's hand 703b. In some embodiments, the input for scrolling scrollable content 702 includes detecting the user's line of sight 713e directed at the scrollable content 702 and hand movement (e.g., air gesture, touch input, or other hand input) 703b while the hand 703b is in a pinch-hand shape with the thumb touching or touching another finger of the hand 703b within a threshold distance (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, or 1 centimeter) ("Hand State C"). For example, in Figure 7D, the computer system 101 detects that the user's hand 703b is moving upward while in a pinch-hand shape, with the user's gaze 713e directed towards the scrollable content 702, and accordingly scrolls the scrollable content 702 downward (for example, by moving the scrollable content 702 upward to reveal additional scrollable content 702 below the scrollable content 702), as shown in Figure 7E. Although Figure 7D shows the user's gaze 713e directed towards a portion of the scrollable content 702 that is not within the scrollable area 704 or 706, in some embodiments the computer system scrolls the scrollable content 702 in response to inputs including the movement of the hand 703b and the user's gaze directed towards one of the scrollable areas 704 or 706 of the scrollable content 702.

[0144] Figure 7E illustrates how the computer system 101 updates the display of scrollable content 702 by scrolling the scrollable content 702 in response to the input shown in Figure 7D, as described above. In Figure 7E, the computer system 101 detects input to scroll the scrollable content 702 upward, provided by the user's hand 703c while the user's gaze 713f is directed towards the scrollable content 702. As shown in Figure 7E, the computer system 101 detects that the hand 703c is moving downward while in a pinch-hand shape (e.g., “hand state C”) while the user's gaze 713f is directed towards the scrollable content 702. In response to the scroll input shown in Figure 7E, the computer system 101 scrolls the scrollable content 702 upward (e.g., by moving the scrollable content 702 downward to show additional scrollable content 702 above the scrollable content 702), as shown in Figure 7F. Figure 7E shows the user's line of sight 713f directed to a portion of the scrollable content 702 that is not within the scrollable area 704 or 706, but in some embodiments, the computer system scrolls the scrollable content 702 in response to inputs including the movement of the hand 703c and the user's line of sight directed to one of the scrollable areas 704 or 706 of the scrollable content 702.

[0145] Figure 7F illustrates how the computer system 101 updates the display of scrollable content 702 by scrolling it in response to the input shown in Figure 7E, as described above. In some embodiments, the computer system 101 scrolls the scrollable content 702 down by a greater amount in response to a scroll input provided by the user's hand than the amount by which the computer system 101 scrolls the scrollable content 702 up in response to the scroll input provided by the user's hand for the same amount of hand movement in the opposite direction (e.g., air gesture, touch input, or other hand input). For example, the amount of hand movement (e.g., air gesture, touch input, or other hand input) 703b shown in Figure 7D is the same as the amount of hand movement (e.g., air gesture, touch input, or other hand input) 703c shown in Figure 7E, but the amount of scroll of the scrollable content 702 in Figure 7E in response to the input in Figure 7D is greater than the amount of scroll of the scrollable content 702 in Figure 7F in response to the input in Figure 7E. In some embodiments, the “amount” of hand movement (e.g., air gesture, touch input, or other hand input) includes the amount of distance, duration, and / or velocity of hand movement (e.g., air gesture, touch input, or other hand input) while in a pinch position, and the user’s gaze is directed towards the scrollable content 702 to provide scroll input directed towards the scrollable content 702.

[0146] In some embodiments, the computer system 101 increases the scrolling speed as the hand moves further from the location where the pinch-hand gesture was initiated, in response to input for scrolling the scrollable content 702 provided by the user's hand, such as the input shown in Figure 7D or Figure 7E. For example, upon detecting a first movement of the hand (e.g., an air gesture, touch input, or other hand input) from the hand's location when the pinch-hand gesture was initiated, the computer system 101 scrolls the scrollable content 702 at a first speed and optionally continues scrolling at the first speed while the hand remains at an updated location after the first movement. In this example, upon detecting a second movement (e.g., an air gesture, touch input, or other hand input) greater than the first movement of the hand from the hand's location when the pinch-hand gesture was initiated, the computer system 101 scrolls the scrollable content 702 at a second speed greater than the first speed and optionally continues scrolling at the second speed while the hand remains at an updated location after the second movement.

[0147] In some embodiments, the computer system 101 scrolls the scrollable content 702 in response to detecting hand movement in a pinch-hand position (e.g., air gesture, touch input, or other hand input) while the user's gaze is directed towards the scrollable content 702, based on a determination that the hand movement (e.g., air gesture, touch input, or other hand input) while the hand is in a pinch-hand position meets one or more criteria. In some embodiments, if the amount of movement (e.g., speed, distance, and / or duration of movement) is less than a predetermined threshold amount, the computer system 101 maintains the display of the scrollable content 702 without scrolling it. Exemplary thresholds are provided below with reference to Method 800 and Figures 8A-8L. In some embodiments, if the hand movement in a pinch position (e.g., air gesture, touch input, or other hand input) is downward and exceeds a threshold speed, the computer system 101 maintains the display of the scrollable content 702 without scrolling it. Exemplary threshold velocities are provided below with reference to Method 800 and Figures 8A to 8L.

[0148] In some embodiments, the computer system 101, while detecting a pinch gesture performed by the user, selects one or more selectable user interface elements displayed via the display generation component 120 in response to detecting the user's gaze directed towards selectable user interface elements. In some embodiments, the one or more selectable user interface elements are selectable options, representations of content items, application icons, user interface containers (e.g., windows), hyperlinks, etc. Exemplary actions performed in response to the selection of these elements include navigating the user interface, presenting content items, saving or opening a file or document, initiating communication with another computer system, changing the settings of the computer system, updating the current input focus, etc.

[0149] Figure 7G shows computer system 101 presenting scrollable content text content 707 in reader mode without displaying additional content 705 of scrollable content 702. The examples shown in Figures 7A to 7F above are examples of computer system 101 presenting scrollable content 702, including text content 707 and additional content 705, in browsing mode. In some embodiments, computer system 101 transitions between displaying content in reader mode and displaying content in browsing mode in response to one or more user inputs.

[0150] In some embodiments, while the computer system 101 is displaying the scrollable text content 707 in reader mode, as shown in Figure 7G, the computer system 101 is configured to scroll the text content 707 in accordance with the user's gaze being directed to a first scrollable area 708 or a second scrollable area 710, in a manner similar to that described above with reference to Figures 7A to 7D with respect to browsing mode. In some embodiments, the computer system 101 is also configured to scroll the text content 707 line by line in response to detecting that the user is reading the text content 707. In some situations, when people read text, once they have finished reading a line of text, they direct their gaze to the beginning of the next line by moving their gaze along the line they just read, from the end of the line they just read to the beginning of the line they just read, before looking at the next line. In Figure 7G, the computer system 101 detects the user's gaze 713h moving from the end of a line of text content 707 to the beginning of a line. In response to detecting the movement of the gaze 713h shown in Figure 7G, the computer system 101 scrolls the text content 707 (for example, by one line) as shown in Figure 7H. In some embodiments, the computer system 101 scrolls the text content 707 in response to the movement of the gaze 713h shown in Figure 7G, regardless of whether the user's hands are detected in a ready state or not.

[0151] Figure 7H shows a computer system 101 that displays the text content 707 after scrolling it in accordance with the movement of the user's gaze 713h shown in Figure 7G. As shown in Figure 7H, in some embodiments, the computer system 101 scrolls the text content 707 by one line in response to the movement of the gaze 713h shown in Figure 7G.

[0152] In some embodiments, the computer system 101 displays a word definition 712 in response to detecting the user's gaze directed at a word for at least a predetermined threshold time. Exemplary time thresholds are provided below with reference to Method 800 and Figures 8A-8L. For example, in Figure 7H, the computer system 101 detects the user's gaze directed at a word 713i for a time threshold and displays a word definition 712 overlaid on the text content 707 accordingly. In some embodiments, the computer system 101 similarly displays a word definition while displaying scrollable content 702, including text content 707 and additional content 705, in the browsing mode shown in Figures 7A-7F. Further explanation of Figures 7A-7H is provided below with reference to Method 800, which is described with respect to Figures 7A-7H.

[0153] Figures 8A to 8L are flowcharts of methods for scrolling scrollable content in response to various user inputs, according to various embodiments. In some embodiments, Method 800 is performed on a computer system (e.g., computer system 101 in Figure 1) that includes display generation components (e.g., display generation components 120 in Figures 1, 3, and 4). In some embodiments, Method 800 is stored on a non-temporary (or temporary) computer-readable storage medium and controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 in Figure 1A). Some operations of Method 800 are optionally combined, and / or the order of some operations is optionally changed.

[0154] In some embodiments, Method 800 is performed in a computer system (e.g., 101) that communicates with a display generating component and one or more input devices (e.g., 314), such as Figure 7A (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device) or a computer). In some embodiments, the display generating component is a display integrated with the computer system (optionally a touchscreen display), an external display such as a monitor, projector, or television, or a hardware component (optionally integrated or external) for projecting a user interface or making a user interface visible to one or more users. In some embodiments, one or more input devices include a computer system or component capable of receiving user input (e.g., capturing user input and / or detecting user input) and transmitting information associated with the user input to the computer system. Examples of input devices include touchscreens, mice (e.g., external), trackpads (optionally integrated or external), touchpads (optionally integrated or external), remote control devices (e.g., external), another mobile device (e.g., separate from the computer system), handheld devices (e.g., external), controllers (e.g., external), cameras, depth sensors, eye-tracking devices, and / or motion sensors (e.g., hand-tracking devices and / or hand motion sensors). In some embodiments, the computer system communicates with the hand-tracking device (e.g., one or more cameras, depth sensors, proximity sensors, and / or touch sensors (e.g., touchscreen or trackpad)). In some embodiments, the hand-tracking device is a wearable device such as a smart glove. In some embodiments, the hand-tracking device is a handheld input device such as a remote control or stylus.

[0155] In some embodiments, such as Figure 7A, a computer system (e.g., 101) displays a user interface (e.g., 702) containing scrollable content (e.g., 705 or 707) via a display generation component (802a). In some embodiments, the scrollable content includes text and / or images. In some embodiments, the scrollable content exceeds the size of the scrollable user interface element in which the scrollable content is displayed. In some embodiments, in response to a request to scroll the scrollable content, the computer system optionally discontinues displaying a first portion of the scrollable content and begins displaying a second portion of the content while maintaining display of a third portion of the content within the scrollable user interface element. In some embodiments, the scrollable content is displayed in a three-dimensional environment. In some embodiments, the three-dimensional environment includes virtual objects such as application windows, operating system elements, representations of other users, and / or representations of content items and physical objects in the physical environment of the computer system. In some embodiments, representations of physical objects are displayed in the three-dimensional environment via a display generation component (e.g., virtual passthrough or video passthrough). In some embodiments, the representation of a physical object is a view of the physical object in the physical environment of the computer system that is visible through the transparency of the display generation components (e.g., true passthrough or reality passthrough). In some embodiments, the computer system displays a three-dimensional environment from the user's viewpoint at a location in the three-dimensional environment corresponding to the physical location of the computer system in the physical environment of the computer system. In some embodiments, the three-dimensional environment is generated, displayed, or otherwise made visible by a device (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment).

[0156] In some embodiments, such as Figure 7A, a computer system (e.g., 101) detects the user's gaze (e.g., 713a) directed towards scrollable content (e.g., 705 or 707) via one or more input devices (e.g., eye-tracking device 314) (802b).

[0157] In some embodiments, such as Figure 7C, upon detecting a user's gaze directed towards scrollable content (e.g., 713d) (802c), and in accordance with the determination that the user's gaze (e.g., 713d) is directed towards a first area of ​​the scrollable content (e.g., 707), the computer system (e.g., 101) maintains the display of the scrollable content (e.g., 707) without scrolling it (e.g., 802d). In some embodiments, the first area of ​​the scrollable content is separated from one or more directions in which the scrollable content is scrollable. For example, if the scrollable content is vertically scrollable, the first area of ​​the scrollable content is the area of ​​the scrollable content between the top and bottom of the scrollable content. As another example, if the scrollable content is horizontally scrollable, the first area of ​​the scrollable content is the area of ​​the scrollable content between the left and right portions of the scrollable content. In some embodiments, the computer system detects the user's gaze directed towards scrollable content, but does not detect any additional input (e.g., via one or more input devices other than an eye-tracking device) that corresponds to a request to scroll the content.

[0158] In some embodiments, such as Figure 7B, upon detecting a user's gaze (e.g., 713b) directed towards scrollable content (e.g., 707) (802c), the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) according to the user's gaze (e.g., 713b) (802e), in response to the detection of a user's gaze (e.g., 713b) directed towards scrollable content (e.g., 706) distinct from a first region of the scrollable content (e.g., 707), and determining that individual parts of the user (e.g., hands or head) meet their respective criteria. In some embodiments, individual parts of the user meet their respective criteria when they are in a predetermined pose relative to the user's torso or another reference point (e.g., within a three-dimensional environment). For example, the user's hand meets its respective criteria when it is at the user's side, in the user's knee, or not otherwise raised (e.g., outside a predetermined region of the three-dimensional environment with an individual spatial orientation relative to the user's torso).

[0159] In some embodiments, such as Figure 7C, upon detecting a user's gaze (e.g., 713c) directed towards scrollable content (e.g., 707) (802c), and in accordance with the determination that the user's gaze (e.g., 713c) is directed towards a second area (e.g., 706) and that individual parts of the user (e.g., 703a) do not meet their respective criteria, the computer system (e.g., 101) maintains the display of the scrollable content (e.g., 707) without scrolling it (802f). In some embodiments, the second area is oriented in one or more directions in which the scrollable content is scrollable. For example, if the scrollable content is vertically scrollable, the second area of ​​the scrollable content is the upper or lower area of ​​the scrollable content. As another example, if the scrollable content is horizontally scrollable, the second area of ​​the scrollable content is the left or right area of ​​the scrollable content. In some embodiments, the computer system scrolls the scrollable content to display a portion of the scrollable content that was not visible when the user's gaze was detected (for example, initially), and displays that portion of the scrollable content in a second area or an area adjacent to the second area. In some embodiments, in response to detecting the user's gaze directed towards a first area of ​​the scrollable content, the computer system scrolls the content in a first direction to display a new portion of the content at a location in the first area or an area adjacent to the first area. In some embodiments, as will be described in more detail below, in response to detecting the user's gaze directed towards a second area of ​​the scrollable content, the computer system scrolls the content in a second direction to display a new portion of the content at a location in the second area or an area adjacent to the second area.

[0160] Scrolling scrollable content according to the user's gaze provides an efficient way to navigate scrollable content and improves user interaction with the computer system by reducing the number of inputs required to perform the action (e.g., scrolling by gaze, in addition to or instead of gaze detection, or scrolling in response to input).

[0161] In some embodiments, such as Figure 7B, each criterion includes a criterion that is met when a distinct part of the user (e.g., 703a) is not detected in a default pose (e.g., the user's hands are not in a ready state and / or the user's hands are not visible) (804). In some embodiments, detecting a default pose includes detecting a distinct part of the user that is in a ready state. In some embodiments, the criterion is met when a distinct part of the user is in a static pose and / or a pose that does not indicate an intention to interact with the computer system. For example, a distinct part of the user is the user's hands, and the criterion is met when the hands are in the user's knees, beside the user, not in the field of view of the hand tracking device, or possibly not lifted, and / or not in a ready state. In some embodiments, while scrolling scrollable content according to the user's gaze, the computer system stops scrolling the scrollable content in response to detecting a distinct part of the user in a default pose (e.g., detecting a ready state) while the user continues to look at a second area.

[0162] Displaying scrollable content without scrolling, in response to detecting individual parts of the user in poses other than the default pose, improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0163] In some embodiments, while displaying a user interface including scrollable content (e.g., 707) (806a), a computer system (e.g., 101) detects input directed to individual user interface elements (e.g., user interface elements within scrollable content) via one or more input devices, and detecting input includes detecting the user's gaze directed to individual user interface elements, such as detecting the gaze 713e in Figure 7D directed to a selectable user interface element and detecting that the hand 703b is performing an individual gesture, and detecting that the user is performing an individual gesture in an individual part of the user (806b). In some embodiments, the input is an air gesture. In some embodiments, detecting that the user is performing an individual gesture in an individual part of the user includes detecting that the user is performing a gesture (e.g., a pinch gesture or a tap gesture) with the hand included in the air gesture input. In some embodiments, an individual part of the user does not meet the respective criteria when the computer system detects an individual gesture. In some embodiments, the input corresponds to a request to select an individual user interface element.

[0164] In some embodiments, while displaying a user interface containing scrollable content (806a), the computer system (e.g., 101) performs an action associated with a particular user interface element (806c) in response to detecting input directed to that particular user interface element. In some embodiments, the action associated with a particular user interface element is an action performed in response to detecting a selection of that particular user interface element. For example, in response to detecting input directed to an option to navigate to a particular user interface, the computer system presents that particular user interface. As another example, in response to detecting input directed to an option to play or pause a content item, the computer system plays or pauses the content item.

[0165] Detecting input directed towards individual user interface elements, including the user's gaze and individual gestures with specific body parts of the user, and then performing actions associated with those individual user interface elements, improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0166] In some embodiments, such as Figure 7A, the second region of the scrollable content (e.g., 706) includes (808) the edges of the scrollable content (e.g., 707). In some embodiments, the second region includes and / or is located adjacent to the upper, lower, left, or right edges of the scrollable content. In some embodiments, the second region includes and / or is located on the edges corresponding to the direction in which the scrollable content is scrollable. For example, the second region includes or is adjacent to the upper or lower edge of the scrollable content in the vertical direction, or the second region includes or is adjacent to the left or right edge of the scrollable content in the horizontal direction. Including the edges of the scrollable content in the second region improves user interaction with the computer system by providing additional control options without disrupting the user interface.

[0167] In some embodiments, such as those shown in Figures 7A and 7B, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a first direction according to the determination that the user's gaze is directed towards a second area (e.g., 706). For example, the computer system scrolls the scrollable content downward according to the determination that the user's gaze is directed towards an area along the bottom of the scrollable content. In another example, the computer system scrolls the scrollable content upward according to the determination that the user's gaze is directed towards an area along the top of the scrollable content.

[0168] In some embodiments, while displaying a user interface (e.g., 702) including scrollable content (e.g., 707) via a display generation component (e.g., 120), the computer system (e.g., 101) detects the user's gaze directed towards the scrollable content (e.g., 707) and, in accordance with the determination that the user's gaze is directed towards a third region of the scrollable content (e.g., region 704 in Figure 7B) and that the third region (e.g., 704) is different from the second region (e.g., 706) and the first region, and that the individual parts of the user meet their respective criteria, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a second direction different from a first direction such as in Figure 7F, according to the user's gaze, the second region (e.g., 706) and the third region (e.g., 704) have different sizes (810b). In some embodiments, the second direction is opposite to the first direction, and the third region is positioned along the edge of the scrollable content opposite to the edge of the scrollable content where the second region is located. In some embodiments, the second and third regions have the same size (e.g., width, length, and / or height) along the first direction and different sizes (e.g., width, length, and / or height) along the second direction. For example, the second and third regions have the same width and different heights.

[0169] Scrolling scrollable content in different directions depending on whether the user's gaze is directed to a second or third area improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0170] In some embodiments, such as FIG. 7A, a second portion (e.g., 706) of the scrollable content (e.g., 707) is located below the scrollable content (e.g., 707) and has a first size (e.g., height, width, or length) (812a). In some embodiments, in response to detecting a user's line of sight directed to the second region while individual portions of the user meet their respective criteria, the computer system scrolls the scrollable content downwards.

[0171] In some embodiments, such as FIG. 7A, a third portion (e.g., 704) of the scrollable content (e.g., 707) is located above the scrollable content (e.g., 707) and has a second size (e.g., height, width, or length) that is smaller than the first size (812b). In some embodiments, in response to detecting a user's line of sight directed to the third region while individual portions of the user meet their respective criteria, the computer system scrolls the scrollable content upwards. In some embodiments, the height of the third region is smaller than the height of the second region. In some embodiments, the widths of the second region and the third region are the same. In some embodiments, the widths of the second region and the third region are different.

[0172] Providing a third region smaller than the second region of the scrollable content at the upper part of the scrollable content improves the user interaction with the computer system by providing additional control options to the user without confusing the user interface.

[0173] In some embodiments, scrolling scrollable content (e.g., 707) in accordance with the user's gaze includes, as shown in Figure 7A, determining that the user's gaze (e.g., 713a) is directed to a location at a first distance from an individual position of the scrollable content (e.g., 707), and then scrolling the scrollable content (e.g., 707) at a first speed in accordance with the user's gaze (814b) (814a), as shown in Figure 7B. In some embodiments, the individual position of the scrollable content is the boundary of a second region and / or the start / end of the scrollable content. In some embodiments, the boundary of the second region of the scrollable content is either the boundary of the second region or adjacent to the boundary of the second region. For example, if the second region is along the bottom of the scrollable content, then the boundary is the bottom region of the scrollable content.

[0174] In some embodiments, scrolling scrollable content (e.g., 707) in accordance with the user's gaze includes, as shown in Figure 7B, determining that the user's gaze (e.g., 713b) is directed to a location at a second distance different from a first distance from an individual position of the scrollable content (e.g., 707), and scrolling the scrollable content (e.g., 707) at a second speed different from a first speed, as shown in Figure 7C (814c) (814a). In some embodiments, the scrolling speed increases as the gaze gets closer to the boundary of the scrollable content. In some embodiments, the scrolling speed changes as the user's gaze moves within a second area of ​​the scrollable content. For example, the scrolling speed gradually increases as the user's gaze moves toward an individual position of the scrollable content.

[0175] Scrolling scrollable content at different speeds according to the distance between the user's line of sight and the individual positions of the scrollable content improves the user interaction with the computer system by providing additional control options without confusing the user interface with additional displayed controls.

[0176] In some embodiments, the user's line of sight (e.g., 713b) is directed to a second region (e.g., 706) of the scrollable content (e.g., 707), and while the individual parts of the user meet their respective criteria and while scrolling the scrollable content (e.g., 707) along the user's line of sight as in FIG. 7B, the computer system (e.g., 101) detects (816a) the user's line of sight (e.g., 713d) directed away from the second region of the scrollable content as in FIG. 7C via one or more input devices. In some embodiments, the computer system detects the user's line of sight directed to the first region of the scrollable content. In some embodiments, the computer system detects the user's line of sight directed to a region of the three-dimensional environment that does not include the scrollable content. In some embodiments, the computer system detects that the user directs their line of sight away from the three-dimensional environment or closes their eyes for a threshold time (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds) associated with a blink.

[0177] In some embodiments, as shown in Figure 7C, in response to detecting a user's gaze (e.g., 713d) directed away from a second area (e.g., 706) of the scrollable content (e.g., 707), the computer system (e.g., 101) reduces the speed at which the scrollable content is scrolling until scrolling of the scrollable content (e.g., 707) is stopped (816b), as shown in Figure 7D. In some embodiments, the computer system stops scrolling the scrollable content in response to detecting a user's gaze directed away from the second area of ​​the scrollable content by slowing down the scrolling speed with simulated inertia until scrolling is stopped. In some embodiments, while the scrolling speed of the scrollable content is slowed down and the scrollable content continues to scroll, the computer system accelerates the scrolling speed of the scrollable content in response to detecting a user's gaze directed towards the second area of ​​the scrollable content while individual parts of the user meet their respective criteria. In some embodiments, in this situation, the computer system increases the scrolling speed until the scrolling speed reaches a predetermined speed (e.g., a speed associated with the location in the second area that the user is looking at, as described above).

[0178] Slowing down the scrolling of scrollable content until scrolling is stopped in response to detecting that the user's gaze is directed away from a second area of ​​the scrollable content improves user interaction with the computer system by providing the user with enhanced visual feedback (for example, informing the user that scrolling will stop if they continue to look away from the second area).

[0179] In some embodiments, as shown in Figure 7A, in response to detecting the user's gaze (e.g., 713a) directed towards scrollable content (e.g., 707), scrolling the scrollable content (e.g., 707) according to the user's gaze, in accordance with the determination that the user's gaze (e.g., 713a) is directed towards a second area (e.g., 706) and that individual parts of the user meet their respective criteria, includes gradually increasing the speed at which the scrollable content (e.g., 707) is scrolled while the user's gaze (e.g., 713a) is directed towards the second area (e.g., 706) and individual parts of the user meet their respective criteria (818). In some embodiments, the computer system gradually increases the scroll speed until the scroll speed reaches a predetermined speed (e.g., a speed associated with a location in the second area that the user is looking at, as described above). In some embodiments, the computer system gradually decreases the scroll speed to zero in response to the user shifting their gaze from the second area to the first area, as described above. In some embodiments, the computer system gradually changes the scrolling speed in response to the user updating their gaze to a location at a different distance from the edge of the content within the second area.

[0180] Gradually increasing the scrolling speed of scrollable content in response to detecting that the user's gaze is directed towards a second area of ​​the scrollable content improves user interaction with the computer system by providing the user with enhanced visual feedback (for example, indicating to the user that scrolling will continue if the user continues to look at the second area).

[0181] In some embodiments, while displaying a user interface (e.g., 702) including scrollable content (e.g., 707) (820a), a computer system (e.g., 101) detects, via one or more input devices (e.g., hand tracking devices), that a distinct part of the user performs a distinct gesture, including the movement of the user's hand (e.g., 703b) while the user's hand is in a pinch-hand shape such as in Figure 7D, and that the distinct part of the user does not meet the respective criteria while performing the distinct gesture (820b). In some embodiments, the distinct gesture includes detecting that the user is making a pinch shape with their hand (e.g., a hand shape in which the thumb is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, or 1 centimeter) of another finger of the hand or touching it) and is moving their hand while maintaining the pinch shape. In some embodiments, upon detecting that the user has stopped performing a pinch gesture with their hand, the computer system stops scrolling the scrollable content according to any further hand movement detected while the hand is not in a pinch position (e.g., air gesture, touch input, or other hand input).

[0182] In some embodiments, as shown in Figure 7E, while a user interface (e.g., 702) containing scrollable content (e.g., 707) is being displayed (e.g., 820a), the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in accordance with the user's hand movement (e.g., air gesture, touch input, or other hand input) (e.g., 703c) (820c), in response to detecting that a separate part of the user (e.g., 703b) has performed a separate gesture. In some embodiments, the computer system scrolls the scrollable content in accordance with the hand movement (e.g., air gesture, touch input, or other hand input) while the hand is in a pinch position. For example, the computer system scrolls the content by an amount corresponding to the amount of movement (e.g., speed, duration, and / or distance) in the same direction in which the hand is moving while in a pinch position. In some embodiments, while scrolling the scrollable content in accordance with an air gesture input, the computer system does not scroll the scrollable content in accordance with the line of sight. For example, while detecting air gesture input (e.g., in response to a request to scroll scrollable content, a different request regarding scrollable content, or a request independent of scrollable content), if the computer system detects the user's gaze directed towards a second area of ​​scrollable content, it may stop scrolling the scrollable content in accordance with the fact that the gaze is directed towards the second area of ​​scrollable content.

[0183] Scrolling scrollable content in accordance with the user's hand movements (e.g., air gestures, touch input, or other manual input) while the user's hand is in a pinch-hand position improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0184] In some embodiments, such as Figure 7D, the movement (e.g., speed, distance, and / or duration) of individual parts of the user (e.g., 703b) has individual magnitudes (822a).

[0185] In some embodiments, upon determination that the movement of a distinct part of the user (e.g., 703b) is in a first direction, such as Figure 7D, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a second direction by a first amount (822b), in response to detecting that the distinct part of the user (e.g., 703b) has performed a distinct gesture, such as Figure 7E. In some embodiments, the second direction in which the computer system scrolls the scrollable content corresponds to the first direction of movement of the distinct part of the user. In some embodiments, the second and first directions are the same direction (e.g., scrolling up by moving the distinct part of the user upwards, or scrolling down by moving the distinct part of the user downwards). In some embodiments, the second and first directions are opposite directions (e.g., scrolling down by moving the distinct part of the user upwards, or scrolling up by moving the distinct part of the user downwards). In some embodiments, the first amount corresponds to a distinct size. When the individual sizes are larger, the first quantity is larger, and when the individual sizes are smaller, the first quantity is smaller.

[0186] In some embodiments, upon determining that the movement of an individual part of the user (e.g., 703c) is in a third direction different from a first direction, such as in Figure 7E, the computer system (e.g., 101) detects that the individual part of the user (e.g., 703c) has performed an individual gesture, and scrolls the scrollable content (e.g., 70) in a fourth direction by a second amount different from the first amount, such as in Figure 7F (822c). In some embodiments, the fourth direction in which the computer system scrolls the scrollable content corresponds to the third direction of movement of the individual part of the user. In some embodiments, the fourth and third directions are the same direction (e.g., scrolling up by moving the individual part of the user upwards, or scrolling down by moving the individual part of the user downwards). In some embodiments, the fourth and third directions are opposite directions (e.g., scrolling down by moving the individual part of the user upwards, or scrolling up by moving the individual part of the user downwards). In some embodiments, the second amount corresponds to an individual size. If the individual size is larger, the second amount is larger, and if the individual size is smaller, the second amount is smaller. In some embodiments, upon detecting downward movement of an individual part of a user having individual size, the computer system scrolls the scrollable content by a smaller amount than the amount the computer system scrolls the scrollable content upon detecting upward movement of an individual part of a user having the same individual size.

[0187] Scrolling scrollable content by different amounts in response to the movement of individual parts of the user with distinct sizes in different directions improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0188] In some embodiments, such as Figure 7E, the movement of the user's hand (e.g., air gesture, touch input, or other hand input) (e.g., 703c) includes the movement of the hand (e.g., air gesture, touch input, or other hand input) (e.g., 703c) from a first location to a second location, where the user's hand (e.g., 703c) maintains a pinch-hand shape while moving from the first location to the second location (824a). In some embodiments, the first location is the location of a particular part of the user's hand when that particular part first makes a pinch-hand shape, such as when the thumb and index finger of the user's hand touch together.

[0189] In some embodiments, as shown in Figure 7E, scrolling scrollable content (e.g., 707) in response to detecting that a separate part of the user (e.g., 703c) has performed a separate gesture includes scrolling the scrollable content (e.g., 707) at a first speed (824c) (824b) according to the determination that the distance between a first location and a second location is a first distance. In some embodiments, the computer system continues to scroll the scrollable content at a first speed while continuing to detect a default part of the user at a second location that is a first distance from the first location.

[0190] In some embodiments, as shown in Figure 7E, scrolling scrollable content (e.g., 707) in response to detecting that a separate part of the user (e.g., 703c) has performed a separate gesture includes (824d) scrolling the scrollable content (e.g., 707) at a second speed greater than the first speed, according to the determination that the distance between a first location and a second location is a second distance greater than the first distance (824b). In some embodiments, the computer system continues to scroll the scrollable content at the second speed while it continues to detect a default part of the user at a second location at a second distance from the first location. In some embodiments, as the user's hand moves while the user's hand is in a pinch position, the computer system changes the scrolling speed of the scrollable content according to the distance between the user's hand's current location and the user's hand's first location.

[0191] Scrolling scrollable content at a speed dependent on the distance between the user's first and second hand locations improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0192] In some embodiments, one or more criteria are met when the user's hand (e.g., 703b) moves a distance (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, 1, 2, 3, or 5 centimeters / second), a distance (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 centimeters), and / or duration (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, or 1 second)) while maintaining a pinch hand shape (826a).

[0193] In some embodiments, as shown in Figure 7C, the computer system (e.g., 101) detects that a separate part of the user (e.g., 703a) has performed a separate gesture, and, according to the determination that the user's hand movement (e.g., air gesture, touch input, or other hand input) (e.g., 703a) does not meet one or more criteria, maintains the display of the scrollable content (e.g., 707) without scrolling it (826b). In some embodiments, if the hand movement (e.g., air gesture, touch input, or other hand input) while the hand is in a pinch position is less than a threshold, the computer system stops scrolling the scrollable content according to the hand movement (e.g., air gesture, touch input, or other hand input) while the hand is in a pinch position.

[0194] Maintaining the display of scrollable content without scrolling it, in response to detecting individual portions of users performing individual gestures that do not meet one or more criteria because their hand movements (e.g., air gestures, touch input, or other manual input) are below a threshold amount, improves user interaction with the computer system by reducing user errors when interacting with the computer system.

[0195] In some embodiments, one or more criteria are not met (828a) when the speed of movement of the user's hand (e.g., air gesture, touch input, or other manual input) (e.g., hand 703a in Figure 7C) is greater than a threshold speed (e.g., 1, 2, 3, 5, 10, 15, 30, or 50 centimeters / second) and the direction of movement of the user's hand (e.g., air gesture, touch input, or other manual input) (e.g., 703a) is downward. In some embodiments, the threshold speed is associated with the speed at which the user drops their hand without the intention of continuing to scroll the scrollable content.

[0196] In some embodiments, as shown in FIG. 7C, in response to detecting that an individual part of a user has performed an individual gesture and in accordance with a determination that one or more criteria are not met, a computer system (e.g., 101) maintains the display of scrollable content (e.g., 707) without scrolling the scrollable content (e.g., 707) (828b). In some embodiments, the computer system scrolls the scrollable content at a speed less than a threshold speed in accordance with a part of a downward movement of a hand (e.g., an air gesture, a touch input, or other hand input). For example, if a hand movement (e.g., an air gesture, a touch input, or other hand input) includes a first part of a downward movement at a speed less than the threshold speed and a second part of a downward movement that exceeds the threshold speed, the computer system scrolls the scrollable content in accordance with the first part of the downward movement without further scrolling the scrollable content in accordance with the second part of the downward movement.

[0197] Maintaining the display of scrollable content without scrolling the scrollable content, in response to detecting an individual part of a user who performs an individual gesture that does not meet one or more criteria because the hand movement (e.g., an air gesture, a touch input, or other hand input) is downward at a speed that exceeds the threshold speed, improves the user interaction with the computer system by reducing user errors when interacting with the computer system.

[0198] In some embodiments, upon detecting a user's gaze (e.g., 713a) directed towards scrollable content (e.g., 707), and determining that the user's gaze (e.g., 713a) is directed towards a second area (e.g., 706) of the scrollable content (e.g., 707) and that individual parts of the user meet their respective criteria, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a first direction according to the user's gaze (e.g., 713a) (830a), as shown in Figure 7A. In some embodiments, the first direction of scrolling corresponds to the location of a second area of ​​the scrollable content within the scrollable content. For example, if the second area is at the bottom of the scrollable content, the computer system scrolls the content downwards (e.g., to show additional content at the bottom of the scrollable content).

[0199] In some embodiments, as shown in Figure 7F, while displaying a user interface (e.g., 120) containing scrollable content (e.g., 707) via a display generation component (e.g., 120), the computer system (e.g., 101) detects the user's gaze directed towards the scrollable content. If the user's gaze is directed towards a third region of the scrollable content (e.g., region 704 in Figure 7A), and the third region (e.g., 704) differs from a second region (e.g., 706), and according to the determination that the individual parts of the user meet their respective criteria, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a second direction opposite to the first direction according to the user's gaze (830b). In some embodiments, the second direction of scrolling corresponds to the location of the third region of the scrollable content within the scrollable content. For example, if the third region is at the top of the scrollable content, the computer system scrolls the content upwards (e.g., showing additional content at the top of the scrollable content). In some embodiments, scrolling scrollable content in response to detection of a user's gaze directed towards a third area of ​​the content includes scrolling the content along an axis different from the axis on which the computer system scrolls the scrollable content in response to detection of a user's gaze directed towards a second area of ​​the scrollable content. For example, the computer system scrolls the scrollable content vertically in response to detection of a user's gaze directed towards an area along the top or bottom of the content, and scrolls the scrollable content horizontally (for example, while individual parts of the user meet one or more criteria) in response to detection of a user's gaze directed towards an area along the left or right of the scrollable content.

[0200] Scrolling scrollable content in different directions depending on the area the user's gaze is directed towards improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0201] In some embodiments, as shown in Figure 7F, in response to the detection of a user's gaze (e.g., 713a) directed towards scrollable content (e.g., 707), the computer system scrolls the scrollable content (e.g., 707) in a first direction, such as in Figure 7B, according to the user's gaze, in response to the detection of a user's gaze (e.g., 713a) directed towards scrollable content (e.g., 707), in response to the determination that the user's gaze (e.g., 713a) is directed towards a second region (e.g., 706) of the scrollable content (e.g., 707), and according to the determination that individual parts of the user meet their respective criteria, scrolling the scrollable content in a first direction, such as in Figure 7B, includes scrolling the scrollable content with a first acceleration (832a). In some embodiments, the first direction of scrolling corresponds to the location of a second region of the scrollable content within the scrollable content. For example, if the second region is at the bottom of the scrollable content, the computer system scrolls the content downward (e.g., to show additional content at the bottom of the scrollable content). In some embodiments, the first acceleration is the acceleration at which the computer system begins scrolling the scrollable content in response to the detection of a user's gaze directed towards a second region of the scrollable content, in response to the determination that individual parts of the user meet their respective criteria. In some embodiments, the computer system scrolls the scrollable content at a first speed in response to detecting the user's gaze directed towards a second area of ​​the scrollable content while individual parts of the user meet their respective criteria.

[0202] In some embodiments, in response to the detection of a user's gaze directed towards scrollable content (e.g., 707), the computer system scrolls the scrollable content in a second direction, such as in Figure 7F, according to the user's gaze, in accordance with the determination that individual parts of the user meet their respective criteria (832b). In some embodiments, the second direction of scrolling corresponds to the location of the third region of the scrollable content within the scrollable content. For example, if the third region is at the top of the scrollable content, the computer system scrolls the content upward (e.g., to show additional content at the top of the scrollable content). In some embodiments, the second acceleration is the acceleration at which the computer system begins scrolling the scrollable content in response to the detection of a user's gaze directed towards the third region of the scrollable content, in accordance with the determination that individual parts of the user meet their respective criteria. In some embodiments, the computer system scrolls the scrollable content at a second speed different from the first speed mentioned above, in response to detecting the user's gaze directed towards a third area of ​​the scrollable content while individual parts of the user meet their respective criteria.

[0203] Scrolling scrollable content at different accelerations when the user's gaze is directed to different areas of the scrollable content improves user interaction with the computer system by providing additional control options without cluttering the user interface with displayed controls.

[0204] In some embodiments, such as Figure 7A, the scrollable content includes text content (e.g., 707) and other content (e.g., 705) (e.g., images, interactive content, and / or interactive user interface elements) (834a). In some embodiments, the other content includes additional text content not included in the text content of the scrollable content. For example, an article includes text content containing the text of the article and other content including advertisements containing the text content of advertisements. In some embodiments, the other content includes multimedia and / or interactive content such as selectable options (e.g., links to other content) for navigating a user interface containing the scrollable content. In some embodiments, the computer system displays the scrollable content containing the text content and other content in a first mode (e.g., browsing mode) and the text content without other content in a second mode (e.g., reader mode). In some embodiments, the computer system transitions between displaying scrollable content in a first mode and displaying the text content of the scrollable content in a second mode, in response to one or more user inputs (e.g., selection of one or more user interface elements, voice input, and / or default gestures performed by a part of the user's body) that respond to a request to change the presentation mode.

[0205] In some embodiments, as shown in Figure 7G, while displaying text content (e.g., 707) of the scrollable content, without displaying other content of the scrollable content (834b), a computer system (e.g., 101) detects movement of the user's gaze (e.g., 713h) via one or more input devices (834c). In some embodiments, the movement of the user's gaze corresponds to the user reading text content of the scrollable content.

[0206] In some embodiments, while displaying text content (e.g., 707) of scrollable content, without displaying other content of the scrollable content (834b), the computer system (e.g., 101) scrolls the text content (e.g., 707) (e.g., Figure 7H) (834e) in response to detecting movement of the user's gaze (e.g., 713h) (834d), according to a determination that the movement of the user's gaze (e.g., 713h) satisfies one or more criteria, including criteria that are satisfied based on the movement of the user's gaze (e.g., 713h) to a line of text in the text content (e.g., 707) such as Figure 7G. In some embodiments, one or more criteria are associated with the user having finished reading a line of text content. In some embodiments, the computer system can detect, based on the detected movement of the user's eyes, whether the user is merely looking at a first part of the text or whether the user is reading a first part of the text item. The computer system optionally compares one or more captured images of the user's eyes to determine whether the movement of the user's eyes matches a movement that matches reading. In some embodiments, people tend to move their gaze from the end of a line they have finished reading to the beginning of that line, or from the end of a line of text to the beginning of the next line. In some embodiments, one or more criteria include criteria that are met when the user's gaze moves from the end of a line to the beginning of the line or the beginning of the next line. In some embodiments, in response to detecting the user's gaze movement corresponding to the user having finished reading a line of text, the computer system scrolls the text content. In some embodiments, the computer system scrolls the text content one line at a time to display the next line at the location in the three-dimensional environment where the line the user had just read was displayed while the user was reading that line of text. For example, the computer system scrolls the text vertically to display individual lines of text at the height where the line the user previously read was previously displayed.As another example, an electronic device scrolls text horizontally to display individual lines of text in the horizontal locations where lines of text previously read by the user were previously displayed. Scrolling text content optionally includes updating the location of lines of text previously read by the user (e.g., moving a first portion of text vertically or horizontally to make space for a second portion of text), or ceasing to display lines of text previously read by the user. In some embodiments, a computer system scrolls text content in response to data collected by an eye-tracking device without receiving additional input from another input device communicating with the computer system (e.g., air gesture input or input detected via a hardware input device).

[0207] In some embodiments, as shown in Figure 7G, while displaying text content (e.g., 707) of scrollable content, without displaying other content of the scrollable content (834b), the computer system (e.g., 101) maintains the display of the text content (e.g., 707) without scrolling the text content, in response to detecting movement of the user's gaze (e.g., 713h) (834d), and determining that the movement of the user's gaze (e.g., 713h) does not meet one or more criteria (834f). In some embodiments, the user's gaze does not meet one or more criteria while the user is reading an individual line of scrollable content (e.g., the beginning or middle of the line). In some embodiments, the user's gaze does not meet one or more criteria when the user reaches the end of a line of text content without moving toward the beginning of the line of text content. For example, the user reads a line of text content and then turns their gaze toward another part of the three-dimensional environment that is different from the beginning of the line of text content or the beginning of the next line of text content.

[0208] Scrolling text content based on whether the user's eye movement meets one or more criteria improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0209] In some embodiments, as shown in Figure 7G, scrolling text content (e.g., 707) in response to detecting user gaze movements (e.g., 713h) that satisfy one or more criteria is independent of whether the individual part of the user is detected in a default pose (836). In some embodiments, the individual part of the user is in a default pose when the user's hands are ready. In some embodiments, while the computer system is displaying the text content of scrollable content without displaying additional content of the scrollable content (e.g., in reader mode), the computer system scrolls the text content according to the user's gaze, regardless of the user's hand pose and / or location. In some embodiments, the computer system scrolls the text content in response to detecting user gaze movements that satisfy one or more criteria while the individual part of the user satisfies each criterion. In some embodiments, the computer system scrolls the text content in response to detecting user gaze movements that satisfy one or more criteria while the individual part of the user does not satisfy each criterion.

[0210] Scrolling text content based on whether the user's eye movements meet one or more criteria, regardless of whether individual parts of the user's gaze are in a default position, improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0211] In some embodiments, while displaying the text content of scrollable content without other content of the scrollable content (838a) (for example, while displaying the text content in the reader mode described above), a computer system (e.g., 101) detects the user's gaze directed at the text content via one or more input devices (838b).

[0212] In some embodiments, while displaying the text content of scrollable content without other content of the scrollable content (838a) (for example, while displaying the text content in the reader mode described above), in response to detecting the user's gaze directed towards the text content (838c), the computer system (e.g., 101) maintains the display of the text content without scrolling the text content (838d) according to the determination that the user's gaze is directed towards a first area of ​​the text content and that the movement of the user's gaze does not meet one or more criteria. In some embodiments, the first area of ​​the text content is away from one or more directions in which the text content is scrollable. For example, if the text content is scrollable vertically, the first area of ​​the text content is the area of ​​the text content between the top and bottom of the text content. As another example, if the text content is scrollable horizontally, the first area of ​​the text content is the area of ​​the text content between the left and right portions of the text content. In some embodiments, the first area of ​​the text content is similar to the first area of ​​scrollable content described above. In some embodiments, the computer system maintains the display of the text content without scrolling it, in response to detecting that the user's gaze is directed towards a first area of ​​the text content, while the user's eye movement does not correspond to the user reading the text content.

[0213] In some embodiments, as shown in Figure 7G, while displaying the text content of scrollable content (e.g., 707) without other content of the scrollable content (e.g., while displaying the text content in the reader mode described above) (838a), the computer system (e.g., 101) detects the user's gaze directed at the text content (838c), and determines that the user's gaze is directed to a second area of ​​the text content (e.g., 710) different from a first area of ​​the text content, that individual parts of the user (e.g., hands or head) meet their respective criteria (e.g., the user's hands are not in a ready state), and that the movement of the user's gaze does not meet one or more criteria, the computer system (e.g., 101) scrolls the text content according to the user's gaze (838e). In some embodiments, scrolling text content according to the user's gaze, based on the determination that the user's gaze is directed to a second area of ​​text content and that individual parts of the user meet their respective criteria, has one or more characteristics in common with the techniques described above for scrolling scrollable content, in response to the detection of the user's gaze directed to a second area of ​​scrollable content while individual parts of the user meet their respective criteria. In some embodiments, the computer system scrolls the text content in response to the detection of the user's gaze directed to a second area of ​​text content, and the movement of the user's gaze corresponds to the user reading the text content. In some embodiments, the computer system scrolls the text content in response to the detection of the user's gaze directed to a second area of ​​text content while the movement of the user's gaze does not correspond to the user reading the text content. In some embodiments, the second area of ​​text content is similar to the second area of ​​scrollable content described above.

[0214] In some embodiments, as shown in Figure 7G, while displaying the text content of scrollable content (e.g., 707) without other content of the scrollable content (838a) (e.g., while displaying the text content in the reader mode described above), the computer system (e.g., 101) scrolls the text content as shown in Figure 7H (838f) in response to detecting the user's gaze directed at the text content (e.g., 713h) (838c), and determining that the user's gaze (e.g., 713h) is directed to a first area of ​​the text content and that the movement of the user's gaze (e.g., 713h) satisfies one or more criteria. In some embodiments, the computer system scrolls the text content according to the user's gaze being directed to a second area of ​​the text content, and scrolls the text content according to the movement of the user's gaze across lines of text content as described above, while the computer system displays the text content of scrollable content without other content of the scrollable content as described above. In some embodiments, there are at least two ways of scrolling text content based on gaze.

[0215] Scrolling text content according to the user's gaze improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0216] In some embodiments, while displaying scrollable content (e.g., 707), such as in Figure 7H, the computer system (e.g., 101) detects a user's gaze (e.g., 713i) directed at the scrollable content (e.g., 707), and determines that the user's gaze (e.g., 713i) has been directed at a word contained in a first region of the scrollable content (e.g., 707) for at least a threshold time (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds), then displays the definition (e.g., 712) of the word contained in the scrollable content (e.g., 707) via a display generation component (e.g., 120) (840). In some embodiments, the word definition is displayed as an overlay on the scrollable content. The computer system stops displaying the word definition if it determines that the user's gaze has been directed at a word contained in a first region of the scrollable content for less than a threshold time. If the computer system determines that the user's gaze is not directed towards a word in the first area of ​​the scrollable content, it will stop displaying the definition of the word.

[0217] Displaying word definitions according to the user's gaze improves user interaction with the computer system by providing additional control options without cluttering the user interface with extra displayed controls.

[0218] In some embodiments, aspects / operations of methods 1000, 1200, 1400, 1600, 1800, 2000, 2200, and / or 2400 may be interchangeable, replaced, and / or added to among these methods. For example, a computer system optionally scrolls content generated via speech input according to method 1000, following one or more steps of method 800. For example, a computer system optionally scrolls content generated via a soft keyboard according to methods 1200, 1400, and / or 1600, following one or more steps of method 800. For brevity, those details will not be repeated here.

[0219] Figures 9A to 9N illustrate exemplary techniques for entering text into a text entry field in response to voice input, according to several embodiments. The user interfaces in Figures 9A to 9N are used to illustrate the processes described below, including the processes in Figures 10A to 10R.

[0220] Figure 9A shows a computer system 101 that displays a three-dimensional environment 901 from the user's viewpoint via a display generation component (e.g., display generation component 120 in Figure 1). As described above with reference to Figures 1 to 6, the computer system 101 optionally includes a display generation component (e.g., a touchscreen) and a plurality of image sensors 314 (e.g., the image sensor 314 in Figure 3). The image sensors 314 optionally include one or more of the following: a visible light camera, an infrared camera, a depth sensor, or any other sensors that the computer system 101 can use to capture one or more images of the user or a part of the user (e.g., one or more of the user's hands) while the user is interacting with the computer system 101. In some embodiments, the user interfaces illustrated and described below may also be implemented on a head-mounted display, which includes a display generating component that displays the user interface or three-dimensional environment to the user, and sensors (e.g., external sensors facing outward from the user) and / or the user's hand movements, such as movement which is interpreted as a gesture by a computer system, such as an air gesture. In some embodiments, it should be understood that one or more of the techniques described herein are applicable to a two-dimensional environment (e.g., on a touch-sensitive display or other display) without departing from the scope of this disclosure.

[0221] Figure 9A shows a computer system 101 displaying a web browsing user interface 902 via a display generation component 120. In some embodiments, the web browsing user interface 902 includes an indication 904 of the URL of a website currently presented by the web browser. For example, in Figure 9A, the web browsing user interface 902 includes a web search website that includes a text entry field 906 to which input specifying one or more search terms is directed, and a selectable option 908 that, when selected, causes the computer system 101 to perform a search using the one or more search terms provided in the text entry field 906. In some embodiments, the computer system 101 is configured to detect input for entering text into the text entry field 906 via a soft keyboard, via a hardware keyboard, or via dictation, according to one or more steps of methods 1200, 1400, and 1600 as described herein.

[0222] In some embodiments, the computer system 101 initiates a process of accepting dictation input directed to the text entry field 906 in response to detecting user attention, including the user's gaze 913a directed to the text entry field 906, via one or more input devices (e.g., an image sensor 314). In some embodiments, the computer system 101 initiates a process of accepting dictation input in response to detecting user attention directed to the text entry field 906, without detecting, or regardless of, additional input, such as input provided using air gestures or hardware input devices. In some embodiments, in response to detecting the user's gaze 913a directed to the text entry field 906, the computer system 101 gradually expands the text entry field. For example, as shown in Figures 9A-9B, in response to the user's gaze 913a being directed to the text entry field 906, the computer system 101 gradually increases the width of the text entry field 906 while the user's gaze 913a is directed to the text entry field 906. In some embodiments, when the user's gaze 913a is directed to the text entry field 906 over a threshold time, the computer system 101 stops expanding the text entry field 906 and begins the process of accepting the utterance input directed to the text entry field. Exemplary time thresholds are provided below in the description of Method 1000 with reference to Figures 10A to 10R.

[0223] Figure 9B shows the web browsing user interface 902 updated in response to the computer system 101 detecting the user's gaze 913a directed at the text entry field 906 over the threshold time referenced above. As shown in Figure 9B, the computer system 101 displays the text entry field 906 with a width greater than the width of the text entry field in Figure 9A when the computer system 101 first detects that the user's gaze 913a is directed at the text entry field 906. Figure 9B also shows the computer system 101 generating an audio output 910a indicating that the computer system 101 is configured to accept utterance input for dictating the text directed at the text entry field 906 in response to the user's gaze 913a being directed at the text entry field 906 over the threshold time. The computer system 101 also highlights the placeholder text 914 displayed in the text entry field 906 in response to the user's gaze 913a being directed at the text entry field 906 for a threshold amount of time, before the computer system 101 detects the user's gaze 913a directed at the text entry field 906.

[0224] Figure 9B shows a computer system 101 that displays a cursor 912 in response to the user's gaze 913a being directed over a text entry field 906 for a threshold time, however, in some embodiments, the computer system 101 does not display the cursor 912 unless and until the user provides an utterance input that dictates the text to be entered into the text entry field 906. In some embodiments, in response to detecting the user's gaze 913a being directed over the text entry field 906 for at least a threshold time (without detecting, or regardless of detecting, additional input such as an air gesture or input detected via a hardware input device), the computer system 101 displays an additional visual indication that the computer system 101 is configured to enter the dictated text provided via the utterance input into the text entry field 906, in a manner similar to how the computer system 101 displays the microphone icon 930 in Figures 9G and 9H below.

[0225] In Figure 9B, while continuously detecting the user's gaze 913a directed towards the text entry field 906, the computer system 101 receives utterance input 916a from the user. In response to the input shown in Figure 9B, the computer system 101 displays a text representation of the utterance input 916a in the text entry field 906 in order to input the text of the utterance input into the text entry field 906, as shown in Figure 9C.

[0226] Figure 9C shows a computer system 101 that displays a text representation 920 of the utterance input shown in Figure 9B in a text entry field 906 in response to the input shown in Figure 9B. In some embodiments, the computer system 101 initiates a process of accepting dictation input for entering text into the text entry field 906 in response to detecting the user's gaze, without detecting utterance input or regardless of whether utterance input is detected, as described above with reference to Figures 9A-9B. In some embodiments, the computer system 101 is configured to accept dictation input for entering text into the text entry field 906, but the computer system 101 enters text into the text entry field 906 in response to utterance input, as shown in Figures 9B-9C, without detecting or regardless of whether input detected via air gesture input and / or hardware input devices is detected. In some embodiments, while detecting utterance input, the computer system 101 generates a glow effect 918 around the text entry field 906 that changes over time based on the volume of the received utterance input. For example, while the computer system 101 is receiving the utterance input, it modifies the size, translucency, color, darkness, or other visual characteristics of the glow effect 918 according to the volume of the utterance input. In some embodiments, the computer system 101 displays a cursor 912 within the text entry field 906 while the utterance input is being received.

[0227] In Figure 9C, the computer system 101 detects the subsequent utterance input 916b while the user's gaze 913b is no longer directed towards the text entry field 906. Although Figure 9C shows the user's gaze 913b directed towards an area of ​​the web browsing user interface that does not include the text entry field 906, in some embodiments the user's gaze may be directed away from the web browsing user interface 902, for example, to a different portion of the display generation component 120 from the portion of the display generation component 120 that includes the text entry field 906, or to be directed away from the display generation component 120. In some embodiments the computer system 101 detects the subsequent utterance input 916b while the user has their eyes closed for a period exceeding a time threshold associated with the user blinking. Exemplary time thresholds are provided below in the description of Method 1000 with reference to Figures 10A to 10R.

[0228] In some embodiments, in response to the detection of a subsequent utterance input 916b while the user's gaze 913b is not directed towards the text entry field 906, the computer system 101 inputs a text-based representation of the subsequent utterance input 916b, as described later with reference to Figure 9D. In some embodiments, in response to the detection of a subsequent utterance input 916b while the user's gaze 913b is not directed towards the text entry field 906, the computer system 101 maintains the display of the text representation 920 of the previously entered text without displaying the text representation of the subsequent utterance input 916b, as described later with reference to Figure 9D. In some embodiments, in response to the detection of a subsequent utterance input 916b while the user's gaze 913b is not directed towards the text entry field 906, the computer system 101 removes (e.g., some, all) the text from the text entry field 906 and stops accepting dictation input directed towards the text entry field 906, as described later with reference to Figure 9E.

[0229] In some embodiments, in response to detecting a subsequent utterance input 916b while the user's gaze 913b is not directed at the text entry field 906, the computer system 101 displays a text representation of the subsequent utterance input 916b in the text entry field, as shown in Figure 9D, if the computer system 101 has already begun accepting dictation input, and refrains from displaying a text representation of the subsequent utterance input 916b, as shown in Figure 9E, if the computer system 101 has not already accepted dictation input. In some embodiments, the computer system 101 removes (e.g., some, all) the text from the text entry field 906 and, because the text entry field 906 is a search text entry field, refrains from displaying a text representation of the subsequent utterance input 916b in the text entry field, as shown in Figure 9E. In some embodiments, the search text entry field is included in a first type of text entry field, which also includes a messaging text entry field and a web browser address field. In some embodiments, if the text entry field is a long text entry field, such as the text entry fields shown in Figures 9F to 9H, and / or requires input in addition to detecting the user's attention directed towards the text entry field to accept utterance input to provide text to the text entry field, the computer system 101 continues to display the previously dictated text but does not display the continuation of the text representation of utterance input detected while the user's gaze was not directed towards the text entry field.

[0230] Figure 9D shows a computer system 101 that updates the text entry field 906 in response to a continuation of the utterance input shown in Figure 9C, according to several embodiments. As described above, in some embodiments, in response to a continuation of the utterance input shown in Figure 9C, the computer system 101 maintains the display of the text representation 920 of the utterance input (e.g., the word "Lorem") in the text entry field 906. In some embodiments, the computer system 101 also displays the text representation of the continuation of the utterance input shown in Figure 9C (e.g., the word "Ipsum"). Figure 9D includes a dashed box around the text representation of the continuation utterance input 916b shown in Figure 9C (e.g., the word "Ipsum"). This is because, in some embodiments, as described above, the computer system 101 discontinues displaying the text representation of the continuation utterance input 916b shown in Figure 9C. In some embodiments, it should be understood that the computer system 101 displays the text representation of the subsequent utterance input 916b without displaying a dashed box around the text representation of the subsequent utterance input 916b. In some embodiments, the computer system 101 discontinues displaying the text representation of the subsequent utterance input 916b and discontinues displaying the dashed box. As described above, in some embodiments, the computer system 101 displays the text representation of the subsequent utterance input 916b in the text entry field 906 in Figure 9D because dictation has already started when the continuation of the utterance input in Figure 9C is received, even if the user's gaze was not directed at the text entry field 906 while the subsequent utterance input 916b was detected.

[0231] In Figure 9D, the computer system 101 displays a glow effect 918 with updated visual characteristics around the text entry field 906 in accordance with the change in the volume level of the utterance input 916b shown in Figure 9C. In some embodiments, the computer system 101 displays the glow effect 918 when it displays the continuation of the utterance input text, and does not display the glow effect 918 when it stops displaying the continuation of the utterance input text.

[0232] In some embodiments, while displaying text 920 in text entry field 906, the computer system 101 detects an utterance input 916c corresponding to a command associated with text entry field 906. For example, since text entry field 906 is a search field, the utterance input 916c contains the word "search". Other examples of utterance commands and their associated text entry fields are provided below in the description of method 1000 with reference to Figures 10A to 10R. In some embodiments, if an utterance input 916c corresponding to a command is received while gaze 913c is directed to text entry field 906, the computer system 101 performs an action corresponding to text entry field 906, such as performing a search for the search term(s) contained in the text entry field when the command was received. In some embodiments, if an utterance input 916c corresponding to a command is received while gaze 913b is not directed to text entry field 906, the computer system 101 refrains from performing an action corresponding to text entry field 906. In some embodiments, the computer system 101 performs an action corresponding to the text entry field 906, such as performing a search against the search term(singular or plural) contained in the text entry field when the command is received, regardless of whether the utterance input 916c corresponding to the command was received while gaze 913c was directed towards the text entry field 906 or while gaze 913b was not directed towards the text entry field 906.

[0233] Figure 9E shows a computer system 101 that updates a text entry field in response to the continuation of the utterance input shown in Figure 9C, according to several embodiments. As described above, in some embodiments, the computer system 101 removes the text corresponding to the utterance input shown in Figure 9B from the text entry field 906, according to the continuation of the utterance input detected while the user's gaze is not directed towards the text entry field 906 shown in Figure 9C. In some embodiments, the computer system 101 removes the text corresponding to the utterance input shown in Figure 9B from the text entry field 906, since the text entry field is a website search field or another text entry field of the same type as the search field, as described above with reference to Figure 9C and later in the description of method 1000 with reference to Figures 10A to 10R. In some embodiments, the computer system 101 cancels the dictation input by removing the text corresponding to the utterance input shown in Figure 9B from the text entry field 906. In some embodiments, the computer system updates the appearance of the text entry field 906 to indicate that the dictation input has been canceled, for example, by deleting text from the text entry field or by restoring the appearance of the text entry field 906 to the appearance of the text entry field 906 in Figure 9A (for example, by reducing the width of the text entry field 906).

[0234] Figures 9F–9H show a computer system 101 displaying a word processing user interface 922, which includes a text entry field 926, a save option 924a, an undo option 924b, a font option 924c, and an option 924d for ceasing to display the word processing user interface 922. In some embodiments, the text entry field 926 of the word processing user interface 922 is a long-form text entry field. In some embodiments, the computer system 101 initiates a process to accept dictation input directed to the long-form text entry field 926 in response to additional input to initiate dictation to the text entry field 926, such as input detected via an air gesture or a hardware input device. Exemplary inputs are described in the following description of method 1000 with reference to Figures 10A–10R. In some embodiments, the computer system initiates dictation upon detecting user attention directed at the text entry field 926, including detecting the user's gaze 913d directed at the text entry field 926 over a threshold time, without receiving, or regardless of receiving, additional input such as air gestures or input detected via a hardware input device. Exemplary threshold times are described below in the description of Method 1000 with reference to Figures 10A–10R. In some embodiments, before dictation is initiated, the computer system 101 displays a cursor 928 in the text entry field 926 indicating the location where text will be inserted in response to input provided via a soft keyboard and / or hardware keyboard according to Methods 1200, 1400, and / or 1600. Once dictation is initiated, as described with reference to Figures 9G–9H, the computer system 101 ceases displaying the cursor 928.

[0235] Figure 9G shows how the computer system 101 updates the word processing user interface 922 in response to the start of dictation. In some embodiments, dictation is initiated based on detecting the user's gaze directed towards the text entry field 926, as shown in Figure 9F, without detecting, or regardless of, additional input such as input detected via air gestures or hardware input devices. In some embodiments, dictation is initiated in response to additional input, as described later in the description of method 1000 with reference to Figures 10A-10R. As shown in Figure 9G, once dictation is initiated, the computer system 101 produces an audio output 910b that is the same as or different from the audio output 910a described above with reference to Figure 9B. Figure 9G also shows that the computer system 101 displays a microphone icon 930 at a location in the text entry field 926 where the dictated text is inserted, in response to detecting utterance input provided by the user. In some embodiments, the microphone icon 930 is displayed at the location in the text entry field where the user's gaze was directed when dictation was initiated. Therefore, in some embodiments, if the user is looking at a different location within the text entry field 926, the computer system 101 displays a microphone icon 930 at that location instead of the location shown in Figure 9G. As shown in Figure 9G, once dictation has begun and the computer system 101 displays the microphone icon 930 at the location to be inserted, the computer system 101 stops displaying the cursor 928 as shown in Figure 9F. In some embodiments, instead of displaying the microphone icon 930 as shown in Figure 9G, the computer system 101 displays a different visual indication at the location within the text entry field 926 where the dictated text will be inserted.

[0236] In Figure 9G, the computer system detects the voice input 916d provided by the user while the user's gaze 913d is directed towards the text entry field 926. In some embodiments, upon receiving the voice input 926d, the computer system 101 displays the text corresponding to the voice input 926d in the text entry field 926, as shown in Figure 9H.

[0237] Figure 9H shows a computer system 101 that displays text 932 corresponding to the voice input shown in Figure 9G in a text entry field. In some embodiments, while displaying text 932 corresponding to voice input, the computer system continues to display a microphone icon 930 (for example, if dictation is still active). As shown in Figure 9H, the microphone icon 930 is displayed after text 932 corresponding to voice input because additional text corresponding to voice input is displayed after text 932 corresponding to voice input. In some embodiments, if the user's gaze is directed away from the text entry field 926 while the user continues to dictate text, the computer system 101 maintains the display of text 932 corresponding to voice input and optionally, since text entry field 926 is a long-form text entry field, enters text corresponding to subsequent voice entry input detected while the user's gaze was directed away from text entry field 926, as described above in the description of method 1000 with reference to Figures 10A to 10R and described in more detail below.

[0238] Figures 9I to 9N illustrate an example in which a computer system 101 inputs text into a text entry field 906 in response to voice input. In Figure 9I, the computer system 101 displays a web browsing user interface 902 that includes a text entry field 906 that accepts user input specifying a website address and / or search terms for a web search. For example, upon detecting that a user has entered text into the text entry field 906, and subsequently detecting input for performing a web search using the text (e.g., selecting search options, performing search gestures, and / or search voice commands), the computer system 101 initiates a web search for content on the Internet corresponding to the text. As shown in Figure 9I, the text entry field 906 includes placeholder text 934. In some embodiments, the placeholder text 934 is displayed in a predetermined pattern, or in colors that animate the hue, darkness, and / or saturation over time according to the changing audio level of detected sounds (e.g., utterances, music, and / or other noises in the environment of the computer system 101). In some embodiments, the computer system 101 displays placeholder text 934 in the text entry field 906 before receiving input to enter text into the text entry field. In some embodiments, the computer system 101 displays placeholder text 934 in the text entry field 906 in response to receiving one or more inputs corresponding to a request to delete existing text from the text entry field, such as the URL of website A currently displayed in the internet browsing user interface 902.

[0239] As shown in Figure 9I, the text entry field 906 is displayed with a background that does not change color according to the changing audio level of the detected audio (e.g., ambient noise or utterance), and is displayed without appearing to grow around the edges of the text entry field 906. In some embodiments, this appearance of the text entry field 906 shown in Figure 9I indicates that the computer system does not input text corresponding to the utterance input into the text entry field 906 in response to receiving the utterance input. For example, if the user speaks one or more words while the computer system 101 is displaying the text entry field 906 as shown in Figure 9I, the computer system 101 maintains the display of placeholder text 934 in the text entry field 90.

[0240] Figure 9I shows a dictation icon 936 contained in a text entry field 906. In some embodiments, the computer system 101 displays the dictation icon 936 in the text entry field 906 in response to detecting user attention directed to the text entry field 906, as described above. As shown in Figure 9I, the computer system 101 detects user attention 913e directed to the dictation icon 936. In response to detecting user attention directed to the dictation icon 936 in Figure 9I, the computer system updates the appearance of the text entry field 906 and inputs text corresponding to the utterance input in response to detecting the utterance input, as described with reference to at least Figure 9J.

[0241] Figure 9J shows a computer system 101 displaying a text entry field 906 with an updated appearance in response to detecting user attention 913e directed towards the dictation icon 936, as described above with reference to Figure 9I. In some embodiments, as shown in Figure 9I, the computer system 101 updates the text entry field 906 to include the dictation icon 938 in a location within the text entry field 906 that is different from the location shown in Figure 9I. In some embodiments, as shown in Figure 9J, the computer system 101 updates the text entry field 906 to be displayed with a background that changes color according to the changing audio level of the detected audio, including the voice input 916e. In some embodiments, as shown in Figure 9J, the computer system 101 updates the text entry field 906 to be displayed with a growing contour 942a that changes color, intensity, and / or radius according to the changing audio level of the detected audio, including the voice input 916e. In some embodiments, as shown in Figure 9J, the computer system 101 updates the text entry field 906 to include an insertion marker 944a that changes color and / or has a growing effect that changes color, intensity, and / or radius, according to a changing audio level of detected audio, including a voice input 916e.

[0242] In some embodiments, while displaying the text entry field 906 as shown in Figure 9J, the computer system 101 receives utterance input 916e while the user's attention 913e is directed to the text entry field 906. In some embodiments, in response to receiving utterance input 916e while displaying the text entry field 906 as shown in Figure 9J and while the user's attention 913e is directed to the text entry field 906, the computer system 101 enters text corresponding to the utterance input into the text entry field 906, as shown in Figure 9K. In some embodiments, if the user's attention 913e is not directed to the text entry field 906 while the computer system 101 detects utterance input 916e, the computer system 101 refrains from entering text corresponding to utterance input 916e into the text entry field. In some embodiments, upon detecting that the user's attention has been directed away from the text entry field 906, the computer system 101 ceases displaying the text entry field 906 in the appearance shown in Figure 9J and displays the text entry field 906 in the appearance shown in Figure 9I.

[0243] Figures 9K and 9L show that the computer system 101 inputs text corresponding to the utterance input 916e in Figure 9J, in response to the utterance input 916e described above with reference to Figure 9J. In some embodiments, the computer system 101 animates the input of the text character by character, as shown in Figures 9K and 9L. As shown in Figure 9K, while inputting text corresponding to the utterance input, the computer system 101 updates the background color of the text entry field 906, the glow effect 942b around the text entry field 906, and / or the color and / or glow of the insertion marker 912 according to the detected audio level (e.g., of the utterance input 916e). Figure 9K shows that while inputting text corresponding to the utterance input 916e, the first portion 946a of the text corresponding to the utterance input 916e is displayed in a first color, and the second portion 948a of the text corresponding to the utterance input 916e is displayed in a second color and / or glow effect. For example, when the computer system 101 displays an additional character corresponding to the utterance input 916e, the computer system 101 displays the character with a color and / or glow effect that changes according to the detected audio level, and then transitions to a first color that is solid. In some embodiments, the background color of the text entry field 906, the glow effect 942b around the text entry field 906, the color and / or glow of the insertion marker 912, and the color of the second portion 948a change in coordination according to the detected audio level.

[0244] Figure 9L shows a continuation of the text entry corresponding to the utterance input 916e shown in Figure 9J. As shown in Figure 9L, as the detected audio level continues to change (for example, as the electronic device continues to detect the utterance input 916e), the computer system 101 updates the background color of the text entry field 906, the glow effect 942c around the text entry field 906, and / or the color and / or glow of the insertion marker 912 according to the detected audio level (for example, of the utterance input 916e). As shown in Figure 9L, when the computer system 101 adds a character to the entered text, the portion of the text 946b that is displayed in solid color includes the additional character and displays the character 948b in a color and / or glow corresponding to the audio level before it was displayed in solid color when the character was added to the text entry field 906.

[0245] In some embodiments, when the computer system 101 no longer detects utterance input 916e, the computer system 101 displays the text entry field 906 having the text corresponding to the utterance input in the appearance shown in Figure 9I. For example, the computer system 101 displays the text entry field 906 with a plain background color that remains the same regardless of the detected audio level, discontinues displaying a glowing effect around the text entry field 906, and discontinues displaying the insertion marker 912 within the text entry field 906. In some embodiments, while displaying the text corresponding to the utterance input 916e in the text entry field, the computer system 101 receives input corresponding to a request to perform an internet search based on the text in the text entry field. In some embodiments, in response to the input, the computer system 101 displays search results related to the text in the text entry field (e.g., the text corresponding to the utterance input 916e).

[0246] In some embodiments, the computer system 101 enters text into the text entry field 906 in response to one or more typed text entry inputs, but the computer system 101 displays the text entry field 906 having the appearance shown in Figure 9I instead of the appearance shown in Figures 9J–9L. For example, Figures 9M–9N show an example in which the computer system 101 enters text into the text entry field 906 in response to input received using a soft keyboard 950. In some embodiments, the computer system 101 similarly enters text into the text entry field 906 in response to input received using a hardware keyboard. In some embodiments, the computer system 101 enters text into the text entry field 906 in response to input directed to the soft keyboard according to one or more steps of Method (single or multiple) 1200, 1400, 1600, and / or 2200. In some embodiments, the computer system 101 enters text into the text entry field 906 in response to input directed to a hardware keyboard according to one or more steps of Method 2400.

[0247] In Figure 9M, the computer system 101 simultaneously displays a text entry field 906 along with a soft keyboard 950. In some embodiments, the soft keyboard 950, when selected, is displayed to the computer system 101 with an option 954 that causes the computer system 101 to input text into the text entry field 906 in response to utterance input. In some embodiments, as shown in Figure 9M, the computer system 101 displays a text entry field 906 with a background color that does not change according to the detected audio level without a glow effect. In some embodiments, the text entry field 906 does not include a dictation icon. Figure 9M shows the text entry field 906 displayed without an insertion marker, but in some embodiments, the text entry field 906 includes an insertion marker. In Figure 9M, the computer system 101 receives input directed to the soft keyboard 950 with a hand 903. In response to the input shown in Figure 9M, the computer system 101 inputs text corresponding to the input directed to the soft keyboard, as shown in Figure 9N.

[0248] Figure 9N shows a computer system 101 displaying a text entry field 906 having text 952 corresponding to the input shown in Figure 9M. In some embodiments, the computer system 101 displays the text 952 in a color that does not change over time and / or according to the detected audio level when the computer system inputs the text 952 and / or after the computer system 952 has input the text. In some embodiments, while and after the text 952 is being entered into the text entry field 906, the computer system 101 displays the text entry field 906 with a background color that does not change according to the detected audio level without a glow effect. In some embodiments, while the text 952 is being displayed in the text entry field 906, the computer system 101 receives input corresponding to a request to perform an internet search based on the text 952 in the text entry field 906, and displays the search results corresponding to the text 952 in response to that input.

[0249] Further explanations regarding Figures 9A to 9N are provided below with reference to Method 1000 described with respect to Figures 9A to 9N.

[0250] Figures 10A to 10R are flowcharts of methods for entering text into a text entry field according to various embodiments. In some embodiments, method 1000 is performed in a computer system (e.g., computer system 101 in Figure 1) which includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., a head-up display, a display, a touchscreen, and / or a projector) and one or more input devices. In some embodiments, method 1000 is stored in a non-temporary (or temporary) computer-readable storage medium and controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 in Figure 1A). Some operations in method 1000 are optionally combined, and / or the order of some operations is optionally changed.

[0251] In some embodiments, such as Figure 9A, Method 1000 is performed in a computer system (e.g., 101) that communicates with a display generation component (e.g., 120) and one or more input devices. In some embodiments, the computer system is the same or similar as the computer system described above with reference to Method 800. In some embodiments, one or more input devices are the same or similar as the one or more input devices described above with reference to Method 800. In some embodiments, the display generation component is the same or similar as the display generation component described above with reference to Method 800.

[0252] In some embodiments, a computer system (e.g., 101) displays a text entry field (e.g., 906) such as Figure 9A via a display generation component (e.g., 120) (1002a). In some embodiments, the text entry field is displayed in the same or similar three-dimensional environment as described above with reference to Method 800. In some embodiments, the text entry field is a bidirectional user interface element that accepts text input. In some embodiments, the three-dimensional environment includes selectable options that, when selected, cause the computer system to perform an action on text (e.g., previously) entered in the text entry field. For example, the text entry field is a web address bar, a search box, a field that accepts a file name, a message field, or a word processor, and the selectable options are, respectively, a navigation option, a search option, a save or load option, an option to send a message, or an option to save the entered text as a document. In some embodiments, the text entry field has one or more of the features of a text entry field described later with reference to Methods 1200, 1400, and / or 1600.

[0253] In some embodiments, while displaying a text entry field, such as in Figure 9A, via a display generation component (e.g., 120) (1002b), the computer system (e.g., 101) detects a first utterance input from the user, such as in Figure 9B (e.g., 916a), via one or more input devices (e.g., a microphone) (1002c). In some embodiments, receiving the first utterance input includes detecting the user uttering a word, number, letter, and / or special character (e.g., a non-character symbol contained in written text). In some embodiments, while detecting the user's gaze directed at the text entry field and the first utterance input, the computer system does not detect any additional input (e.g., via one or more input devices other than an eye-tracking device and / or a microphone) corresponding to a request to enter text into the text entry field.

[0254] In some embodiments, while displaying a text entry field via a display generation component (e.g., 120) (1002b), in response to detecting a first utterance input from the user (e.g., 916a) as in Figure 9A (1002d), when the first utterance input from the user (e.g., 916a) is received, the computer system displays a text representation (e.g., 920) of the first utterance input in the text entry field (e.g., 906) as in Figure 9B (1002e), in accordance with the determination that the user's attention (e.g., including gaze 913a) is directed to a text entry field (e.g., 906) as in Figure 9B (e.g., the user's gaze or a proxy of the user's gaze is maintained over a threshold period before the first utterance input is detected, as will be described in more detail below). In some embodiments, the text representation of the first utterance input is a descriptive representation of the word and / or character spoken by the user. In some embodiments, before receiving a first utterance input, the computer system presents individual text within a text entry field, and displaying a font-based text representation of the first utterance input includes replacing the individual text with the text representation of the first utterance input. For example, the individual text may indicate the purpose of the text entry field (e.g., “Message” or similar text in a messaging text entry field, “Search” or “Enter search term here” in a search text entry field) or include text associated with a previous or current function of the application associated with the text entry field (e.g., the URL of a website presented in a web browser when the first utterance input is received). In some embodiments, the font-based text representation of the first utterance input within the text entry field is added to the individual text, such as adding text to a document in a word processing application.

[0255] In some embodiments, while displaying a text entry field via a display generation component (e.g., 120) (1002b), as in Figure 9A, a first utterance input from the user (e.g., 916a) is detected (1002d), and as in Figure 9C, the computer system (e.g., 101) stops displaying a text representation of the first utterance input in the text entry field (e.g., 906) (e.g., 1002f), as in Figure 9E, in response to the detection of a first utterance input from the user (e.g., 913b) when the first utterance input from the user (e.g., 916b) is received (e.g., the user's gaze is not directed at the text entry field (e.g., 902) (e.g., the user's gaze is not directed at the text entry field, or the gaze is maintained for less than a threshold period described in more detail below and before the first utterance input is detected). In some embodiments, discontinuing the display of the text representation of the first utterance input within the text entry field includes maintaining the display of individual texts displayed within the text entry field while the first utterance input is detected.

[0256] As described above, displaying a textual representation of the first utterance input in the text entry field improves user interaction with the computer system by providing additional control techniques (e.g., utterance input) without cluttering the user interface with additional displayed controls.

[0257] In some embodiments, as shown in Figure 9B, the determination that the user's attention (e.g., 913a) is directed to a text entry field (e.g., 906) when a first utterance input from the user (e.g., 916a) is received includes the determination that the user's gaze (e.g., 913a) has been directed to the text entry field (e.g., 906) for at least a time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds) (1004a). In some embodiments, the computer system uses an eye-tracking device included in one or more input devices to determine the location of the user's gaze. In some embodiments, the determination that the user's attention is not directed to a text entry field includes the determination that the user's gaze is not directed to a text entry field, or that the time the user's gaze has been directed to a text entry field is less than a time threshold. Displaying a textual representation of a first utterance in the text entry field, based on detecting the user's gaze directed at the text entry field over a time threshold, improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0258] In some embodiments, such as Figure 9A, detecting that the user's attention (e.g., 913a) is directed towards a text entry field (e.g., 906) includes detecting that the user's gaze (e.g., 913a) has been directed towards the text entry field (e.g., 906) for longer than a time threshold (e.g., 0.1, 0.2, 0, 3, 0.5, 1, 2, or 3 seconds) (1006a). In some embodiments, such as Figure 9B, while the text entry field (e.g., 906) is being displayed, in response to detecting the user's gaze (e.g., 913a) directed towards the text entry field (e.g., 906), the computer system (e.g., 101) presents an indication (e.g., 910a and / or 914) of the duration the user's gaze was directed towards the text entry field (1006b). In some embodiments, the computer system modifies the indication of the duration the user's gaze was directed towards the text entry field when the user's gaze continues to be directed towards the text entry field. In some embodiments, the computer system presents an indication in response to detecting the user's gaze directed at a text entry field over a time threshold. In some embodiments, the indication is a visual indication displayed via a display generating component. In some embodiments, the indication is an audio indication presented via one or more audio output devices communicating with the computer system. In some embodiments, the visual indication is a gradual expansion (e.g., horizontally) of the text entry field. In some embodiments, the visual indication is a progress bar. In some embodiments, the visual indication is a gradual change in the color and / or contour of the text entry field.

[0259] Presenting an indication of the duration the user's gaze was directed towards a text entry field improves user interaction with the computer system by providing the user with enhanced feedback.

[0260] In some embodiments, as shown in Figure 9B, while displaying a text entry field (e.g., 906), the computer system (e.g., 101) detects a user's gaze (e.g., 913a) directed towards the text entry field (e.g., 906) (1008a), and, according to a determination of the duration for which the user's gaze (e.g., 913a) was directed towards the text entry field (e.g., 906) (e.g., meeting or exceeding a time threshold), presents a second indication (e.g., 910a and / or 914) indicating that a first utterance input (e.g., 916a) is directed towards the text entry field (e.g., 906) (1008b). In some embodiments, presenting a second indication indicating that a first utterance input is directed towards the text entry field includes expanding the text entry field. For example, the computer system increases the width of the text entry field. In some embodiments, presenting a second indication that a first utterance input is directed to a text entry field includes initiating the display of a visual indication (e.g., an icon or image such as a microphone or speech bubble). In some embodiments, the second indication that a first utterance input is directed to a text entry field is displayed at an insertion location in the text within the text entry field where the text of the first utterance input is entered in response to the first utterance input. In some embodiments, the second indication that a first utterance input is directed to a text entry field is an audio indication presented via one or more audio output devices communicating with a computer system.

[0261] In some embodiments, as shown in Figure 9A, while displaying a text entry field (e.g., 906), the computer system (e.g., 101) detects that the user's gaze (e.g., 913a) is directed towards the text entry field (e.g., 906) (1008a), and, according to a determination that the duration for which the user's gaze (e.g., 913a) was directed towards the text entry field (e.g., 906) is less than a time threshold, it discontinues displaying the second indication (1008c). In some embodiments, the computer system maintains the display of the visual indication that the user's gaze is directed towards the text entry field, regardless of whether the user's gaze was directed towards the text entry field over a time threshold, in response to detecting that the user's gaze has been directed towards the text entry field over a time threshold. In some embodiments, the computer system discontinues displaying the visual indication that the user's gaze is directed towards the text entry field, in response to detecting that the user's gaze has been directed towards the text entry field over a time threshold.

[0262] Presenting a second indication that a first utterance input is directed to the text entry field in response to the user's gaze being directed to the text entry field over a time threshold improves user interaction with the computer system by providing the user with enhanced feedback.

[0263] In some embodiments, as shown in Figure 9B, the computer system (e.g., 101) detects a first utterance input from the user (e.g., 916a) while displaying a text entry field (e.g., 906) (1010a), and in accordance with the determination that the user's attention (e.g., 913a) is directed to the text entry field (e.g., 906), the computer system (e.g., 101) displays a text cursor (e.g., 912) within the text entry field (e.g., 906) via a display generation component (e.g., 120), and the text representation of the first utterance input (e.g., 920) is inserted into the text entry field (e.g., 906) at the location of the text cursor (e.g., 912) within the text entry field (e.g., 906), as shown in Figure 9C (1010b). In some embodiments, the computer system does not display a text cursor in the text entry field unless it detects a first utterance input from the user while the user's attention is directed to the text entry field, and until it does. In some embodiments, the text cursor is an insertion marker. In some embodiments, after displaying the text representation of a first utterance input within a text entry field, the computer system maintains the display of the text cursor at an updated location within the text entry field (e.g., the end of the text representation of the first utterance input) according to the determination that the user's attention is still directed to the text entry field. In some embodiments, the text cursor is a visual indication displayed via a display-generating component that shows the location within the text entry field where text is entered in response to an input (e.g., dictation input, soft keyboard input by methods 1200, 1400, and / or 1600, or hardware keyboard input) that corresponds to a request to enter text into the text entry field.In some embodiments, the computer system updates the position of the text cursor in the text entry field while entering individual text into the text entry field in response to an input in response to a request to enter text, in order to indicate that subsequent text entered in response to a subsequent input in response to a request to enter text into the text entry field is entered after the individual text.

[0264] In some embodiments, as shown in Figure 9B, in response to detecting a first utterance input from the user (e.g., 916a) while a text entry field (e.g., 906) is being displayed (1010a), the computer system (e.g., 101) discontinues displaying a text cursor within the text entry field (e.g., 906), such as in Figure 9A, based on the determination that the user's attention (e.g., 913b) is not directed towards the text entry field (e.g., 906), such as in Figure 9C (1010c).

[0265] Displaying a text cursor in a text entry field improves user interaction with the computer system by providing users with enhanced visual feedback.

[0266] In some embodiments, as shown in Figure 9A, detecting that the user's attention (e.g., 913a) is directed towards a text entry field (e.g., 906) includes detecting that the user's gaze (e.g., 913a) has been directed towards the text entry field (e.g., 906) for longer than a time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds) (1012a).

[0267] In some embodiments, as shown in Figure 9A, while the user's attention (e.g., 913a) is directed away from the text entry field (e.g., 906), the computer system (e.g., 101) displays the text entry field (e.g., 906) having visual properties (e.g., color, opacity, line style, and / or size) having a first value via a display generating component (e.g., 120) (1012b).

[0268] In some embodiments, as shown in Figure 9A, while a text entry field (e.g., 906) having a visual characteristic having a first value is displayed via a display generation component (e.g., 120), a computer system (e.g., 101) detects a user's gaze (e.g., 913a) directed towards a text entry field (e.g., 913a) as shown in Figure 9A (1012c) via one or more input devices (e.g., 314).

[0269] In some embodiments, as shown in Figure 9B, upon detecting a user's gaze (e.g., 913a) directed towards a text entry field (e.g., 906), the computer system (e.g., 101) gradually modifies the display of the text entry field (e.g., 906) having a visual characteristic having a first value via a display generation component (e.g., 120) according to the duration that the user's gaze (e.g., 913a) is directed towards the text entry field (e.g., 906) (1012d). In some embodiments, the value of the visual characteristic changes over time as the user's gaze remains directed towards the text entry field. For example, visual characteristics may be color, size, border, or brightness, and the computer system displays the text entry field with a first color, size, border, or brightness while the user's gaze is not directed at the text entry field, and gradually changes the color, size, border, or brightness of the text entry field while the user's gaze remains directed at the text entry field, transitioning to display the text entry field with a second color, size, border, or brightness in response to detecting the user's gaze directed at the text entry field over a time threshold.

[0270] Gradually adjusting the visual properties of a text entry field in response to detecting the user's gaze directed at the field improves user interaction with the computer system by providing the user with enhanced visual feedback.

[0271] In some embodiments, as shown in Figure 9B, while displaying a text entry field (e.g., 906) via a display generation component (e.g., 120), the computer system (e.g., 101) detects a first utterance input (e.g., 916b) from the user, based on the determination that the user's attention (e.g., 913a) is directed to the text entry field (e.g., 906) when a first utterance input (e.g., 916a) from the user is received, and in response to the detection of the first utterance input (e.g., 916b) from the user, the computer system (e.g., 101) displays the text entry field (e.g., 906) via the display generation component (e.g., 120) as shown in Figures 9B-9C, having visual characteristics (e.g., size, color, opacity, contour style, and / or visual effects such as glow or shadow) that have individual values ​​that change over time according to the changes over time of the characteristics (e.g., volume, tone, and / or frequency) of the first utterance input (e.g., 916a) (1014). In some embodiments, the visual characteristic is a glow effect displayed around the text entry field. In some embodiments, the intensity of glow (and / or other visual properties) (e.g., color darkness, brightness, saturation, thickness, and / or opacity) changes over time according to the audio level of the first speech input.

[0272] Displaying a text entry field with visual characteristics that have individual values ​​that change over time according to the characteristics of the first speech input improves user interaction with the computer system by providing the user with enhanced visual feedback.

[0273] In some embodiments, as shown in Figure 9B, after detecting a first utterance input from the user (e.g., 916a) via one or more input devices, while displaying a text entry field (e.g., 906) via a display generation component (e.g., 120) (1016a), the computer system (e.g., 101) detects a second utterance input (e.g., 916b) from the user via one or more input devices (e.g., 314) (1016b), while the user's attention (e.g., 913b) is not directed towards the text entry field (1016b). In some embodiments, the start of the second utterance input from the user is detected within a time threshold (e.g., 0.5, 1, 2, 3, 4, or 5 seconds) that detects the end of the first utterance input. For example, the computer system detects that the user has not spoken for less than the time threshold between the first and second utterance inputs. In some embodiments, the user's attention is directed to an area of ​​the three-dimensional environment other than the text entry field. In some embodiments, the user's attention is not directed to the three-dimensional environment. In some embodiments, the user's eyes are closed for a longer period than a time threshold associated with blinking (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds).

[0274] In some embodiments, as shown in Figure 9B, after detecting a first utterance input from the user (e.g., 916a) via one or more input devices, the computer system (e.g., 101) displays a text representation of the second utterance input (e.g., 920) in the text entry field (e.g., 906) via a display generation component (e.g., 120) (1016a), as shown in Figure 9C, in response to detecting a second utterance input from the user (e.g., 916b) while the user's attention (e.g., 913b) was not directed to the text entry field (e.g., 906) (1016c), as shown in Figure 9B, according to the determination that the user's attention (e.g., 913a) was directed to the text entry field (e.g., 906) when the first utterance input from the user (e.g., 916a) was received, the computer system (e.g., 101) displays a text representation of the second utterance input (e.g., 920) in the text entry field (e.g., 906) via a display generation component (e.g., 120) (1016d). In some embodiments, the computer system displays a text representation of a first utterance in a text entry field while the user is providing a second utterance. In some embodiments, the computer system displays a text representation of a second utterance in the text entry field simultaneously with the text representation of the first utterance. In some embodiments, the computer system, upon detecting that the user's attention is directed towards the text entry field, initiates a process of presenting text representations of utterances in the text entry field and continues to input text representations of additional utterances even if additional utterances are detected while the user's attention is no longer directed towards the text entry field.

[0275] In some embodiments, as shown in Figure 9B, after detecting a first utterance input from the user (e.g., 916a) via one or more input devices, and while displaying a text entry field (e.g., 906) via a display generation component (e.g., 120) (1016a), as shown in Figure 9C, in response to detecting a second utterance input from the user (e.g., 916b) while the user's attention (e.g., 913b) is not directed to the text entry field (e.g., 906) (1016c), the computer system (e.g., 101) stops displaying a text representation of the second utterance input in the text entry field (e.g., 906) via the display generation component (e.g., 120) (1016e), according to the determination that the user's attention was not directed to the text entry field (e.g., 906) when the first utterance input was received (1016e). In some embodiments, because the user's attention was not directed to the text entry field when the first utterance input was received, the computer system refrains from displaying a text representation of the first utterance input in the text entry field and displays the text entry field without a text representation of the first utterance input while the second utterance input is received (for example, regardless of where the user is looking while the computer system detects the second utterance input). In some embodiments, the computer system does not initiate the process of entering a text representation of the utterance input into the text entry field until the computer system detects, and does not detect, the user's attention directed to the text entry field.

[0276] Displaying a textual representation of a second utterance in a text entry field improves user interaction with the computer system by executing an action when a set of conditions are met, without requiring further user input.

[0277] In some embodiments, as shown in Figure 9C, while the computer system (e.g., 101) displays a text representation (e.g., 920) of a first utterance in the text entry field (e.g., 906) in response to detecting a first utterance from the user while the user's attention is directed to the text entry field (1018a), as shown in Figure 9C, the computer system (e.g., 101) receives a second utterance (e.g., 916b) via one or more input devices (e.g., 120) which is a continuation of the first utterance from the user while the user's attention (e.g., 906) is directed away from the text entry field (1018b). In some embodiments, the second utterance received while the user's attention is directed away from the text entry field is similar to the second utterance received while the user's attention is directed away from the text entry field described above.

[0278] In some embodiments, as shown in Figure 9C, while the first utterance input from the user is detected while the user's attention is directed to the text entry field, a text representation of the first utterance input (e.g., 920) is displayed in the text entry field (e.g., 906) (1018a),

[0279] As shown in Figure 9D, upon receiving a second utterance input (e.g., 916b in Figure 9C), the computer system (e.g., 101) displays a text representation of the second first utterance input (e.g., 920) via a display generation component (e.g., 120) (1018c). In some embodiments, the computer system, upon detecting user attention directed towards the text entry field while providing the first utterance input as described above, continues to input a text representation of the user utterance after inputting the text representation of the first utterance input.

[0280] Displaying a continuation of the initial utterance in a text entry field improves user interaction with the computer system by performing an action when a set of conditions are met, without requiring further user input.

[0281] In some embodiments, as shown in Figure 9C, while a text representation of a first utterance input (e.g., 920) is displayed in a text entry field (e.g., 906) via a display generation component (e.g., after detecting the first utterance input while the user's attention is directed to the text entry field), the computer system (e.g., 101) detects a second utterance input (e.g., 916b) via one or more input devices (e.g., 314) which is a continuation of the first utterance input from the user while the user's attention (e.g., 913b) is not directed to the text entry field (e.g., 906) (1020a). In some embodiments, the second utterance input from the user detected while the user's attention is not directed to the text entry field is similar to the second utterance input from the user detected while the user's attention is not directed to the text entry field as described above.

[0282] In some embodiments, as shown in Figure 9E, in response to detecting a second utterance input from the user (e.g., 916b) while the user's attention (e.g., 913b) is not directed to the text entry field (e.g., 906), the computer system (e.g., 101) stops displaying the text representation (e.g., 920) of the first utterance input in the text entry field (e.g., 906) via a display generation component (e.g., 120) (1020b). In some embodiments, in response to detecting that the user's attention has been directed away from the text entry field (e.g., detecting that the user has averted their gaze from the text entry field), the computer system deletes the text in the text entry field that was previously entered via dictation. In some embodiments, upon detecting that the user's attention has shifted away from the text entry field (for example, detecting that the user looks away from the text entry field), the computer system deletes the text in the text entry field that was entered via dictation without performing any actions associated with the text entry field (for example, searching for the search term entered in the text entry field, sending the message entered in the text entry field, and / or navigating to the website entered in the text entry field).

[0283] Stopping the display of the text representation of the first utterance in the text entry field in response to detecting a second utterance while the user's attention is diverted away from the text entry field improves user interaction with the computer system by reducing the number of inputs required to perform the action (e.g., removing the text representation of the first utterance from the text entry field).

[0284] In some embodiments, as shown in Figure 9B, upon detection of a first utterance input from the user (e.g., 916a), and in accordance with the determination that the user's attention (e.g., 913a) is directed to the text entry field (e.g., 906) when the first utterance input (e.g., 916a) is received, the computer system (e.g., 101) displays the text entry field (e.g., 906) having visual characteristics having a first value (e.g., color, size, opacity, text style such as font style, text size, and / or text highlighting, and / or border style) via a display generation component (e.g., 120) (1022a). In some embodiments, the computer system displays the text entry field having visual characteristics having a first value while detecting speech input while the user's attention is directed to the text entry field. In some embodiments, the computer system displays a highlighted text representation of the speech input in response to (e.g., and during) the detection of speech input while the user's attention is directed to the text entry field. In some embodiments, the computer system performs the display.

[0285] In some embodiments, as shown in Figure 9C, while a text entry field (e.g., 906) having a visual characteristic having a first value is displayed via a display generation component (e.g., 120), the computer system (e.g., 101) detects via one or more input devices (e.g., 314) that the user's attention (e.g., 913b) is not directed towards the text entry field (e.g., 906) (1022b). In some embodiments, the user's attention is directed towards an area of ​​the three-dimensional environment other than the text entry field. In some embodiments, the user's attention is directed away from the three-dimensional environment (e.g., away from the display generation component). In some embodiments, the user closes their eyes for longer than a time threshold associated with blinking (e.g., 0.5, 1, 2, 3, or 5 seconds).

[0286] In some embodiments, as shown in Figure 9A, in response to detecting that the user's attention is not directed at the text entry field (e.g., 906), the computer system (e.g., 101) displays the text entry field (e.g., 906) via a display generation component (e.g., 120) that has a visual characteristic having a distinct value that changes over time until it reaches a second value different from a first value (1022c). In some embodiments, the value of the visual characteristic changes gradually over time until it reaches a second value in response to detecting that the user's attention is not directed at the text entry field. For example, highlighting on the text contained in the text entry field gradually fades out. Transitioning to displaying a text entry field with a visual characteristic having a second value in response to detecting that the user's attention is not directed at the text entry field improves user interaction with the computer system by providing the user with enhanced visual feedback.

[0287] In some embodiments, as shown in Figure 9D, while displaying a text representation (e.g., 920) of a first utterance input within a text entry field (e.g., 906), the computer system (e.g., 101) detects a second utterance input (e.g., 916c) via one or more input devices (e.g., 314) (1024b).

[0288] In some embodiments, as shown in Figure 9D, while a text representation of a first utterance input (e.g., 920) is displayed in a text entry field (e.g., 906), the computer system (e.g., 101) detects a first utterance input from the user (1024a), as shown in Figure 9D, a second utterance input (e.g., 916c), and, according to the determination that the second utterance input (e.g., 916c) corresponds to a request to perform an action on the text representation of the first utterance input (e.g., 920) in the text entry field (e.g., 906), and that one or more criteria are met (e.g., including criteria that are met when the user's attention is directed to the text entry field), the computer system (e.g., 101) performs an action on the text representation of the first utterance input (e.g., 920) in the text entry field (e.g., 906) (1024d). In some embodiments, the second utterance input is a predetermined utterance associated with the action, or includes one such utterance. For example, a text entry field is a message configuration field, where the second utterance input is "Send" or "Send it," and the action is to send a message containing the text expression of the first utterance input. Another example is a text entry field that is a search field, where the second utterance input is "Search" or "Proceed," and the action is to perform a search that includes the text expression of the first utterance input as the search term.

[0289] In some embodiments, as shown in Figure 9D, while a text representation of a first utterance input (e.g., 920) is displayed in a text entry field (e.g., 906), the computer system (e.g., 101) detects a first utterance input from the user (1024a) or a second utterance input (e.g., 916c), and determines that the second utterance input does not correspond to a request to perform an action on the text representation of the first utterance input (e.g., 920) in the text entry field (e.g., 906), or that one or more criteria are not met, then refrains from performing an action on the text representation of the first utterance input (e.g., 920) in the text entry field (e.g., 906) (1024e). In some embodiments, the second utterance input does not contain a predetermined utterance associated with an action. In some embodiments, the computer displays a text representation of a second utterance in a text entry field in response to a second utterance that does not correspond to a request to perform an action (for example, instead of or in addition to the text representation of a first utterance).

[0290] Performing an action on the textual representation of the first utterance in a text entry field in response to a second utterance improves user interaction with the computer system by providing additional controls without cluttering the user interface with additional displayed controls.

[0291] In some embodiments, such as Figure 9D, the determination that a second utterance input (e.g., 916c) corresponds to a request to perform an action on the text representation (e.g., 920) of the first utterance input within the text entry field (e.g., 906), based on the determination that the text entry field (e.g., 906) is a first type of text entry field, is based on one or more first criteria (1026a). In some embodiments, one or more first criteria include criteria that are met when the second utterance input contains the first utterance. For example, if the text entry field is a search field, one or more first criteria include criteria that are met when the second utterance input contains "search," "proceed," etc. In some embodiments, the computer system refrains from performing an action on the text representation of the first utterance input within the text entry field, based on the determination that the second utterance input corresponds to an action associated with a second type of text entry field that is different from the first type of text entry field.

[0292] In some embodiments, the determination that a second utterance input (e.g., 916c) corresponds to a request to perform an action on the text representation (e.g., 920) of the first utterance input in the text entry field is based on one or more second criteria different from one or more first criteria (1026b), according to the determination that the text entry field is a second type of text entry field different from a first type of text entry field (e.g., a text entry field different from text entry field 906 in Figure 9D). In some embodiments, the one or more second criteria include criteria that are satisfied when the second utterance input contains a second utterance. For example, if the text entry field is a messaging field, the one or more first criteria include criteria that are satisfied when the second utterance input contains “send,” “send it,” etc. In some embodiments, according to the determination that a second utterance input corresponds to an action associated with a first type of text entry field different from a second type of text entry field, the computer system refrains from performing an action on the text representation of the first utterance input in the text entry field.

[0293] Evaluating a second utterance input according to different criteria depending on the type of text entry field improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0294] In some embodiments, one or more criteria include a criterion that is met when the user's gaze (e.g., 913c) is directed towards the text entry field (e.g., 906) while the computer system (e.g., 101) is detecting a second utterance input (e.g., 916c) (e.g., for at least a time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds)) (1028). In some embodiments, according to the determination that the user's gaze is not directed towards the text entry field while the computer system is detecting the second input, the computer system refrains from performing an action on the text representation of the first utterance input in the text entry field, regardless of whether the second utterance input meets one or more additional criteria for determining that the second utterance input corresponds to a request to perform an action on the text representation of the first utterance input in the text entry field, such as the first utterance input containing a default utterance associated with the action.

[0295] Determining that the second utterance corresponds to a request to perform an action on the text representation of the first utterance in the text entry field, based on the user's gaze being directed towards the text entry field while the second utterance is being detected, improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.

[0296] In some embodiments, such as Figure 9A, the computer system (e.g., 101) displays individual text within a text entry field (e.g., 906) via a display generation component (e.g., 120) (1030a), as shown in Figure 9A, before detecting a first utterance input from the user (e.g., 916a), as shown in Figure 9B. In some embodiments, the individual text is previously entered in response to a second utterance input similar to the first utterance input described above, under the same or similar conditions as described above. In some embodiments, the individual text is previously entered by the user via a different input modality, such as using a soft keyboard or a hardware keyboard, according to one or more of the methods 1200, 1400, or 1600 described below. In some embodiments, the individual text is placeholder text that is automatically displayed by the computer system without receiving input in response to a request to enter placeholder text into the text entry field.

[0297] In some embodiments, as shown in Figure 9B, upon detection of a first utterance input from the user (e.g., 916a), and in accordance with the determination that the user's attention (e.g., 913a) is directed to a text entry field (e.g., 906), the computer system (e.g., 101), via a display generation component (e.g., 120), suspends the display of individual text in the text entry field (e.g., 906) and displays a text representation (e.g., 920) of the first utterance input in the text entry field (e.g., 906) (1030b). In some embodiments, the computer system replaces the individual text with a text representation of the first utterance input in the text entry field upon detection of a first utterance input from the user while the user's attention is directed to the text entry field. In some embodiments, the computer system suspends the display of individual text in the text entry field in response to a first utterance input without detecting additional input corresponding to a request to suspend the display of individual text in the text entry field. Stopping the display of individual text and displaying a textual representation of the first utterance in the text entry field upon detection of a first utterance improves user interaction with the computer system by performing an action when a set of conditions are met, without requiring further user input.

[0298] In some embodiments, as shown in Figure 9F, before detecting a first utterance from the user, the computer system (e.g., 101) displays individual text and a cursor (e.g., 928) at a first location in a text entry field (e.g., 926) via a display generation component (e.g., 120) (1032a). In some embodiments, the computer system displays the cursor in accordance with the determination that it is possible to edit the individual text. In some embodiments, the computer system displays the cursor in accordance with the determination that it is possible to edit the individual text in a way other than replacing the entire individual text (e.g., adding text or deleting part of the individual text without detecting the entire individual text). In some embodiments, in accordance with the determination that it is not possible to edit the individual text, the computer system (optionally, depending on whether the first utterance has been detected while the user's attention is directed to the text entry field) discontinues displaying the cursor before detecting the first utterance.

[0299] In some embodiments, as shown in Figure 9G, upon detection of a first utterance input from the user (e.g., 916d), and in accordance with the determination that the user's attention (e.g., 913d) is directed to a text entry field (e.g., 926) (1032a), the computer system (e.g., 101) maintains the display of individual text within the text entry field (e.g., 926) via a display generation component (e.g., 120) (1032b). In some embodiments, upon detection of a first utterance input while the user's attention is directed to the text entry field, and in accordance with the determination that it is not possible to edit the individual text, the computer system discontinues the display of individual text within the text entry field and displays a cursor or visual indication, which is described in more detail below.

[0300] In some embodiments, as shown in Figure 9G, upon detecting a first utterance input from the user (e.g., 916d), and in accordance with the determination that the user's attention (e.g., 913d) is directed to a text entry field (e.g., 926) (1032a), the computer system (e.g., 101) stops displaying the cursor in the text entry field (e.g., 926) via a display generation component (e.g., 120) (1032d).

[0301] In some embodiments, as shown in Figure 9G, upon detecting a first utterance input from the user (e.g., 916d), and in accordance with the determination that the user's attention (e.g., 913d) is directed to a text entry field (e.g., 926) (1032a), the computer system (e.g., 101) displays a visual indication (e.g., 930) at a second location within the text entry field (e.g., 926) (e.g., the same as or different from the first location) via a display generation component (e.g., 120), and the textual representation of the first utterance input is added to the individual text at the second location within the text entry field (e.g., 926) (1032e). In some embodiments, the visual indication is different from a cursor. In some embodiments, the visual indication is the same as a cursor. In some embodiments, after entering a text representation of a first utterance into a text entry field, the computer system displays a visual indication adjacent to (e.g., immediately after) (e.g., after) the text representation of the first utterance. In some embodiments, the visual indication is a microphone or speech bubble or an image of a person speaking. In some embodiments, in response to detecting that the user's attention has been directed away from the text entry field without detecting any continuation of the first utterance, the computer system ceases displaying the visual indication and begins displaying a cursor (e.g., at the location in the text entry field corresponding to the text representation of the first utterance).

[0302] Displaying a visual indication at the location within the text entry field where a text representation of the first utterance is added, in response to the detection of the first utterance from the user, improves user interaction with the computer system by providing the user with enhanced visual feedback.

[0303] In some embodiments, as shown in Figures 9G to 9H, a second location where a textual representation of an utterance is added to individual text is adjacent to (e.g., nearby or adjacent to) the first portion of text (1034a) according to the determination that the user's gaze (e.g., 913d) is directed to a first portion of text within a text entry field (e.g., 926) while a first utterance input (e.g., 916d) from the user is detected.

[0304] In some embodiments, while a first utterance input from the user is detected, the second location to which the textual representation of the utterance is added to the individual text is adjacent to (e.g., nearby or adjacent to) the second portion of the text in a different text entry field (e.g., 926) than the first location in the text entry field (e.g., the user's gaze 913d in Figure 9G is at a location other than the location shown in Figure 9G) (1034b). In some embodiments, the computer system displays a visual indication at the location in the text entry field that the user is looking at while the user's attention is directed to the text entry field. In some embodiments, the computer system updates the position of the visual indication in accordance with the user's gaze moving from one location in the text entry field to another location in the text entry field before the user provides the first utterance input. In some embodiments, once a computer system displays a visual indication at a second location, the computer system maintains the display of the visual indication at the second location even if the user's gaze moves away from the second location, until one or more criteria are met (for example, the user diverts their attention from a text entry field, or the user provides input to a user interface element other than a text entry field, or the user provides input to stop typing text into a text entry field based on a first utterance input).

[0305] Displaying visual indicators and allowing text input based on the user's gaze location improves user interaction with the computer system by providing additional control options without cluttering the user interface with extra displayed controls.

[0306] In some embodiments, such as Figure 9C, upon receiving a first utterance input from a user, the computer system (e.g., 101) detects a second utterance input (e.g., 916b) via one or more input devices (e.g., 314) that is not directed to the text entry field (e.g., 906), in response to detecting a first utterance input from a user, the computer system (e.g., 101) detects a second utterance input (e.g., 916b) that is a continuation of the first utterance input from the user via one or more input devices, as in Figure 9C, while detecting a second utterance input (e.g., 916b) via one or more input devices (e.g., 314) that is not directed to the text entry field (e.g., 906) (e.g., 913b). In some embodiments, a second utterance, which is a continuation of a first utterance detected while the user's attention is not directed to the text entry field, is similar to the second utterance, which is a continuation of a first utterance detected while the user's attention is not directed to the text entry field, as described in more detail above.

[0307] In some embodiments, such as Figure 9C, when a first utterance input from a user is received, the computer system (e.g., 101) detects the first utterance input from the user and, in accordance with the determination that the user's attention is directed to a text entry field (e.g., 906), displays a text representation (e.g., 920) of the first utterance input within the text entry field (e.g., 906) via a display generation component (e.g., 120) (1036a). In other embodiments, such as Figure 9G, the computer system (e.g., 101) detects the user's attention (e.g., 913b) that is not directed to a text entry field (e.g., 926), and, in accordance with the determination that the text entry field (e.g., 906) is a first type of text entry field, displays a continuation of the text representation (e.g., 920) of the first utterance input (e.g., 920) within the text entry field (e.g., 906) via a display generation component (e.g., 120), as shown in Figure 9D (1036d). In some embodiments, the computer system displays a text representation of a second utterance input while simultaneously maintaining the display of the text representation of the first utterance input in the text entry field. In some embodiments, the computer system discontinues the display of the text representation of the first utterance input in the text entry field and replaces it with the text representation of the second utterance input in the text entry field. In some embodiments, the first type of text entry field is a long text entry field where the computer system requires input, in addition to detecting the user's gaze directed towards the text entry field to begin dictation, such as a memo field, a word processing application field, or an email configuration field. In some embodiments, in addition to detecting the user's gaze directed towards the text entry field, the input is a selection of user interface elements associated with the dictation input, individual gestures performed using a part of the user's body, and / or individual utterance input (e.g., "Hi, voice assistant, start dictation").

[0308] In some embodiments, such as Figure 9C, upon receiving a first utterance input from the user, the computer system (e.g., 101) detects the first utterance input from the user and, in accordance with the determination that the user's attention is directed to the text entry field (e.g., 906), displays a text representation (e.g., 920) of the first utterance input in the text entry field (e.g., 906) via a display generation component (e.g., 120), as shown in Figure 9E (1036a). Upon detecting the user's attention not directed to the text entry field (e.g., 906) (e.g., 913b), the computer system (e.g., 101) stops displaying the text representati...

Claims

1. It is a method, In a computer system that communicates with a display generation component and one or more input devices, To display a user interface including scrollable content via the aforementioned display generation component, The detection of the user's gaze directed towards the scrollable content via one or more input devices, In response to detecting the user's gaze directed towards the scrollable content, In accordance with the determination that the user's gaze is directed towards the first area of ​​the scrollable content, the display of the scrollable content is maintained without scrolling the scrollable content. The user's gaze is directed to a second area of ​​the scrollable content that is different from the first area, and the scrollable content is scrolled according to the user's gaze, in accordance with the determination that the individual parts of the user meet the respective criteria. A method comprising: maintaining the display of the scrollable content without scrolling it, in accordance with the determination that the user's gaze is directed toward the second area and that the individual parts of the user do not meet the respective criteria.

2. The method according to claim 1, wherein each of the aforementioned criteria includes a criterion that is satisfied when the individual part of the user is not detected in a default pose.

3. While the user interface including the scrollable content is displayed, The detection of input directed to individual user interface elements via one or more input devices, wherein the detection of input includes detecting the user's gaze directed to the individual user interface elements and detecting that the user is performing individual gestures in individual parts of the user. The method of claim 1 or 2, further comprising: detecting the input directed to the individual user interface element and performing an action associated with the individual user interface element in response to the detection of the input directed to the individual user interface element.

4. The method according to any one of claims 1 to 3, wherein the second region of the scrollable content includes the edge of the scrollable content.

5. The computer system scrolls the scrollable content in a first direction according to the determination that the user's gaze is directed towards the second area, and the method While the user interface including the scrollable content is being displayed via the display generation component, In response to detecting the user's gaze directed towards the scrollable content, The method according to any one of claims 1 to 4, wherein the user's line of sight is directed to a third region of the scrollable content, the third region being different from the second region and also different from the first region, and the scrollable content is scrolled in a second direction different from the first direction in accordance with the user's line of sight, in accordance with the determination that the individual parts of the user satisfy the respective criteria, the second region and the third region being of different sizes.

6. The second region of the scrollable content is located at the bottom of the scrollable content and has a first size. The method according to claim 5, wherein the third region of the scrollable content is located at the top of the scrollable content and has a second size smaller than the first size.

7. Scrolling the scrollable content according to the user's gaze is, In accordance with the determination that the user's gaze is directed to a location at a first distance from an individual position of the scrollable content, the scrollable content is scrolled at a first speed corresponding to the user's gaze. The method according to any one of claims 1 to 6, comprising: scrolling the scrollable content at a second speed different from the first speed in response to the user's gaze, based on the determination that the user's gaze is directed to a location at a second distance from the individual position of the scrollable content which is different from the first distance.

8. The user's gaze is directed toward the second area of ​​the scrollable content, and while the individual parts of the user satisfy the respective criteria, the user's gaze is being scrolled according to the user's gaze, and the user's gaze is being detected, via one or more input devices, to move away from the second area of ​​the scrollable content. The method according to any one of claims 1 to 7, further comprising detecting the user's gaze directed away from the second area of ​​the scrollable content, and reducing the speed at which the scrollable content is scrolling until the scrolling of the scrollable content is stopped.

9. In response to detecting the user's gaze directed towards the scrollable content, and in accordance with the determination that the user's gaze is directed towards the second area and that the individual parts of the user satisfy the respective criteria, scrolling the scrollable content according to the user's gaze is: The method according to any one of claims 1 to 8, comprising gradually increasing the speed at which the scrollable content is scrolled while the user's gaze is directed toward the second area and the individual parts of the user satisfy the respective criteria.

10. While the user interface including the scrollable content is displayed, The individual parts of the user perform individual gestures, including the movement of the user's hand, while the user's hand is in a pinch-hand shape, via one or more input devices, and the individual parts of the user detect that the respective criteria are not met while performing the individual gestures. The method according to any one of claims 1 to 9, further comprising: scrolling the scrollable content in accordance with the movement of the user's hand (e.g., an air gesture, touch input, or other manual input) in accordance with a determination that one or more criteria are met in response to detection that the individual part of the user has performed the individual gesture.

11. The movement of the individual parts of the user has individual dimensions. In accordance with the determination that the movement of the individual part of the user is in a first direction, the computer system, in response to detecting that the individual part of the user has performed the individual gesture, scrolls the scrollable content in a second direction by a first amount. The method according to claim 10, wherein, in accordance with the determination that the movement of the individual part of the user is in a third direction different from the first direction, the computer system, in response to detecting that the individual part of the user has performed the individual gesture, scrolls the scrollable content in a fourth direction by a second amount different from the first amount, the fourth direction being different from the second direction.

12. The user's hand movement (e.g., air gesture, touch input, or other manual input) includes the movement of the hand from a first location to a second location (e.g., air gesture, touch input, or other manual input), and the user's hand maintains the pinch-hand shape while moving from the first location to the second location, and the scrollable content is scrolled in response to the detection that the individual part of the user has performed the individual gesture. The scrollable content is scrolled at a first speed according to the determination that the distance between the first location and the second location is a first distance, The method according to claim 10 or 11, comprising scrolling the scrollable content at a second speed greater than the first speed, according to the determination that the distance between the first location and the second location is a second distance greater than the first distance.

13. The one or more criteria mentioned above include a criterion that is met when the user's hand moves by at least a threshold amount while maintaining the pinch hand shape, and the method is The method according to any one of claims 10 to 12, further comprising detecting that the individual part of the user has performed the individual gesture, and determining that the movement of the user's hand (e.g., an air gesture, touch input, or other manual input) does not meet one or more of the criteria, maintaining the display of the scrollable content without scrolling the scrollable content.

14. If the speed of the user's hand movement (e.g., air gesture, touch input, or other manual input) is greater than a threshold speed and the direction of the user's hand movement (e.g., air gesture, touch input, or other manual input) is downward, then one or more of the criteria are not met, and the method is The method according to any one of claims 10 to 13, further comprising, in response to detecting that the individual part of the user has performed the individual gesture, maintaining the display of the scrollable content without scrolling it, in accordance with the determination that one or more criteria are not met.

15. In response to detecting the user's gaze directed towards the scrollable content, and in accordance with the determination that the user's gaze is directed towards the second area of ​​the scrollable content and that the individual parts of the user satisfy the respective criteria, the computer system scrolls the scrollable content in the first direction according to the user's gaze, and the method While the user interface including the scrollable content is being displayed via the display generation component, In response to detecting the user's gaze directed towards the scrollable content, The method according to any one of claims 1 to 14, further comprising the user's gaze being directed to a third area of ​​the scrollable content, the third area being different from the second area, and the scrollable content being scrolled in a second direction opposite to the first direction in accordance with the user's gaze, in accordance with the determination that the individual parts of the user meet the respective criteria.

16. In response to the detection of the user's gaze directed towards the scrollable content, scrolling the scrollable content in the first direction according to the user's gaze, in accordance with the determination that the user's gaze is directed towards the second region of the scrollable content and that the individual parts of the user satisfy the respective criteria, includes scrolling the scrollable content at a first acceleration. The method according to claim 15, wherein, in response to the detection of the user's gaze directed toward the scrollable content, the user's gaze is directed toward the third region of the scrollable content, and, in accordance with the determination that the individual portion of the user satisfies the respective criteria, scrolling the scrollable content in a second direction according to the user's gaze includes scrolling the scrollable content with a second acceleration different from the first acceleration.

17. The scrollable content includes text content and other content, and the method is While the text content of the scrollable content is being displayed, without displaying any other content of the scrollable content, The movement of the user's gaze is detected via one or more input devices. In response to detecting the movement of the user's gaze, Scrolling the text content in accordance with a determination that the movement of the user's gaze satisfies one or more criteria, including criteria that are met based on the movement of the user's gaze to lines of text within the text content. The method according to any one of claims 1 to 16, further comprising: maintaining the display of the text content without scrolling it, in accordance with a determination that the movement of the user's gaze does not satisfy one or more of the criteria.

18. The method according to claim 17, wherein scrolling the text content in response to the detection of the movement of the user's gaze that satisfies one or more of the aforementioned criteria is independent of whether the individual parts of the user are detected in a predetermined pose.

19. While the text content of the scrollable content is displayed without the other content of the scrollable content, The detection of the user's gaze directed towards the text content via one or more input devices, In response to detecting the user's gaze directed towards the text content, In accordance with the determination that the user's gaze is directed to a first area of ​​the text content and that the movement of the user's gaze does not satisfy one or more of the criteria, the display of the text content is maintained without scrolling the text content. The user's gaze is directed to a second area of ​​the text content that is different from the first area of ​​the text content, and the individual parts of the user satisfy the respective criteria, and the movement of the user's gaze does not satisfy one or more of the criteria, and the text content is scrolled according to the user's gaze. The method according to claim 17 or 18, further comprising: scrolling the text content in accordance with a determination that the user's gaze is directed to a first area of ​​the text content and that the movement of the user's gaze satisfies one or more of the criteria.

20. While the scrollable content is being displayed, in response to detecting the user's gaze directed towards the scrollable content, The method according to any one of claims 1 to 19, further comprising displaying the definition of the word contained in the scrollable content via the display generating component, based on the determination that the user's gaze has been directed to the word contained in the first region of the scrollable content for at least a threshold time.

21. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 20.

22. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 20.

23. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 1 to 20.

24. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are A user interface including scrollable content is displayed via the aforementioned display generation component. The user's gaze directed towards the scrollable content is detected via one or more input devices. In response to detecting the user's gaze directed towards the scrollable content, In accordance with the determination that the user's gaze is directed towards the first area of ​​the scrollable content, the display of the scrollable content is maintained without scrolling the scrollable content. The user's gaze is directed to a second area of ​​the scrollable content that is different from the first area, and according to the determination that the individual parts of the user meet the respective criteria, the scrollable content is scrolled according to the user's gaze. A non-temporary computer-readable storage medium including instructions for maintaining the display of the scrollable content without scrolling it, in accordance with the determination that the user's gaze is directed toward the second area and that the individual parts of the user do not meet the respective criteria.

25. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are A user interface including scrollable content is displayed via the aforementioned display generation component. The user's gaze directed towards the scrollable content is detected via one or more input devices. In response to detecting the user's gaze directed towards the scrollable content, In accordance with the determination that the user's gaze is directed towards the first area of ​​the scrollable content, the display of the scrollable content is maintained without scrolling the scrollable content. The user's gaze is directed to a second area of ​​the scrollable content that is different from the first area, and according to the determination that the individual parts of the user meet the respective criteria, the scrollable content is scrolled according to the user's gaze. A computer system including an instruction to maintain the display of the scrollable content without scrolling it, in accordance with the determination that the user's gaze is directed toward the second area and that the individual parts of the user do not meet the respective criteria.

26. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A means for displaying a user interface including scrollable content via the aforementioned display generation component, Means for detecting the user's gaze directed towards the scrollable content via one or more input devices, In response to detecting the user's gaze directed towards the scrollable content, In accordance with the determination that the user's gaze is directed towards the first area of ​​the scrollable content, the display of the scrollable content is maintained without scrolling the scrollable content. The user's gaze is directed to a second area of ​​the scrollable content that is different from the first area, and according to the determination that the individual parts of the user meet the respective criteria, the scrollable content is scrolled according to the user's gaze. A computer system comprising: means for maintaining the display of the scrollable content without scrolling the scrollable content, in accordance with the determination that the user's gaze is directed toward the second area and that the individual parts of the user do not meet the respective criteria.

27. It is a method, In a computer system that communicates with a display generation component and one or more input devices, To display a text entry field via the aforementioned display generation component, While the text entry field is being displayed via the display generation component, To detect a first utterance input from a user of the computer system via one or more input devices, In response to detecting the first utterance input from the user, When the first utterance input from the user is received, and it is determined that the user's attention is directed to the text entry field, the text representation of the first utterance input is displayed in the text entry field via the display generation component. A method comprising: stopping the display of the text representation of the first utterance input in the text entry field in accordance with the determination that the user's attention is not directed to the text entry field when the first utterance input from the user is received.

28. The method of claim 27, wherein the determination that the user's attention is directed to the text entry field when the first utterance input from the user is received includes the determination that the user's gaze is directed to the text entry field for at least a time threshold.

29. Detecting that the user's attention is directed to the text entry field includes detecting that the user's gaze is directed to the text entry field for a longer period than a time threshold, and the method is The method according to claim 27 or 28, further comprising, in response to detecting the user's gaze directed at the text entry field while the text entry field is being displayed, displaying an indication of the duration for which the user's gaze was directed at the text entry field.

30. While the text entry field is being displayed, in response to the detection of the user's gaze directed towards the text entry field, In accordance with the determination that the duration for which the user's gaze was directed towards the text entry field exceeds the time threshold, a second indication is presented indicating that a first utterance input is directed towards the text entry field. The method of claim 29, further comprising: discontinuing to present the second indication in accordance with the determination that the duration for which the user's gaze was directed toward the text entry field is less than the time threshold.

31. While the text entry field is being displayed, in response to detecting the first utterance input from the user, In accordance with the determination that the user's attention is directed to the text entry field, a text cursor is displayed in the text entry field via the display generation component, wherein the text representation of the first utterance input is inserted into the text entry field at the location of the text cursor within the text entry field. The method according to any one of claims 27 to 30, further comprising: discontinuing to display the text cursor in the text entry field in accordance with the determination that the user's attention is not directed to the text entry field.

32. Detecting that the user's attention is directed to the text entry field includes detecting that the user's gaze is directed to the text entry field for a longer period than a time threshold, and the method is While the user's attention is directed away from the text entry field, the display generation component displays the text entry field having a visual characteristic having a first value, While the text entry field having the visual characteristics having the first value is displayed via the display generation component, the user's gaze directed towards the text entry field is detected via one or more input devices. In response to detecting the user's gaze directed towards the text entry field, The method according to any one of claims 27 to 31, further comprising: gradually modifying, via the display generation component, the display of the text entry field having the visual characteristics having the first value, to the display of the text entry field having the visual characteristics having a second value different from the first value, according to the duration for which the user's gaze is directed toward the text entry field.

33. The method according to any one of claims 27 to 32, further comprising, while the text entry field is being displayed via the display generation component, detecting the first utterance input from the user, and in accordance with the determination that the user's attention is directed to the text entry field when the first utterance input from the user is received, displaying the text entry field via the display generation component, which has visual characteristics having individual values ​​that change over time in accordance with the changes over time in the characteristics of the first utterance input.

34. After detecting the first utterance input from the user via one or more input devices, while displaying the text entry field via the display generation component, To detect a second utterance input from the user, which is a continuation of the first utterance input, via one or more input devices, while the user's attention is not directed towards the text entry field. In response to detecting the second utterance input from the user while the user's attention is not directed towards the text entry field, In accordance with the determination that the user's attention was directed to the text entry field when the first utterance input from the user was received, the text representation of the second utterance input is displayed in the text entry field via the display generation component. The method according to any one of claims 27 to 33, further comprising: in accordance with the determination that the user's attention was not directed to the text entry field when the first utterance input was received, the display generation component refrains from displaying the text representation of the second utterance input in the text entry field.

35. In response to detecting the first utterance input from the user while the user's attention is directed to the text entry field, while the text representation of the first utterance input is displayed in the text entry field, Receiving a second utterance input from the user, which is a continuation of the first utterance input, via one or more input devices, while the user's attention is directed away from the text entry field. The method according to any one of claims 27 to 34, further comprising displaying a text representation of the second utterance input via the display generation component in response to receiving the second utterance input.

36. While the text representation of the first utterance input is displayed in the text entry field via the display generation component, a second utterance input, which is a continuation of the first utterance input from the user while the user's attention is not directed towards the text entry field, is detected via one or more input devices. The method according to any one of claims 27 to 34, further comprising: detecting the second utterance input from the user while the user's attention is not directed to the text entry field, and ceasing to display the text representation of the first utterance input in the text entry field via the display generation component.

37. In response to detecting the first utterance input from the user, and in accordance with the determination that the user's attention is directed to the text entry field when the first utterance input is received, the display generation component shall display the text entry field having visual characteristics having a first value, While the text entry field having the visual characteristics having the first value is being displayed via the display generation component, it is detected via one or more input devices that the user's attention is not directed towards the text entry field. The method according to any one of claims 27 to 34 or 36, further comprising: detecting that the user's attention is not directed toward the text entry field, displaying the text entry field having the visual characteristics having individual values ​​that change over time until a second value different from the first value is reached, via the display generating component.

38. In response to detecting the first utterance input from the user, while the text representation of the first utterance input is displayed in the text entry field, To detect a second speech input via one or more of the aforementioned input devices, In response to the detection of the second speech input, The second utterance input corresponds to a request to perform an action on the text expression of the first utterance input in the text entry field, and the action on the text expression of the first utterance input in the text entry field is performed according to the determination that one or more criteria are met, The method according to any one of claims 27 to 37, further comprising: the determination that the second utterance input does not correspond to the request to perform the action on the text representation of the first utterance input in the text entry field, or that one or more of the criteria are not met, and that the action on the text representation of the first utterance input in the text entry field is not performed.

39. The determination that the second utterance input corresponds to the request to perform the action on the text representation of the first utterance input in the text entry field, in accordance with the determination that the text entry field is a first type of text entry field, is based on one or more first criteria. The method of claim 38, wherein the determination that the second utterance input corresponds to the request to perform the action on the text representation of the first utterance input in the text entry field, based on the determination that the text entry field is a second type of text entry field different from the first type of text entry field, is based on one or more second criteria different from the one or more first criteria.

40. The method according to claim 38 or 39, wherein the one or more criteria include criteria that are met when the user's gaze is directed towards the text entry field while the computer system is detecting the second utterance input.

41. Before detecting the first utterance input from the user, the individual text is displayed in the text entry field via the display generation component, The method according to any one of claims 27 to 40, further comprising: detecting the first utterance input from the user, and, in accordance with the determination that the user's attention is directed to the text entry field, ceasing to display the individual text in the text entry field via the display generation component and displaying the text representation of the first utterance input in the text entry field.

42. Before detecting the first utterance input from the user, the individual text and cursor are displayed at a first location within the text entry field via the display generation component, In response to the detection of the first utterance input from the user, and in accordance with the determination that the user's attention is directed to the text entry field, The display generation component maintains the display of the individual texts in the text entry field, The display generation component shall be used to discontinue the display of the cursor in the text entry field, The method according to any one of claims 27 to 41, further comprising displaying a visual indication at a second location in the text entry field via the display generation component, wherein the text representation of the first utterance input is added to the individual text at the second location in the text entry field.

43. In accordance with the determination that the user's gaze is directed towards the first portion of the text in the text entry field while the first utterance input from the user is detected, the second location where the text representation of the utterance is added to the individual text is close to the first portion of the text, The method of claim 42, wherein, while the first utterance input from the user is detected, the determination is made that the user's gaze is directed to a second portion of the text in the text entry field, which is different from the first location in the text entry field, the second location where the text representation of the utterance is added to the individual text is close to the second portion of the text.

44. In accordance with the determination that the user's attention is directed to the text entry field when the first utterance input from the user is received, and in response to detecting the first utterance input from the user, the display generation component displays the text representation of the first utterance input in the text entry field, While detecting a second utterance input, which is a continuation of the first utterance input from the user, via one or more input devices, it is detected via one or more input devices that the user's attention is not directed towards the text entry field. In response to detecting that the user's attention is not directed towards the text entry field, In accordance with the determination that the text entry field is a first type of text entry field, the continuation of the text representation of the first utterance input is displayed in the text entry field via the display generation component, The method according to any one of claims 27 to 43, further comprising: determining that the text entry field is a second type of text entry field different from the first type of text entry field; and, via the display generation component, deactivating the display of the text representation of the second utterance input in the text entry field.

45. In response to detecting that the user's attention is not directed towards the text entry field, The method according to claim 44, further comprising: maintaining the display of the text representation of the first utterance input in the text entry field in accordance with the determination that the text entry field is a first type of text entry field; and discontinuing the display of the text representation of the first utterance input in the text entry field in accordance with the determination that the text entry field is a second type of text entry field.

46. The method according to claim 44 or 45, in accordance with the determination that the text entry field is a text entry field of the second type, the computer system, upon receiving the first utterance input from the user, regardless of whether it detects a separate text entry input different from the first utterance input via the one or more input devices before detecting the first utterance input, the computer system, upon detecting the first utterance input from the user in accordance with the determination that the user's attention is directed to the text entry field, displays the text representation of the first utterance input in the text entry field via the display generation component.

47. Displaying the text representation of the first utterance input within the text entry field via the display generation component in accordance with the determination that the text entry field is a text entry field of the first type, in response to detecting a separate text entry input different from the first utterance input via one or more input devices before detecting the first utterance input, the method The method according to any one of claims 44 to 46, further comprising, in response to the detection of the first utterance input from the user, in accordance with the determination that the text entry field is a text entry field of the first type, and in accordance with the determination that the individual text entry inputs were not detected before the detection of the first utterance input from the user, the display generation component, in which case the display of the text representation of the first utterance input in the text entry field is discontinued.

48. Displaying the text representation of the utterance input includes displaying the text representation of the utterance input in a first appearance, and the method is Receiving typed text entry input directed to the text entry field via one or more input devices, The method according to any one of claims 27 to 47, further comprising, in response to receiving the typed text entry input, displaying a text representation of the typed text entry input in the text entry field via the display generation component, wherein the text representation of the typed text entry input is displayed in a second appearance different from the first appearance.

49. The method according to claim 48, wherein displaying the text representation of the utterance input in the first appearance includes displaying the text representation of the utterance input with a growing effect, and displaying the text representation of the typed text entry input in the text entry field in the second appearance includes displaying the text representation of the typed text entry input in the text entry field without the growing effect.

50. Displaying the text representation of the speech input in the first appearance is, The display generation component displays individual parts of the text representation of the utterance input within the text entry field, and then displays them in one or more colors that change over time. The method according to claim 48 or 49, further comprising, after the period has elapsed, displaying the individual parts of the text representation of the utterance input in individual colors that do not change over time, via the display generation component.

51. The method according to claim 50, wherein displaying the individual parts of the text representation of the speech input in a color that changes over time includes displaying the individual parts of the text representation of the speech input in a color that changes over time in response to changes in the audio level of the speech input over time.

52. The display generation component further includes displaying a text insertion marker in the text entry field indicating the location in the text entry field where additional text is added in response to receiving text entry input via the display generation component, While the first speech input is being detected, the text insertion marker is displayed with individual visual effects. The method according to any one of claims 27 to 50, wherein the text insertion marker is displayed without the individual visual effects while the first utterance input is not detected.

53. The method according to claim 52, wherein the individual visual effects include visual characteristics that change over time in response to changes in the audio level of the first speech input over time.

54. Displaying the aforementioned text entry field means While the first speech input is being detected, the text entry field is displayed with individual visual effects, The method according to any one of claims 27 to 53, comprising displaying the text entry field without the individual visual effects while the first utterance input is not being detected.

55. The method according to claim 54, wherein the individual visual effect is a glowing visual effect.

56. The method according to claim 55, wherein the growing visual effect includes a visual characteristic having a value that changes over time in response to a change in the audio level of the first speech input over time.

57. The method according to any one of claims 54 to 56, wherein displaying the text entry field with the individual visual effects includes displaying the text entry field in a first color, and displaying the text entry field without the individual visual effects includes displaying the text entry field in a second color different from the first color.

58. The method according to claim 57, wherein displaying the text entry field in the first color includes changing the color of the text entry field over time in response to changes in the audio level of the first utterance input over time.

59. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 27 to 58.

60. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 27 to 58.

61. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 27 to 58.

62. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are The text entry field is displayed via the aforementioned display generation component. While the text entry field is being displayed via the display generation component, The system detects a first utterance input from a user of the computer system via one or more input devices. In response to detecting the first utterance input from the user, When the first utterance input from the user is received, and it is determined that the user's attention is directed to the text entry field, the text representation of the first utterance input is displayed in the text entry field via the display generation component. A non-temporary computer-readable storage medium including a command to stop displaying the text representation of the first utterance input in the text entry field, in accordance with the determination that the user's attention is not directed to the text entry field when the first utterance input from the user is received.

63. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are The text entry field is displayed via the aforementioned display generation component. While the text entry field is being displayed via the display generation component, The system detects a first utterance input from a user of the computer system via one or more input devices. In response to detecting the first utterance input from the user, When the first utterance input from the user is received, and it is determined that the user's attention is directed to the text entry field, the text representation of the first utterance input is displayed in the text entry field via the display generation component. A computer system including an instruction to refrain from displaying the text representation of the first utterance input in the text entry field, in accordance with the determination that the user's attention is not directed to the text entry field when the first utterance input from the user is received.

64. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is The means for displaying a text entry field via the aforementioned display generation component, While the text entry field is being displayed via the display generation component, The system detects a first utterance input from a user of the computer system via one or more input devices. In response to detecting the first utterance input from the user, When the first utterance input from the user is received, and it is determined that the user's attention is directed to the text entry field, the text representation of the first utterance input is displayed in the text entry field via the display generation component. A computer system comprising: means for stopping the display of the text representation of the first utterance input in the text entry field when the user's attention is not directed to the text entry field when the first utterance input from the user is received.

65. It is a method, In a computer system that communicates with a display generation component and one or more input devices, Through the display generation component, a first object, the first object including a text entry field, displays a three-dimensional environment from an individual viewpoint including the first object at an individual location within the three-dimensional environment, While displaying the three-dimensional environment from a specific viewpoint that includes the first object containing the text entry field at the specific location within the three-dimensional environment, a first input corresponding to the selection of the text entry field is detected via one or more input devices. In response to detecting the first input, In accordance with the determination that the individual location in the three-dimensional environment is a first location greater than the threshold distance from the individual viewpoint, the display generation component displays a keyboard at the keyboard location in the three-dimensional environment according to the first input, wherein the keyboard is for entering text into the text entry field, and the keyboard location in the three-dimensional environment is less than the threshold distance from the individual viewpoint. A method comprising: displaying the keyboard at the keyboard location in the three-dimensional environment according to the first input via the display generation component, based on the determination that the individual location in the three-dimensional environment is a second location different from the first location, and the second location is greater than the threshold distance from the individual viewpoint.

66. The method according to claim 65, further comprising, in response to detecting the first input, determining that the individual location in the three-dimensional environment is a third location less than the threshold distance from the individual viewpoint, and via the display generation component, displaying the keyboard at a second keyboard location in the three-dimensional environment according to the first input, wherein the second keyboard location is closer to the individual viewpoint than the keyboard location.

67. The method according to claim 65 or 66, further comprising maintaining the display of the first object at the individual location via the display generation component in response to the detection of the first input.

68. The method according to any one of claims 65 to 67, wherein displaying the first object includes displaying the first object at a first angle with respect to a specific reference in the three-dimensional environment via the display generation component, and displaying the keyboard includes displaying the keyboard at a second angle different from the first angle with respect to the specific reference in the three-dimensional environment via the display generation component.

69. The method according to any one of claims 65 to 68, wherein, if selected, displaying the keyboard in response to the detection of the first input includes displaying user interface elements associated with the keyboard, which cause the computer system to initiate a process of repositioning the keyboard in the three-dimensional environment.

70. Detecting input directed to the user interface element, which includes a request to reposition the keyboard in the three-dimensional environment, via one or more input devices, and which includes a request to update the distance between the keyboard and the individual viewpoints in the three-dimensional environment from the current distance to the updated distance; In response to the input directed to the user interface element, In accordance with the determination that the updated distance is within the first distance range, the keyboard is displayed via the display generation component at individual locations in the three-dimensional environment that are at the first distance from the user's viewpoint. The method according to claim 69, further comprising: determining that the updated distance is within a second distance range different from the first distance range, and then, via the display generation component, displaying the keyboard at a separate location in the three-dimensional environment that is at a second distance different from the first distance from the user's viewpoint.

71. Detecting input directed to the user interface element, which includes a request to reposition the keyboard in the three-dimensional environment, via one or more input devices, and which includes a request to update the distance between the keyboard and the individual viewpoints in the three-dimensional environment from the current distance to the updated distance; In response to the input directed to the user interface element, In accordance with the determination that the updated distance is a first distance from the user's viewpoint, the keyboard is displayed in the three-dimensional environment at a first angle with respect to an individual reference in the three-dimensional environment via the display generation component, The method according to claim 69 or 70, further comprising: determining that the updated distance is a second distance different from the first distance from the user's viewpoint, displaying the keyboard in the three-dimensional environment via the display generation component at a second angle different from the first angle with respect to the individual reference in the three-dimensional environment.

72. The method according to any one of claims 65 to 71, wherein, if selected, displaying the keyboard in response to the detection of the first input, the method includes displaying a user interface element that causes the computer system to initiate a process of resizing the keyboard in the three-dimensional environment.

73. The method according to any one of claims 65 to 72, wherein detecting the first input includes detecting the user's attention directed to the text entry field and a predetermined gesture performed by an individual part of the user via one or more input devices.

74. While displaying the three-dimensional environment from a specific viewpoint that includes the first object containing the text entry field at the specific location within the three-dimensional environment, a second input is detected via one or more input devices in response to a request to initiate a process for dictating text input directed to the text entry field. The method according to any one of claims 65 to 73, further comprising: detecting the second input, initiating the process for dictating the text input directed to the text entry field without displaying the keyboard via the display generation component.

75. The method according to any one of claims 65 to 74, wherein displaying the keyboard in response to the first input includes displaying a representation of a portion of the first object, including at least a portion of the text entry field, via the display generation component.

76. While the keyboard is displayed in response to the first input, the display generation component displays the cursor in the text entry field at a first location in the text entry field, and displays the representation of the cursor in the representation of the part of the first object at the corresponding first location in the representation of the part of the first object, While displaying the representation of a portion of the first object, including the representation of the cursor, via the display generation component, one or more inputs directed to the keyboard via one or more input devices are detected in response to a request to enter text into the text entry field. In response to the one or more inputs, Displaying the text in the text entry field and the representation of the text in the representation of the part of the first object, including, via the display generation component, displaying the cursor at a second location in the text entry field based on one or more inputs corresponding to the request to input the text into the text entry field, and displaying the representation of the cursor in the representation of the part of the first object at a corresponding second location in the representation of the part of the first object, The method according to claim 75, further comprising updating a separate portion of the first object contained in the representation of the portion of the first object to maintain the display of the representation of the cursor at the corresponding second location in the representation of the portion of the first object via the display generation component.

77. While the three-dimensional environment is displayed from a specific viewpoint that includes the first object containing the text entry field at the specific location within the three-dimensional environment, a second input corresponding to a request to enter text into the text entry field is detected via the hardware keyboard of one or more input devices. In response to detecting the second input, The text is displayed in the text entry field via the display generation component, The method according to claim 75 or 76, further comprising displaying the representation of a portion of the first object, including the representation of the text entered via the hardware keyboard, without displaying the keyboard, via the display generation component.

78. Displaying the keyboard includes displaying the keyboard at a first angle with respect to individual references in the three-dimensional environment via the display generation component, The method according to any one of claims 75 to 77, wherein displaying the representation of the portion of the first object includes displaying the representation of the portion of the first object at a third angle different from the second angle with respect to the individual reference in the three-dimensional environment.

79. Displaying the representation of the part of the first object is In accordance with the determination that the spatial relationship between the individual viewpoint of the user of the computer system and the representation of a part of the first object is a first spatial relationship, the representation of a part of the first object is displayed at a first angle with respect to an individual reference in the three-dimensional environment via the display generation component, The method according to any one of claims 75 to 78, comprising: displaying the representation of the part of the first object at a second angle different from the first angle with respect to an individual reference plane in the three-dimensional environment via the display generation component, in accordance with the determination that the spatial relationship between the individual viewpoint of the user and the representation of the part of the first object is a second spatial relationship.

80. The method according to any one of claims 75 to 79, wherein displaying the first object includes displaying the first object at a first angle with respect to a specific reference in the three-dimensional environment via the display generation component, and displaying the representation of the part of the first object includes displaying the representation of the part of the first object at a second angle different from the first angle with respect to the specific reference in the three-dimensional environment via the display generation component.

81. Displaying the first object includes displaying selectable options included in the first object via the display generation component, and displaying the representation of a portion of the first object includes displaying the representation of the selectable options within the representation of a portion of the first object via the display generation component, and the method is To detect a second input directed to the selectable option included in the first object via one or more input devices, In response to detecting the second input, the individual actions associated with the selectable options are performed, To detect a third input directed to the representation of the selectable option within the representation of the part of the first object via one or more input devices, The method according to any one of claims 75 to 80, further comprising detecting the third input and ceasing to perform the individual operation associated with the selectable option.

82. While displaying the representation of a portion of the first object, which includes at least the portion of the text entry field, via the display generation component, a second input is detected via one or more input devices, which is a second input targeting the representation of the individual text within the representation of the portion of the first object, wherein the second input corresponds to a request to select a specific portion of the individual text. In response to detecting the second input, The display generation component updates the display of the representation of the individual text to indicate the selection of the individual part of the individual text, The method according to any one of claims 75 to 81, further comprising updating the display of the text entry field to indicate the selection of the individual portion of the individual text via the display generation component.

83. While the first object, the representation of a part of the first object, and the keyboard are being displayed via the display generation component, one or more inputs directed to the keyboard are detected via one or more input devices in response to a request to enter text into the text entry field. The method according to any one of claims 75 to 82, further comprising displaying, via the display generation component, the text in the text entry field and the representation of the text in the representation of the part of the first object, in response to one or more of the inputs.

84. The method according to any one of claims 75 to 83, wherein displaying the keyboard in response to the first input includes displaying a plurality of selectable options associated with a text operation directed to the text entry field via the display generation component, the plurality of selectable options being displayed between the representation of the portion of the first object and the keyboard in the three-dimensional environment.

85. The keyboard location is at a first distance from the individual viewpoint, and the method is In response to detecting the first input, The method according to any one of claims 65 to 84, further comprising: displaying the keyboard via the display generation component at a fourth location that is a second distance from the user's individual viewpoint, according to the determination that the individual location in the three-dimensional environment is a third location, and the third location is less than the threshold distance from the individual viewpoint; and displaying the keyboard via the display generation component at a fifth location that is a third distance different from the second distance from the user's individual viewpoint, according to the determination that the individual location in the three-dimensional environment is a fourth location different from the third location, and the fourth location is less than the threshold distance from the individual viewpoint.

86. The first location in the three-dimensional environment has a first vertical position in the three-dimensional environment, and the second location in the three-dimensional environment has a second vertical position different from the first vertical position in the three-dimensional environment. Displaying the keyboard at the keyboard location in the three-dimensional environment, in accordance with the determination that the individual location in the three-dimensional environment is the first location, includes displaying the keyboard at a third vertical position according to the first vertical position of the first location via the display generation component. The method according to any one of claims 65 to 85, wherein displaying the keyboard at the keyboard location in the three-dimensional environment, in accordance with the determination that the individual location in the three-dimensional environment is the second location, includes displaying the keyboard via the display generation component at a fourth vertical position different from the third vertical position, according to the second vertical position of the second location.

87. The third vertical position has a distinct angular offset from the first location with respect to the individual viewpoint in the three-dimensional environment, The method according to claim 85, wherein the fourth vertical position has the individual angular offset from the second location to the individual viewpoint in the three-dimensional environment.

88. Detecting the first input includes detecting the user's attention directed to a first location in the text entry field via one or more input devices, In response to the first input, displaying the keyboard at the keyboard location in the three-dimensional environment is: In accordance with the determination that the first location within the text entry field has a first horizontal position in the three-dimensional environment, the keyboard is displayed in a second horizontal position according to the first horizontal position via the display generation component. The method according to any one of claims 65 to 87, comprising: displaying the keyboard via the display generation component at a fourth horizontal position different from the second horizontal position according to the second horizontal position, based on the determination that the first location in the text entry field has a third horizontal position different from the first horizontal position in the three-dimensional environment.

89. While the keyboard is displayed at a third location in the three-dimensional environment that is within a second threshold distance of the individual viewpoints via the display generation component, Receiving text entry input directed to the keyboard via one or more of the aforementioned input devices, Upon receiving the aforementioned text entry input, While the default portion of the user of the computer system is within the direct input threshold distance of the physical location corresponding to the keyboard in the three-dimensional environment, the determination that the text entry input involves performing a first gesture with the default portion of the user, and the input of text into the text entry field according to the text entry input, The method according to any one of claims 65 to 88, further comprising: when the default portion of the user is farther than the direct input threshold distance of the physical location corresponding to the keyboard in the three-dimensional environment, the user refrains from entering the text into the text entry field in accordance with the text entry input, in accordance with the determination that the text entry input involves performing a second gesture with the default portion of the user.

90. While the keyboard is displayed at a third location in the three-dimensional environment between the second threshold distance and the third threshold distance of the individual viewpoints via the display generation component, Receiving text entry input directed to the keyboard via one or more of the aforementioned input devices, Upon receiving the aforementioned text entry input, While the default portion of the user of the computer system is within the direct input threshold distance of the physical location corresponding to the keyboard in the three-dimensional environment, the determination that the text entry input involves performing a first gesture with the default portion of the user, and the input of text into the text entry field according to the text entry input, The method according to any one of claims 65 to 89, further comprising: entering text into the text entry field according to the text entry input, in accordance with the determination that the text entry input involves performing a second gesture with the default portion of the user, while the default portion of the user is farther than the direct input threshold distance of the physical location corresponding to the keyboard in the three-dimensional environment.

91. While the keyboard is displayed in a third location within the three-dimensional environment that is greater than the second threshold distance of the individual viewpoints via the display generation component, Receiving text entry input directed to the keyboard via one or more of the aforementioned input devices, Upon receiving the aforementioned text entry input, The text entry input includes performing a first gesture with a default part of the user of the computer system, and entering text into the text entry field according to the text entry input, based on the determination that the default part of the user is farther than the direct input threshold distance of the physical location corresponding to the keyboard in the three-dimensional environment. The method according to any one of claims 65 to 90, further comprising: while the default portion of the user is within the direct input threshold distance of the physical location corresponding to the keyboard in the three-dimensional environment, ceasing to input the text into the text entry field in accordance with the text entry input, in accordance with the determination that the text entry input involves performing a second gesture with the default portion of the user.

92. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 65 to 91.

93. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 65 to 91.

94. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 65 to 91.

95. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are Through the display generation component, a first object, the first object including a text entry field, displays a three-dimensional environment from an individual viewpoint including the first object at an individual location within the three-dimensional environment. While displaying the three-dimensional environment from a specific viewpoint that includes the first object containing the text entry field at the specific location within the three-dimensional environment, a first input corresponding to the selection of the text entry field is detected via one or more input devices. In response to detecting the first input, In accordance with the determination that the individual location in the three-dimensional environment is a first location greater than a threshold distance from the individual viewpoint, the display generation component displays a keyboard at the keyboard location in the three-dimensional environment according to the first input, wherein the keyboard is for entering text into the text entry field, and the keyboard location in the three-dimensional environment is less than the threshold distance from the individual viewpoint. A non-temporary computer-readable storage medium including a command to display the keyboard at the keyboard location in the three-dimensional environment according to the first input via the display generation component, based on the determination that the individual location in the three-dimensional environment is a second location different from the first location, and the second location is greater than the threshold distance from the individual viewpoint.

96. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are Through the display generation component, a first object, the first object including a text entry field, displays a three-dimensional environment from an individual viewpoint including the first object at an individual location within the three-dimensional environment. While displaying the three-dimensional environment from a specific viewpoint that includes the first object containing the text entry field at the specific location within the three-dimensional environment, a first input corresponding to the selection of the text entry field is detected via one or more input devices. In response to detecting the first input, In accordance with the determination that the individual location in the three-dimensional environment is a first location greater than a threshold distance from the individual viewpoint, the display generation component displays a keyboard at the keyboard location in the three-dimensional environment according to the first input, wherein the keyboard is for entering text into the text entry field, and the keyboard location in the three-dimensional environment is less than the threshold distance from the individual viewpoint. A computer system including an instruction to display the keyboard at the keyboard location in the three-dimensional environment according to a first input via the display generation component, based on the determination that the individual location in the three-dimensional environment is a second location different from the first location, and the second location is greater than the threshold distance from the individual viewpoint.

97. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is The means for displaying a three-dimensional environment from an individual viewpoint including the first object, via the display generation component, where the first object includes a text entry field, at an individual location within the three-dimensional environment, Means for detecting a first input corresponding to the selection of the text entry field via one or more input devices while the three-dimensional environment is being displayed from a particular viewpoint that includes the first object containing the text entry field at the particular location within the three-dimensional environment, In response to detecting the first input, In accordance with the determination that the individual location in the three-dimensional environment is a first location greater than a threshold distance from the individual viewpoint, the display generation component displays a keyboard at the keyboard location in the three-dimensional environment according to the first input, wherein the keyboard is for entering text into the text entry field, and the keyboard location in the three-dimensional environment is less than the threshold distance from the individual viewpoint. A computer system comprising: means for displaying the keyboard at the keyboard location in the three-dimensional environment according to a first input via the display generation component, based on the determination that the individual location in the three-dimensional environment is a second location different from the first location, and the second location is greater than the threshold distance from the individual viewpoint.

98. It is a method, In a computer system that communicates with a display generation component and one or more input devices, The display generation component enables the display of a three-dimensional environment including a keyboard having a plurality of keys, wherein the keyboard is displayed at a first location in the three-dimensional environment, and the plurality of keys extend from a region corresponding to the surface of the keyboard at a first distance. While the three-dimensional environment, including the keyboard, is displayed at the first location within the three-dimensional environment, the system receives a first input via one or more input devices, which includes the movement of a part of the user's body toward an individual key of the plurality of keys on the keyboard. In response to receiving the first input, the movement toward the individual key includes a movement toward a location that corresponds to the first key and is less than a threshold distance from the surface of the keyboard, and the threshold distance is determined to be closer to the keyboard than the first distance from the surface of the keyboard. The first key is moved toward the surface of the keyboard at the first location in the three-dimensional environment by a second distance, the second distance being closer to the surface of the keyboard than the location, A method comprising performing one or more actions corresponding to the selection of the first key.

99. In response to receiving the first input, and in accordance with the determination that the movement toward the individual key includes a movement toward the location corresponding to the first key, moving the first key toward the surface of the keyboard by the second distance is: While detecting a portion of the movement of the user's body, including movement from the surface of the keyboard to the threshold distance, the first key is moved toward the surface of the keyboard according to the portion of the movement. The method of claim 98, comprising: moving the first key closer to the keyboard by the remainder of the second distance in response to the movement of the part of the user's body toward the first key reaching the threshold distance from the surface of the keyboard, wherein moving the first key by the remainder of the second distance is independent of any further movement of the part of the user's body.

100. In response to receiving the first input, the determination that the movement toward the individual key includes a movement toward a second location which corresponds to the first key and is greater than the threshold distance from the surface of the keyboard and less than the first distance from the surface of the keyboard, The first key is moved by a third distance in accordance with the movement of a part of the user's body toward the surface of the keyboard at the first location in the three-dimensional environment, The method according to claim 98 or 99, further comprising: discontinuing to perform one or more operations corresponding to the selection of the first key.

101. After detecting the movement of the part of the user's body included in the first input, To detect a second movement of the part of the user's body away from the individual keys via one or more input devices, The method of claim 100, further comprising detecting the second movement of the part of the user's body, and in accordance with the determination that the movement toward the individual key includes a movement toward the second location corresponding to the first key, moving the first key away from the surface of the keyboard in accordance with the second movement of the part of the user's body.

102. The method according to any one of claims 98 to 101, further comprising, in response to receiving the first input, stopping the movement of the individual key toward the surface of the keyboard according to a determination that the movement toward the individual key involves a movement to a second location greater than a first distance from the surface of the keyboard.

103. In response to receiving the first input, the determination that the movement toward the individual key corresponds to a second key different from the first key and includes a movement to a second location less than the threshold distance from the surface of the keyboard, Moving the second key by the second distance toward the surface of the keyboard at the first location in the three-dimensional environment, The method according to any one of claims 98 to 102, further comprising performing one or more actions corresponding to the selection of the second key.

104. In response to receiving the first input, the determination that the movement toward the individual key corresponds to a second key different from the first key and includes a movement to a second location which is greater than the threshold distance from the surface of the keyboard and less than the first distance from the surface of the keyboard, The second key is moved by a third distance in accordance with the movement of a part of the user's body toward the surface of the keyboard at the first location in the three-dimensional environment, The method according to any one of claims 98 to 103, further comprising: discontinuing to perform one or more operations corresponding to the selection of the second key.

105. While the keyboard is displayed at the first location in the three-dimensional environment, The display generation component displays a selectable option at a second location in the three-dimensional environment, wherein the selectable option extends a third distance from a backplane different from the surface of the keyboard. To detect a second input via one or more input devices, including the movement of the part of the user's body toward the selectable option, Upon receiving the second input, The determination that the movement toward the selectable option corresponds to the movement of the selectable option toward the backplane by at least the third distance, and the execution of one or more actions corresponding to the selection of the selectable option, The method according to any one of claims 98 to 104, further comprising: determining that the movement toward the selectable option corresponds to a movement of the selectable option less than the third distance toward the backplane, and then ceasing to perform the one or more actions corresponding to the selection of the selectable option.

106. Displaying the three-dimensional environment including the keyboard means The display generation component includes displaying a simulated shadow corresponding to the part of the user's body overlaid on a second key among the plurality of keys of the keyboard, In accordance with the determination that the location of the user's body part in the three-dimensional environment corresponds to a third key among the plurality of keys on the keyboard, the simulated shadow is overlaid and displayed on the third key. The method according to any one of claims 98 to 105, wherein, according to the determination that the location of the part of the user's body in the three-dimensional environment corresponds to a fourth key among the plurality of keys on the keyboard, the simulated shadow is displayed overlaid on the fourth key.

107. Displaying the three-dimensional environment including the keyboard means This includes displaying a simulated shadow of the user's body overlaid on a second key among the plurality of keys on the keyboard via the display generation component, According to the determination that the location of the part of the user's body in the three-dimensional environment is at a second distance from the second key, the simulated shadow is displayed with visual characteristics having a first value. The method according to any one of claims 98 to 106, wherein, in accordance with the determination that the location of the part of the user's body in the three-dimensional environment is at a third distance different from the second key, the simulated shadow is displayed with the visual characteristics having a second value different from the first value.

108. Displaying the three-dimensional environment including the keyboard is done via the display generation component, A simulated shadow corresponding to the part of the user's body is overlaid on a second key among the plurality of keys of the keyboard, The method according to any one of claims 98 to 107, comprising simultaneously displaying a simulated shadow corresponding to a second part of the user's body, which is overlaid on a third key, different from the second key, among the plurality of keys of the keyboard.

109. In response to receiving the first input, according to the determination that the movement toward the individual key includes a movement toward the location corresponding to the first key, The method according to any one of claims 98 to 108, further comprising displaying an animation of a first portion of the keyboard including the first key via the display generating component, wherein the animation indicates that the first key has been selected without modifying the display of a second portion of the keyboard outside the first portion of the keyboard.

110. While the three-dimensional environment, including the keyboard, is displayed at the first location within the three-dimensional environment, While detecting the second movement of the part of the user's body toward the second key, the movement of the second part of the user's body toward the third key is detected. While detecting the second movement of the part of the user's body, in response to detecting the movement of the second part of the user's body toward the third key, According to the determination that the second movement of the part of the body includes movement to a third location which corresponds to the second key and is less than the threshold distance from the surface of the keyboard, and according to the determination that the movement of the second part of the user's body includes movement to a fourth location which corresponds to the third key and is less than the threshold distance from the surface of the keyboard, Moving the second key and the third key by the second distance toward the surface of the keyboard at the first location in the three-dimensional environment, The method according to any one of claims 98 to 109, further comprising performing one or more actions corresponding to the selection of the second key and the third key.

111. The first input is detected while the keyboard is displayed in a first mode that does not include displaying a cursor overlaid on the keyboard, and the method is While the keyboard is being displayed in the first mode, it is detected that one or more criteria associated with displaying the keyboard in a second mode different from the first mode are met, Displaying the keyboard in the three-dimensional environment in the second mode via the display generation component, including, upon detection that one or more of the criteria associated with displaying the keyboard in the second mode are met, displaying a cursor overlaid on a second key of the plurality of keys on the keyboard corresponding to the location of the part of the user's body in the three-dimensional environment via the display generation component, While the keyboard is displayed in the second mode, A second input, which includes a gesture performed using the part of the user's body via one or more input devices, the second input receiving a second input that satisfies one or more criteria, Upon receiving the second input, According to the determination that the second key is the third key, Moving the third key toward the surface of the keyboard, Performing one or more actions corresponding to the selection of the third key, According to the determination that the second key is the fourth key, Moving the fourth key toward the surface of the keyboard, The method according to any one of claims 98 to 110, further comprising performing one or more actions corresponding to the selection of the fourth key.

112. Displaying the second key among the plurality of keys of the keyboard at a distance of the first distance from the surface of the keyboard is in accordance with the determination that the individual location of the part of the user's body does not meet one or more criteria associated with the second key, and the method is The method according to any one of claims 98 to 111, further comprising updating the keyboard via the display generating component to display the second key at a third distance from the surface of the keyboard, in accordance with a determination that the individual location of the part of the user's body satisfies one or more criteria associated with the second key, including criteria that are satisfied when the individual location of the part of the user's body is within a threshold distance of the location corresponding to the second key, wherein the third distance is greater than the first distance.

113. In response to receiving the first input, according to the determination that the movement toward the individual key includes a movement toward the location corresponding to the first key, The method according to any one of claims 98 to 112, further comprising presenting an audio indication of the selection of the first key via one or more output devices that communicate with the computer system.

114. The first input is received while the keyboard is in a first mode, and the method is While displaying the three-dimensional environment including the keyboard in a second mode different from the first mode, receiving a second input via one or more input devices directed to the individual keys, wherein the second input includes a gesture performed using the part of the user's body, and does not include the movement of the part of the user's body to a location corresponding to the individual key, Upon receiving the second input, and in accordance with the determination that the second input satisfies one or more criteria and that the second input is directed to the first key, Moving the first key toward the surface of the keyboard, Performing one or more actions corresponding to the selection of the first key, The method according to claim 113, further comprising presenting a second audio indication of the selection of the first key, which is different from the audio indication of the selection of the first key, via one or more output devices.

115. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 98 to 114.

116. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 98 to 114.

117. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 98 to 114.

118. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are The display generation component displays a three-dimensional environment including a keyboard having a plurality of keys, wherein the keyboard is displayed at a first location in the three-dimensional environment, and the plurality of keys extend from a region corresponding to the surface of the keyboard at a first distance. While the three-dimensional environment, including the keyboard, is displayed at the first location within the three-dimensional environment, a first input is received via one or more input devices, including the movement of a part of the user's body toward an individual key of the plurality of keys on the keyboard. In response to receiving the first input, the movement toward the individual key includes a movement toward a location that corresponds to the first key and is less than a threshold distance from the surface of the keyboard, and the threshold distance is determined to be closer to the keyboard than the first distance from the surface of the keyboard. Move the first key toward the surface of the keyboard at the first location in the three-dimensional environment by a second distance, the second distance being closer to the surface of the keyboard than the location; A non-temporary computer-readable storage medium containing instructions for performing one or more actions corresponding to the selection of the first key.

119. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are The display generation component displays a three-dimensional environment including a keyboard having a plurality of keys, wherein the keyboard is displayed at a first location in the three-dimensional environment, and the plurality of keys extend from a region corresponding to the surface of the keyboard at a first distance. While the three-dimensional environment, including the keyboard, is displayed at the first location within the three-dimensional environment, a first input is received via one or more input devices, including the movement of a part of the user's body toward an individual key of the plurality of keys on the keyboard. In response to receiving the first input, the movement toward the individual key includes a movement toward a location that corresponds to the first key and is less than a threshold distance from the surface of the keyboard, and the threshold distance is determined to be closer to the keyboard than the first distance from the surface of the keyboard. Move the first key toward the surface of the keyboard at the first location in the three-dimensional environment by a second distance, the second distance being closer to the surface of the keyboard than the location; A computer system including instructions for performing one or more actions corresponding to the selection of the first key.

120. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is Means for displaying a three-dimensional environment, including a keyboard having a plurality of keys, wherein the keyboard is displayed at a first location in the three-dimensional environment, and the plurality of keys extend from a region corresponding to the surface of the keyboard at a first distance, via the display generation component; Means for receiving a first input via one or more input devices, including the movement of a part of the user's body toward an individual key of the plurality of keys on the computer system, while the three-dimensional environment, including the keyboard, is displayed at the first location within the three-dimensional environment. In response to receiving the first input, the movement toward the individual key includes a movement toward a location that corresponds to the first key and is less than a threshold distance from the surface of the keyboard, and the threshold distance is determined to be closer to the keyboard than the first distance from the surface of the keyboard. Move the first key toward the surface of the keyboard at the first location in the three-dimensional environment by a second distance, the second distance being closer to the surface of the keyboard than the location; A computer system comprising means for performing one or more actions corresponding to the selection of the first key.

121. It is a method, In a computer system that communicates with a display generation component and one or more input devices, The display generation component enables the display of a three-dimensional environment including a keyboard having a plurality of keys, wherein the keyboard is displayed at a first location in the three-dimensional environment, and the keyboard is displayed without displaying a cursor for selecting one or more of the plurality of keys. Without displaying the cursor, while the three-dimensional environment, including the keyboard, is displayed at the first location in the three-dimensional environment, a first input is received via one or more input devices, including a change in the position of one or more parts of the user of the computer system. Upon receiving the first input, A method comprising displaying the cursor overlaid on a portion of the plurality of keys on the keyboard via the display generation component, wherein the cursor indicates the portion of the plurality of keys that has current focus.

122. While the keyboard in the three-dimensional environment and the cursor overlaid on a part of the keyboard are displayed via the display generation component, Receiving a second input directed to the keyboard, including input from each of the one or more parts of the user, via the one or more input devices; Upon receiving the second input, In accordance with the determination that a portion of the plurality of keys currently having focus is a first key among the plurality of keys, To perform the function associated with the first key among the plurality of keys, In accordance with the determination that a portion of the plurality of keys currently having focus is a second key among the plurality of keys, The method according to claim 121, further comprising performing a function associated with the second key among the plurality of keys.

123. While the display generation component displays the keyboard in the three-dimensional environment and the cursor overlaid on the part of the keyboard, the part of the keyboard corresponding to an individual key among the plurality of keys, Receiving a second input directed to the keyboard via one or more input devices, wherein the second input includes a gesture performed by each of the one or more parts of the user that meets one or more criteria, The method according to claim 121 or 122, further comprising: receiving the second input, performing a function associated with the individual key among the plurality of keys that currently have focus.

124. The cursor indicates the portion of the plurality of keys that currently have focus, based on a first portion of each of the one or more portions of the user, and the method is Upon receiving the first input, The method according to any one of claims 121 to 123, further comprising displaying a second cursor overlaid on a second portion of the plurality of keys of the keyboard via the display generation component, wherein the second cursor indicates the second portion of the plurality of keys that currently has a second focus based on the second portion of each of the one or more portions of the user, and the second cursor is displayed simultaneously with the cursor.

125. While the keyboard in the three-dimensional environment, the cursor overlaid on the first key among the plurality of keys, and the second cursor overlaid on the second key among the plurality of keys are displayed via the display generation component, Receiving a sequence of one or more inputs directed to each of several keys on the keyboard, including the simultaneous selection of the first key and the second key, via one or more input devices, The method according to any one of claims 121 to 124, further comprising: receiving the sequence of one or more inputs; and performing one or more functions associated with each of the multiple keys on the keyboard.

126. The method according to any one of claims 121 to 125, wherein the change in the position of each of the one or more parts of the user of the computer system included in the first input includes a change in the relative orientation between one or more wrists of the user of the computer system.

127. The method according to any one of claims 121 to 126, further comprising displaying a simulated shadow of the cursor via the display generation component in response to receiving the first input, wherein the simulated shadow of the cursor is displayed on the portion of the plurality of keys of the keyboard having the current focus.

128. While the keyboard and cursor are being displayed, the display generation component further includes displaying the keyboard's backplane, wherein the plurality of keys of the keyboard are overlaid on the keyboard's backplane in the three-dimensional environment. According to the determination that the cursor is overlaid on the first portion of the plurality of keys and not overlaid on the second portion of the plurality of keys, The first portion of the plurality of keys is displayed with a first amount of visual separation from the backplane of the keyboard, The second portion of the plurality of keys is represented by a second amount of visual separation from the backplane of the keyboard, and the second amount of visual separation is smaller than the first amount of visual separation. According to the determination that the cursor is overlaid on the second portion of the plurality of keys and not overlaid on the first portion of the plurality of keys, The second portion of the plurality of keys is displayed with a first amount of visual separation from the backplane of the keyboard. The method according to any one of claims 121 to 127, wherein the first portion of the plurality of keys is displayed with a second amount of visual separation from the backplane of the keyboard.

129. The portion of the plurality of keys on the keyboard is based on the location of one or more portions of the user in the three-dimensional environment, and the method is While the keyboard and the cursor are displayed overlaid on a portion of the plurality of keys, the movement of one or more portions of the user from a location in the three-dimensional environment associated with a portion of the plurality of keys of the keyboard to a location in the three-dimensional environment associated with a second portion of the plurality of keys of the keyboard is detected. In response to detecting the movement of one or more parts of the user, The method according to any one of claims 121 to 128, further comprising updating the three-dimensional environment via the display generation component to display the cursor overlaid on a second portion of the plurality of keys, without displaying the cursor overlaid on a portion of the plurality of keys.

130. While the keyboard in the three-dimensional environment and the cursor overlaid on a part of the keyboard are displayed via the display generation component, Receiving one or more sequences of inputs via one or more input devices, which include detecting the movement of each of the one or more parts of the user through a sequence of locations associated with a separate set of keys, while each of the one or more parts of the user is in a predetermined shape. The method according to any one of claims 121 to 129, further comprising performing an action associated with the individual set of the plurality of keys in response to receiving the second input.

131. While the keyboard in the three-dimensional environment and the cursor overlaid on a part of the keyboard are displayed via the display generation component, Receiving a second input directed to a portion of the plurality of keys of the keyboard via one or more input devices, The method according to any one of claims 121 to 130, further comprising, in response to receiving the second input, displaying an animation of the second part of the keyboard, including the portion of the plurality of keys of the keyboard, via the display generation component, wherein the animation indicates that the portion of the plurality of keys has been selected without modifying the display of the third part of the keyboard outside the second part of the keyboard.

132. While the keyboard in the three-dimensional environment and the cursor overlaid on a part of the keyboard are displayed via the display generation component, Receiving a second input via one or more of the aforementioned input devices that corresponds to a request to change the keyboard input mode from cursor input mode to non-cursor input mode, The method according to any one of claims 121 to 131, further comprising: maintaining the display of the keyboard via the display generation component and discontinuing the display of the cursor via the display generation component in response to receiving the second input.

133. The method according to claim 132, wherein receiving the second input includes detecting a change in the orientation of one or more wrists of the user of the computer system via the one or more input devices.

134. While the keyboard in the three-dimensional environment and the cursor overlaid on a part of the keyboard are displayed via the display generation component, Receiving a second input directed to the keyboard via one or more of the aforementioned input devices, Upon receiving the second input, Activating a portion of the plurality of keys that currently have focus, The method according to any one of claims 121 to 133, further comprising generating a first audio indication corresponding to the selection of the plurality of keys via one or more output devices that communicate with the computer system.

135. While the keyboard is displayed in the three-dimensional environment without displaying the cursor, To detect a third input directed to a portion of the plurality of keys on the keyboard via one or more input devices, Upon receiving the third input, Activating some of the multiple keys on the keyboard, The method according to claim 134, further comprising generating a second audio indication different from the first audio indication corresponding to the selection of the plurality of keys via one or more output devices that communicate with the computer system.

136. While the keyboard is displayed in the three-dimensional environment via the display generation component without displaying the cursor, The receiving of a second input directed to a second portion of the plurality of keys of the keyboard via one or more input devices, wherein the second input is provided by each of the one or more portions of the user. Upon receiving the second input, The system performs an action associated with the second part of the plurality of keys in accordance with the determination that the second input includes one or more of the user within a threshold distance of the keyboard, and discontinues performing the action associated with the second part of the plurality of keys in accordance with the determination that the second input includes one or more of the user that is further than the threshold distance from the keyboard. While the keyboard is displayed in the three-dimensional environment with the cursor overlaid on some of the keys via the display generation component, A third input directed to the keyboard via one or more input devices, the third input being received by each of the one or more parts of the user while each of the one or more parts of the user is within the threshold distance of the keyboard. The method according to any one of claims 121 to 135, further comprising performing an action associated with the portion of the plurality of keys on the keyboard in response to receiving the third input.

137. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 121 to 136.

138. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 121 to 136.

139. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 121 to 136.

140. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are The display generation component displays a three-dimensional environment including a keyboard having a plurality of keys, the keyboard being displayed at a first location in the three-dimensional environment, and the keyboard being displayed without displaying a cursor for selecting one or more of the plurality of keys. Without displaying the cursor, while the three-dimensional environment, including the keyboard, is displayed at the first location in the three-dimensional environment, a first input is received via one or more input devices, including a change in the position of one or more parts of the user of the computer system. Upon receiving the first input, A non-temporary computer-readable storage medium includes instructions for displaying the cursor overlaid on a portion of the plurality of keys of the keyboard via the display generation component, wherein the cursor indicates the portion of the plurality of keys that has current focus.

141. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are The display generation component displays a three-dimensional environment including a keyboard having a plurality of keys, the keyboard being displayed at a first location in the three-dimensional environment, and the keyboard being displayed without displaying a cursor for selecting one or more of the plurality of keys. Without displaying the cursor, while the three-dimensional environment, including the keyboard, is displayed at the first location in the three-dimensional environment, a first input is received via one or more input devices, including a change in the position of one or more parts of the user of the computer system. Upon receiving the first input, A computer system comprising instructions for displaying the cursor overlaid on a portion of the plurality of keys on the keyboard via the display generation component, wherein the cursor indicates the portion of the plurality of keys that has current focus.

142. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is Means for displaying a three-dimensional environment including a keyboard, via the display generation component, wherein the keyboard is displayed at a first location in the three-dimensional environment, and the keyboard is displayed without displaying a cursor for selecting one or more of the multiple keys, Means for receiving a first input via one or more input devices, including a change in the position of one or more parts of a user of the computer system, while the three-dimensional environment, including the keyboard, is displayed at the first location in the three-dimensional environment without displaying the cursor, Upon receiving the first input, A computer system comprising means for displaying the cursor overlaid on a portion of the plurality of keys of the keyboard via the display generation component, wherein the cursor indicates the portion of the plurality of keys that has current focus.

143. It is a method, In a computer system that communicates with a display generation component and one or more input devices, The three-dimensional environment including a first region including a cursor is displayed via the aforementioned display generation component, To detect a first movement of an individual part of the user of the computer system via one or more input devices, In response to detecting the first movement of the individual part of the user, When the first movement of the individual part of the user is detected, the determination that the user's attention is directed to the first region of the three-dimensional environment restricts the movement of the cursor to the first region, while moving the cursor in accordance with the first movement of the individual part of the user, A method comprising displaying the cursor at a location located within the second region but outside the first region, in accordance with a determination that one or more criteria are met, including criteria that are met based on the user's attention being directed to a second region of the three-dimensional environment different from the first region of the three-dimensional environment when the first movement of the individual part of the user is detected.

144. The one or more criteria mentioned above include criteria that are met when the movement of the individual parts of the user exceeds a predetermined threshold amount of movement, and the method is In response to detecting the first movement of the individual part of the user while the user's attention is directed to the second area, Displaying the cursor at a location that is within the second region and outside the first region, in accordance with the determination that one or more criteria are met, including the first movement of the individual part of the user which includes a movement amount exceeding the predetermined threshold amount, The method according to claim 143, further comprising: maintaining the display of the cursor in the first area in accordance with the determination that one or more criteria are not met because the first movement of the individual part of the user includes a movement amount less than the predetermined threshold amount.

145. The one or more of the above criteria include a criterion that is met when the individual part of the user does not provide input for drawing with the cursor, and the method is In response to detecting the first movement of the individual part of the user while the user's attention is directed to the second area, Displaying the cursor at a location located within the second region and outside the first region, in accordance with the determination that one or more criteria are met, including the fact that the individual part of the user does not provide the input for drawing using the cursor, The method according to claim 143 or 144, further comprising maintaining the display of the cursor in the first area in accordance with the determination that one or more of the criteria are not met, because the individual part of the user provides the input for drawing using the cursor.

146. In response to detecting the first movement of the individual part of the user, Moving the cursor in accordance with the first movement of the individual part of the user, based on the determination that the cursor is performing a drawing operation while the individual part of the user is performing the first movement, includes moving the cursor by a first amount. The method according to any one of claims 143 to 145, wherein moving the cursor in accordance with the first movement of the individual part of the user, based on the determination that the cursor is not performing a drawing operation while the individual part of the user is performing the first movement, includes moving the cursor by a second amount greater than the first amount.

147. In response to detecting the first movement of the individual part of the user, The method according to any one of claims 143 to 146, further comprising, while the first movement is being performed, displaying a drawing having a profile corresponding to the movement of the cursor via the display generation component, based on the determination that the individual part of the user is an individual shape, and the individual shape corresponds to an individual shape that corresponds to a request to draw in the three-dimensional environment using the cursor.

148. While the cursor is displayed in the aforementioned three-dimensional environment, Receiving individual inputs corresponding to requests to make selections using the cursor via one or more input devices, The method according to any one of claims 143 to 147, further comprising: receiving the individual input, and determining that the cursor is within a threshold distance of a selectable user interface element in the three-dimensional environment when the individual input is received, and performing an action according to the selection of the selectable user interface element.

149. The method according to any one of claims 143 to 148, wherein the user's attention is determined by smoothing the gaze data to remove one or more high-frequency changes in gaze location over individual periods.

150. In response to detecting the first movement of the individual part of the user, according to the determination that one or more criteria are met, The method according to any one of claims 143 to 149, further comprising displaying the movement of the cursor from a first location in the three-dimensional environment to a second location in the three-dimensional environment via the display generating component, in accordance with the determination that the movement of the user's attention satisfies one or more respective criteria for the first movement of the individual part of the user, and in accordance with the determination that while the first movement is being performed, the individual part of the user is an individual shape, and the individual shape is an individual shape corresponding to a request to move the cursor, wherein the movement of the cursor is based on the movement of the user's attention and the movement of the individual part of the user.

151. In response to detecting the first movement of the individual part of the user, according to the determination that one or more criteria are met, The method according to claim 150, further comprising, in accordance with the determination that the movement of the user's attention satisfies one or more of the respective criteria for the first movement of the individual part of the user, and while the individual part of the user is performing the first movement, displaying a drawing in the three-dimensional environment via the display generation component from the first location to the second location of the cursor in the first area of ​​the three-dimensional environment, in accordance with the determination that the first shape is a first shape corresponding to a request to draw in the three-dimensional environment using the cursor.

152. The method according to any one of claims 143 to 151, further comprising detecting the first movement of the individual part of the user, and in accordance with the determination that the user's attention is directed to the first region of the three-dimensional environment when the first movement of the individual part of the user is detected, and in accordance with the determination that the amount of the first movement of the individual part of the user corresponds to the movement of the cursor outside the first region of the three-dimensional environment, moving the cursor to the boundary of the first region in the three-dimensional environment in accordance with the first movement of the individual part of the user.

153. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 143 to 152.

154. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 143 to 152.

155. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 143 to 152.

156. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are A three-dimensional environment including a first region including a cursor is displayed via the aforementioned display generation component. The first movement of an individual part of the user of the computer system is detected via one or more input devices. In response to detecting the first movement of the individual part of the user, When the first movement of the individual part of the user is detected, and in accordance with the determination that the user's attention is directed to the first region of the three-dimensional environment, the movement of the cursor to the first region is restricted, while the cursor is moved in accordance with the first movement of the individual part of the user. A non-temporary computer-readable storage medium, including instructions for displaying the cursor at a location within the second region but outside the first region, in accordance with a determination that one or more criteria are met, including criteria that are met based on the user's attention being directed to a second region of the three-dimensional environment different from the first region of the three-dimensional environment when the first movement of the individual part of the user is detected.

157. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are A three-dimensional environment including a first region including a cursor is displayed via the aforementioned display generation component. The first movement of an individual part of the user of the computer system is detected via one or more input devices. In response to detecting the first movement of the individual part of the user, When the first movement of the individual part of the user is detected, and in accordance with the determination that the user's attention is directed to the first region of the three-dimensional environment, the movement of the cursor to the first region is restricted, while the cursor is moved in accordance with the first movement of the individual part of the user. A computer system including an instruction to display the cursor at a location within the second region but outside the first region, in accordance with a determination that one or more criteria are met, including criteria that are met based on the user's attention being directed to a second region of the three-dimensional environment different from the first region of the three-dimensional environment when the first movement of the individual part of the user is detected.

158. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A means for displaying a three-dimensional environment including a first region including a cursor via the aforementioned display generation component, Means for detecting the first movement of individual parts of the user of the computer system via one or more input devices, In response to detecting the first movement of the individual part of the user, When the first movement of the individual part of the user is detected, and in accordance with the determination that the user's attention is directed to the first region of the three-dimensional environment, the movement of the cursor to the first region is restricted, while the cursor is moved in accordance with the first movement of the individual part of the user. A computer system comprising: means for displaying the cursor at a location within the second region but outside the first region, in accordance with a determination that one or more criteria are met, including criteria that are met based on the user's attention being directed to a second region of the three-dimensional environment different from the first region of the three-dimensional environment when the first movement of the individual part of the user is detected.

159. It is a method, In a computer system that communicates with a display generation component and one or more input devices, To simultaneously display a user interface including a text entry field and a text entry element configured to input text into the text entry field, via the aforementioned display generation component, While the text entry element and the user interface are being displayed simultaneously via the display generation component, The text entry input is to receive a text entry input directed to the text entry element via one or more input devices, the text entry input including a speech input. A method comprising: updating the display of the text entry element to include a text representation of the utterance input via the display generation component in response to receiving the text entry input, without entering text into the text entry field.

160. The user interface is the user interface of the application, the text entry field is the text entry field of the application, and the text entry element is a system user interface element. The method according to claim 159, wherein the application has no access to the text representation of the utterance input while the computer system is displaying the text representation of the utterance input contained in the text entry element without entering the text into the text entry field.

161. The user interface is the user interface of the first application, the text entry field is the text entry field of the first application, and the method is The display generation component simultaneously displays a user interface for the second application, which is different from the first application, including a second text entry field for the second application, and a text entry element, which is configured to input text into the second text entry field. While the text entry element and the user interface are being displayed simultaneously via the display generation component, A second text entry input directed to the text entry element via one or more input devices, wherein the second text entry input receives a second text entry input including a second utterance input, The method according to claim 159 or 160, further comprising: updating the display of the text entry element to include a text representation of the second utterance input via the display generation component in response to receiving the second text entry input, without entering text into the second text entry field.

162. The method according to any one of claims 159 to 161, wherein the computer system displays the user interface and the text entry elements in the environment via the display generation component, and the simultaneous display of the user interface and the text entry elements includes displaying the text entry elements between the text entry field of the user interface and the user's viewpoint of the computer system in the environment via the display generation component.

163. Updating the display of the text entry element to include the text representation of the utterance input means that In response to detecting a first portion of the utterance input corresponding to a first amount of text, the display generation component displays the text entry element in a first size according to the first amount of text, The method according to any one of claims 159 to 162, comprising: detecting the first portion of the utterance input and the second portion of the utterance input corresponding to a second amount of text different from the first amount of text, and then displaying the text entry element via the display generation component in a second size different from the first size according to the second amount of text.

164. Updating the display of the text entry element to include the text representation of the utterance input means that In response to detecting the first portion of the utterance input, and in accordance with the determination that the first amount of text corresponds to displaying the text entry element at a third size which includes displaying the text entry element beyond the boundary of the text entry field, the text entry element is to be displayed at a predetermined fourth size which includes displaying the text entry element within the boundary of the text entry field, The method according to claim 163, comprising: displaying the text entry element within the boundaries of the text entry field at a predetermined fourth size, in accordance with the determination that, in response to detecting the first and second portions of the utterance input, the second amount of text corresponds to displaying the text entry element at a fifth size, which includes displaying the text entry element beyond the boundaries of the text entry field.

165. In response to receiving the text entry input, while displaying the text representation of the utterance input within the text entry element, it is detected that the user of the computer system has stopped providing the text entry input via one or more input devices, The method according to any one of claims 159 to 164, further comprising detecting that the user has ceased providing the text entry input, and inputting the text representation of the utterance input into the text entry field.

166. While the text entry element, which includes the text representation of the utterance input, is being displayed, Entering the text representation of the utterance input into the text entry field according to a determination that one or more criteria are met, including criteria that are met in response to the detection of text commit input via one or more input devices, The method according to any one of claims 159 to 165, further comprising: determining that one or more of the above criteria are not met, and ceasing to input the text representation of the utterance input into the text entry field.

167. The method according to claim 166, wherein detecting the commit input includes detecting the user's attention directed to the text representation of the utterance input within the text entry element.

168. The method according to claim 166 or 167, wherein detecting the commit input includes detecting a second utterance input that satisfies one or more second criteria.

169. The method further includes displaying text entry options via the display generation component while the text entry element and the user interface are being displayed simultaneously. The method according to claim 168, wherein the one or more second criteria include criteria that are satisfied when the computer system detects the user's attention directed to the text entry option via the one or more input devices while the computer system is detecting the second utterance input.

170. The method according to any one of claims 166 to 169, further comprising discontinuing the display of the text representation of the utterance input in the text entry element in accordance with the determination that one or more of the above criteria are not met.

171. The user interface is the user interface of the application, and the text entry field is the text entry field of the application. Entering the text representation of the utterance input into the text entry field includes providing the application with access to the text representation of the utterance input, The method according to any one of claims 166 to 170, wherein ceasing to input the text representation of the utterance input into the text entry field includes ceasing to provide the application with access to the text representation of the utterance input.

172. While the user interface including the text entry field and the text entry elements is being displayed simultaneously, a visual indication is displayed via the display generation component that the computer system is configured to input text in response to the utterance input, based on the determination that one or more criteria are met. The method according to any one of claims 159 to 171, further comprising: discontinuing the display of the visual indication that the computer system is configured to input the text in response to the utterance input, in accordance with the determination that one or more of the above criteria are not met.

173. The method according to claim 172, wherein the one or more criteria include criteria that are satisfied in response to detecting an input state corresponding to a user of the computer system who intends to dictate text to be entered into the text entry field.

174. The method according to claim 173, wherein detecting the input state includes detecting the user's attention to the computer system directed to the visual indication that the computer system is configured to input the text in response to the utterance input.

175. The method according to any one of claims 172 to 174, wherein the visual indication is a visual characteristic having a value that changes over time in accordance with a change in the characteristics of the speech input.

176. The method according to any one of claims 159 to 175, wherein displaying the text entry element includes displaying at least a portion of the text entry element in a semi-transparent manner.

177. Displaying the text representation of the utterance input within the text entry element means that Displaying the cursor at a default location for the text representation of the utterance input, Upon receiving the first portion of the text entry input, the cursor is displayed at the first location within the text entry element. The method according to any one of claims 159 to 176, comprising displaying the cursor at a second location different from the first location within the text entry element in response to receiving the first portion and the second portion of the text entry input.

178. The method according to claim 177, wherein displaying the cursor includes displaying the cursor along with a visual indication that the computer system is configured to input text into the text entry element in response to receiving the utterance input.

179. The method according to claim 178, wherein the visual indication is an animated visual feature having a value that changes over time according to the characteristics of the speech input.

180. While the user interface including the text entry field is displayed without displaying the text entry element, The computer system detects that the user's attention is directed to the text entry field via one or more input devices and that one or more criteria are met, and in response to the detection that the user's attention is directed to the text entry field and that one or more criteria are met, the display generation component simultaneously displays the text entry element and the user interface. While the text entry element and the user interface are displayed simultaneously, In response to receiving the aforementioned text entry input, and in accordance with the determination that the user's attention was not directed at the text entry element while the text entry input was detected, the method further includes: The method according to any one of claims 159 to 179, wherein updating the display of the text entry element to include the text representation of the utterance input without entering text into the text entry field in response to receiving the text entry input is in accordance with the determination that the user's attention is directed to the text entry element while the text entry input is being detected.

181. While the user interface including the text entry field is displayed without displaying the text entry element, The detection of whether the user's attention of the computer system is directed to the text entry field via one or more input devices and whether one or more criteria are met, In response to the user's attention being directed to the text entry field and the detection that one or more criteria are met, the display generation component simultaneously displays the text entry element and the user interface. To detect that while the text entry element and the user interface are displayed simultaneously, the user's attention is directed away from the text entry field and one or more second criteria are met, The method according to any one of claims 159 to 180, further comprising: discontinuing the display of the text entry element in response to detecting that the user's attention is directed away from the text entry field and that one or more of the second criteria are met.

182. In response to detecting the aforementioned text entry input, it is detected that while the text representation of the utterance input is displayed within the text entry element, the user's attention is directed away from the text entry field, and one or more of the second criteria are met. The method according to claim 181, further comprising: directing the user's attention away from the text entry field and detecting that one or more of the second criteria are met, ceasing to display the text entry element and the text representation of the utterance input without entering the text into the text entry field.

183. To simultaneously display the user interface including the text entry field and the soft keyboard including the text dictation element via the display generation component, While the user interface and the soft keyboard are being displayed simultaneously via the display generation component, A second text entry input directed to the dictation element via one or more input devices, wherein the second text entry input receives a second text entry input including a second utterance input, Upon receiving the second text entry input, the text representation of the second utterance input is displayed via the display generation component, The further includes receiving input via one or more input devices in response to a request to dictate text into the text entry field while the user interface including the text entry field is displayed, without displaying the soft keyboard and without displaying the text entry element, The method according to any one of claims 159 to 181, wherein the simultaneous display of the user interface and the text entry element is performed in response to the input corresponding to the request to dictate the text into the text entry field, and the simultaneous display of the user interface and the text entry element is performed without displaying the soft keyboard.

184. While simultaneously displaying the text entry field having the text representation of the utterance input and the user interface without the soft keyboard via the display generation component, Detecting that one or more criteria are met, including criteria that are met when the user's attention to the computer system is directed away from the text entry field, via one or more input devices, In response to detecting that one or more of the above criteria are met, the display of the text representation of the speech input is stopped, While the soft keyboard, the user interface, and the text representation of the second speech input are simultaneously displayed via the display generation component, To detect that one or more criteria are met via the one or more input devices, The method according to claim 183, further comprising: detecting that one or more of the above criteria are met; maintaining the display of the text representation of the second speech input.

185. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 159 to 184.

186. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 159 to 184.

187. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 159 to 184.

188. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are The display generation component simultaneously displays a user interface including a text entry field and a text entry element configured to input text into the text entry field. While the text entry element and the user interface are being displayed simultaneously via the display generation component, A text entry input directed to the text entry element via one or more input devices, the text entry input receives a text entry input including a speech input, A non-temporary computer-readable storage medium including instructions for updating the display of the text entry element to include a text representation of the utterance input via the display generation component in response to receiving the text entry input, without entering text into the text entry field.

189. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are The display generation component simultaneously displays a user interface including a text entry field and a text entry element configured to input text into the text entry field. While the text entry element and the user interface are being displayed simultaneously via the display generation component, A text entry input directed to the text entry element via one or more input devices, the text entry input receives a text entry input including a speech input, A computer system including instructions to update the display of a text entry element to include a text representation of an utterance input via a display generation component, in response to receiving the text entry input, without entering text into the text entry field.

190. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A means for simultaneously displaying, via the aforementioned display generation component, a user interface including a text entry field, and a text entry element configured to input text into the text entry field, While the text entry element and the user interface are being displayed simultaneously via the display generation component, A text entry input directed to the text entry element via one or more input devices, the text entry input receives a text entry input including a speech input, A computer system comprising: means for updating the display of the text entry element to include a text representation of the utterance input via the display generation component in response to receiving the text entry input, without entering text into the text entry field.

191. It is a method, In a computer system that communicates with a display generation component and one or more input devices, The display generation component simultaneously displays a soft keyboard including multiple keys and a user interface element including a text representation, wherein the text representation corresponds to the text contained in a text entry field. While the soft keyboard and user interface elements are displayed, Receiving selection input via one or more input devices, In response to receiving the aforementioned selection input, In accordance with the determination that the selection input includes the user's attention directed to a first key among the plurality of keys on the soft keyboard, The display of the text representation is updated via the display generation component to include the first character corresponding to the first key, In accordance with the determination that the selection input includes the user's attention directed to a second key, which is different from the first key among the plurality of keys on the soft keyboard, The display of the text is updated via the display generation component to include a second character corresponding to the second key, wherein the second character is different from the first character. In accordance with the determination that the selection input includes the user's attention directed towards a part of the user interface element, A method comprising updating the display of the representation of the text via the display generation component in order to remove one or more characters from the representation of the text.

192. The method according to claim 191, wherein the part of the user interface element is the end of the expression of the text.

193. The method according to claim 191 or 192, wherein the user interface element includes a cursor displayed in relation to the representation of the text, and the portion of the user interface element is the cursor.

194. Before receiving the selection input, the system detects, via one or more input devices, that the user's attention is directed towards the part of the user interface element. The method according to any one of claims 191 to 193, further comprising detecting that the user's attention is directed to the portion of the user interface element, and displaying a visual indication via the display generating component that the selection of the portion of the user interface element causes the removal of one or more characters from the representation of the text.

195. Upon receiving the selection input, and in accordance with the determination that the selection input includes the user's attention directed to the delete key among the multiple keys on the soft keyboard, The method according to any one of claims 191 to 194, further comprising updating the display of the representation of the text via the display generation component in order to remove one or more characters from the representation of the text.

196. In response to receiving the selection input, and in accordance with the determination that the selection input includes the user's attention directed to a part of the user interface element, the display of the text is updated to remove one or more characters from the text, and then a second selection input is received via one or more input devices, including the user's attention directed to a part of the user interface element. The method according to any one of claims 191 to 195, further comprising updating the display of the representation of the text via the display generation component to remove one or more additional characters from the representation of the text in response to receiving the second selection input.

197. The method according to any one of claims 191 to 196, further comprising, in response to receiving the selection input, ceasing to display the representation of the text via the display generating component in accordance with a determination that the selection input includes the user's attention directed to move away from the soft keyboard.

198. The method of claim 197, further comprising, in response to receiving the selection input, ceasing to display the soft keyboard via the display generation component in accordance with the determination that the selection input includes the user's attention directed to a part of the user interface that does not have a text entry field, wherein the user interface includes the text entry field.

199. The method according to claim 197 or 198, further comprising, in response to receiving the selection input, ceasing to display the representation of the text via the display generation component while maintaining the display of the soft keyboard, in accordance with the determination that the selection input includes the user's attention directed to the second text entry field.

200. Updating the display of the text representation to include the first character corresponding to the first key includes scrolling the text representation in accordance with the determination that the space between the text representation and the default boundary within the user interface element is insufficient to display the first character. The method according to any one of claims 191 to 199, wherein updating the display of the representation of the text to include the second character corresponding to the second key includes scrolling the representation of the text in accordance with the determination that the space between the representation of the text and the default boundary in the user interface element is insufficient to display the second character.

201. While the soft keyboard, the user interface elements, and the text contained in the text entry field are being displayed, Receiving a second input via one or more input devices that corresponds to a request to select a portion of the text contained in the text entry field, Upon receiving the second input, Updating the display of a portion of the text contained in the text entry field so that it is displayed with a first visual characteristic having a first value, wherein before the second input was detected, the portion of the text contained in the text entry field was displayed with a first visual characteristic having a second value different from the first value. The method according to any one of claims 191 to 200, further comprising updating the display of a portion of the text representation corresponding to the portion of the text contained in the text entry field so that it is displayed with a second visual characteristic having a third value, wherein before the second input was detected, the portion of the text representation was displayed with a second visual characteristic having a fourth value different from the third value.

202. The method according to any one of claims 191 to 201, wherein displaying the representation of the text includes displaying a portion of the representation of the text that is within a threshold distance of the boundary of the user interface element having a visual characteristic having a first value, and displaying a portion of the representation of the text that is further than the threshold distance from the boundary of the user interface element having a visual characteristic having a second value different from the first value.

203. Displaying the aforementioned expression in the aforementioned text means In accordance with the determination that the portion of the text within the threshold distance of the boundary of the user interface is currently selected, a visual indication that the visual indication is displayed with the visual characteristic having the first value, along with the display of the portion of the text within the threshold distance of the boundary of the user interface element, The method of claim 202, further comprising: displaying the portion of the text that is further than the threshold distance from the boundary of the user interface element together with the currently selected visual indication, in accordance with the determination that the portion of the text that is further than the threshold distance from the boundary of the user interface element is currently selected, wherein the visual indication is displayed with the visual characteristic having the second value.

204. The method according to any one of claims 191 to 203, wherein displaying the representation of the text includes displaying a portion of the representation of the text having a first orientation with respect to an insertion marker included in the user interface element with a visual characteristic having a first value, and displaying a portion of the representation of the text having a second orientation with respect to the insertion marker with a visual characteristic having a second value different from the first value.

205. Receiving text entry input via one or more input devices, which includes utterance input and the user's attention directed to the text or the expression in the text entry field, Upon receiving the aforementioned text entry input, The display of the text representation is updated to include the first text representation of the utterance input via the display generation component, The method according to any one of claims 191 to 204, further comprising updating the display of the text contained in the text entry field via the display generation component to include a second text representation of the utterance input.

206. Receiving text entry input, including utterance input, via one or more of the aforementioned input devices, Upon receiving the aforementioned text entry input, In accordance with the determination that the text entry input includes the user's attention directed to the text entry field, The display of the text representation is updated to include the first text representation of the utterance input via the display generation component, The display generation component updates the display of the text contained in the text entry field to include the second text representation of the utterance input, In accordance with the determination that the text entry input includes the user's attention directed to the representation of the text, To discontinue updating the display of the text representation to include the first text representation of the utterance input via the display generation component, The method according to any one of claims 191 to 205, further comprising: stopping updating the display of the text contained in the text entry field via the display generation component to include the second text representation of the utterance input.

207. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 191 to 206.

208. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 191 to 206.

209. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 191 to 206.

210. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are The display generation component simultaneously displays a soft keyboard including multiple keys and a user interface element including a text representation, wherein the text representation corresponds to the text contained in a text entry field. While the soft keyboard and user interface elements are displayed, The system receives a selection input via one or more input devices. In response to receiving the aforementioned selection input, In accordance with the determination that the selection input includes the user's attention directed to a first key among the plurality of keys on the soft keyboard, The display of the text representation is updated via the display generation component to include the first character corresponding to the first key. In accordance with the determination that the selection input includes the user's attention directed to a second key, which is different from the first key among the plurality of keys on the soft keyboard, The display of the text is updated via the display generation component to include a second character corresponding to the second key, wherein the second character is different from the first character. In accordance with the determination that the selection input includes the user's attention directed towards a part of the user interface element, A non-temporary computer-readable storage medium including instructions for updating the display of the representation of the text via the display generation component in order to remove one or more characters from the representation of the text.

211. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are The display generation component simultaneously displays a soft keyboard including multiple keys and a user interface element including a text representation, wherein the text representation corresponds to the text contained in a text entry field. While the soft keyboard and user interface elements are displayed, The system receives a selection input via one or more input devices. In response to receiving the aforementioned selection input, In accordance with the determination that the selection input includes the user's attention directed to a first key among the plurality of keys on the soft keyboard, The display of the text representation is updated via the display generation component to include the first character corresponding to the first key. In accordance with the determination that the selection input includes the user's attention directed to a second key, which is different from the first key among the plurality of keys on the soft keyboard, The display of the text is updated via the display generation component to include a second character corresponding to the second key, wherein the second character is different from the first character. In accordance with the determination that the selection input includes the user's attention directed towards a part of the user interface element, A computer system including instructions for updating the display of the representation of the text via the display generation component in order to remove one or more characters from the representation of the text.

212. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A means for simultaneously displaying, via the display generation component, a soft keyboard including a plurality of keys, and a user interface element including a text representation, wherein the text representation corresponds to text contained in a text entry field. While the soft keyboard and user interface elements are displayed, The system receives a selection input via one or more input devices. In response to receiving the aforementioned selection input, In accordance with the determination that the selection input includes the user's attention directed to a first key among the plurality of keys on the soft keyboard, The display of the text representation is updated via the display generation component to include the first character corresponding to the first key. In accordance with the determination that the selection input includes the user's attention directed to a second key, which is different from the first key among the plurality of keys on the soft keyboard, The display of the text is updated via the display generation component to include a second character corresponding to the second key, wherein the second character is different from the first character. In accordance with the determination that the selection input includes the user's attention directed towards a part of the user interface element, A computer system comprising means for updating the display of the representation of the text via the display generation component in order to delete one or more characters from the representation of the text.

213. It is a method, In a computer system that communicates with a display generation component and one or more input devices, The method involves displaying a user interface element containing a text entry field within the environment via the aforementioned display generation component, In accordance with the determination that one or more of the aforementioned input devices have a first location with respect to the environment, the user interface elements are displayed in a second location within the environment that has a first spatial relationship with respect to the hardware input devices. In accordance with the determination that the hardware input device has a third location in the environment that is different from the first location in the environment, the user interface element is displayed at a fourth location in the environment that has the first spatial relationship to the hardware input device. While the user interface elements are displayed in the environment in the first spatial relationship with respect to the hardware input device, The hardware input device receives text entry input, Upon receiving the aforementioned text entry input, A method comprising updating the text entry field to include text corresponding to the text entry input.

214. Displaying the user interface element including the text entry field includes displaying the selectable options included in the user interface element, and the method is The system receives input corresponding to the selection of the selectable options via one or more input devices, The method according to claim 213, further comprising: receiving the input corresponding to the selection of the selectable option, and performing an operation according to the selectable option.

215. The aforementioned selectable options include the indication of the first text, The method according to claim 214, wherein performing the operation according to the selectable options includes updating the text entry field to include the first text.

216. Performing the operation according to the selectable options includes configuring the computer system to accept dictation input directed to the text entry field, and the method is Receiving speech input via one or more input devices, In response to receiving the aforementioned speech input, The determination that the computer system is configured to accept the dictation input directed to the text entry field, and the updating of the text entry field to include a textual representation of the utterance input, The method of claim 214, further comprising: determining that the computer system is not configured to accept the dictation input directed to the text entry field, and then ceasing to update the text entry field to include the text representation of the utterance input.

217. The method according to claim 214, wherein performing the operation according to the selectable options includes displaying a soft keyboard in the environment via the display generation component.

218. The method according to any one of claims 214 to 217, wherein receiving the input corresponding to the selection of the selectable option includes detecting, via one or more input devices, that the default portion of the user is performing a default gesture while the default portion of the user is within a threshold distance of a location corresponding to the selectable option.

219. The method according to any one of claims 214 to 217, wherein receiving the input corresponding to the selection of the selectable option includes detecting, via one or more input devices, that the default portion of the user performs a default gesture while the user's attention to the selectable option is directed to the selectable option and the default portion of the user is farther away from a threshold distance to the location corresponding to the selectable option.

220. The computer system receives selection input via one or more input devices, which includes the default portion of the user performing a default gesture while the user's attention is directed to the selectable option and the default portion of the user is farther away from the location corresponding to the selectable option than a threshold distance; In response to receiving the aforementioned selection input, The operation is performed according to the selectable options, based on the determination that the selected input was received while the hardware input device was not detecting an input. The method according to any one of claims 214 to 219, further comprising: discontinuing to perform the operation according to the selectable option in accordance with the determination that the selection input was received while the hardware input device was detecting the input.

221. The method according to any one of claims 214 to 220, wherein receiving the input corresponding to the selection of the selectable option includes detecting the activation of an element of the hardware input device while the user's attention is directed to the selectable option.

222. The method according to any one of claims 213 to 221, wherein the surface of the hardware input device has a first orientation with respect to the viewpoint of a user of the computer system in the environment, and displaying the user interface element in the environment includes displaying the user interface element having an orientation angle with respect to the viewpoint, wherein the second orientation is different from the first orientation.

223. Displaying the user interface elements based on the determination that, via one or more input devices, the hardware input device is detected within a default area of ​​the environment and that the hardware input device is communicating with the computer system. The method according to any one of claims 213 to 222, further comprising: discontinuing the display of the user interface element in accordance with the failure to detect the hardware input device communicating with the computer system within the default area of ​​the environment via the one or more input devices.

224. The further includes displaying a visual indicator of the status of the hardware input device, In accordance with the determination that the hardware input device has the first location relative to the environment, the visual indication is displayed at a fifth location in the environment that has a second spatial relationship with the hardware input device. The method according to any one of claims 213 to 223, wherein, in accordance with the determination that the hardware input device has the third location with respect to the environment, the visual indication is displayed at a sixth location in the environment that has a second spatial relationship with respect to the hardware input device, different from the fifth location.

225. While the hardware input device has a sixth location in relation to the environment, and the fifth location has a first spatial relationship with respect to the hardware input device, the user interface element is displayed in the fifth location in the environment from the first viewpoint of the user of the computer system via the display generation component, To detect the user's viewpoint shift from the first viewpoint to a second viewpoint different from the first viewpoint, In response to detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, The method according to any one of claims 213 to 224, further comprising maintaining the display of the user interface element at the fifth location in the environment via the display generation component, in accordance with the determination that the hardware input device has the sixth location in the environment.

226. While the user interface element including the text entry field is being displayed, the user interface including a second text entry field having the current focus of the hardware input device is displayed via the display generation component, The method according to any one of claims 213 to 225, further comprising updating the second text entry field to include the text corresponding to the text entry input in response to the receipt of the text entry input.

227. The user interface elements are displayed at a fifth location in the environment via the display generation component, and while the user interface including the second text entry field is displayed, Receiving input via one or more input devices that corresponds to a request to update the location of the user interface, including the second text entry field, The method according to claim 226, further comprising updating the location of the user interface including the second text entry field in the environment while maintaining the display of the user interface element at the fifth location in the environment, in response to receiving the input corresponding to the request to update the location of the user interface including the second text entry field.

228. While the user interface element is displayed at a fifth location in the environment via the display generation component, and while the second text entry field has the current focus of the hardware input device, Receiving input via one or more input devices corresponding to a request to update the current focus of the hardware input device from the second text entry field to the third text entry field, The method according to claim 226 or 227, further comprising updating the current focus of the hardware input device from the second text entry field to the third text entry field in response to receiving the input corresponding to the request to update the current focus of the hardware input device from the second text entry field to the third text entry field, while maintaining the display of the user interface element at the fifth location.

229. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices, wherein the one or more programs include instructions for performing the method according to any one of claims 213 to 228.

230. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 213 to 228.

231. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A computer system comprising means for performing the method described in any one of claims 213 to 228.

232. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, wherein the one or more programs are The display generation component displays a user interface element containing a text entry field within the environment. In accordance with the determination that one or more of the aforementioned input devices have a first location with respect to the environment, the user interface elements are displayed in a second location within the environment that has a first spatial relationship with respect to the hardware input devices. In accordance with the determination that the hardware input device has a third location in the environment that is different from the first location in the environment, the user interface element is displayed in a fourth location in the environment that has the first spatial relationship to the hardware input device. While the user interface elements are displayed in the environment in the first spatial relationship with respect to the hardware input device, The hardware input device receives text entry input, Upon receiving the aforementioned text entry input, A non-temporary, computer-readable storage medium including instructions for updating a text entry field to include text corresponding to the text entry input.

233. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is One or more processors, The system comprises a memory that stores one or more programs configured to be executed by one or more processors, and the one or more programs are The display generation component displays a user interface element containing a text entry field within the environment. In accordance with the determination that one or more of the aforementioned input devices have a first location with respect to the environment, the user interface elements are displayed in a second location within the environment that has a first spatial relationship with respect to the hardware input devices. In accordance with the determination that the hardware input device has a third location in the environment that is different from the first location in the environment, the user interface element is displayed in a fourth location in the environment that has the first spatial relationship to the hardware input device. While the user interface elements are displayed in the environment in the first spatial relationship with respect to the hardware input device, The hardware input device receives text entry input, Upon receiving the aforementioned text entry input, A computer system including instructions for updating a text entry field to include text corresponding to the text entry input.

234. A computer system that communicates with a display generation component and one or more input devices, wherein the computer system is A means for displaying a user interface element including a text entry field within the environment via the aforementioned display generation component, In accordance with the determination that one or more of the aforementioned input devices have a first location with respect to the environment, the user interface elements are displayed in a second location within the environment that has a first spatial relationship with respect to the hardware input devices. The means by which the user interface element is displayed at a fourth location in the environment having a first spatial relationship with the hardware input device, in accordance with the determination that the hardware input device has a third location in the environment that is different from the first location in the environment, While the user interface elements are displayed in the environment in the first spatial relationship with respect to the hardware input device, The hardware input device receives text entry input, Upon receiving the aforementioned text entry input, A computer system comprising means for updating the text entry field to include text corresponding to the text entry input.