DEVICE, METHOD, AND GRAPHICAL USER INTERFACE FOR NAVIGATING AND INPUT OR MODIFYING CONTENT - Patent application
The computer system addresses inefficiencies in VR/AR interaction by using gaze, gesture, and voice inputs to streamline navigation and editing, improving usability and conserving power in battery-operated devices.
Patent Information
- Application Number
- JP2024540029
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-09-24
- Filing Date
- 2023-01-03
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-01-03
AI Technical Summary
Existing methods for interacting with virtual and augmented reality environments are cumbersome, inefficient, and error-prone, leading to cognitive burdens and excessive energy consumption, particularly in battery-operated devices.
Implementing a computer system with improved user interfaces that utilize gaze-based, gesture-based, and voice-based inputs, along with enhanced interaction techniques such as scrollable content management, soft keyboard interaction, and cursor control, to reduce the number and complexity of user inputs.
Enhances user interaction efficiency, reduces errors, conserves power, and extends battery life by minimizing unnecessary inputs and providing intuitive navigation and editing tools within XR environments.
Smart Images

Figure 0007794985000001 
Figure 0007794985000002 
Figure 0007794985000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 266,357, filed January 3, 2022, U.S. Provisional Patent Application No. 63 / 337,539, filed May 2, 2022, and U.S. Provisional Patent Application No. 63 / 377,025, filed September 24, 2022, the contents of which are incorporated by reference herein in their entirety for all purposes.
[0002] The present disclosure relates generally to computer systems that provide computer-generated experiences, including, but not limited to, electronic devices that provide real and mixed reality experiences via display generation components. [Background technology]
[0003] The development of computer systems for augmented reality has progressed significantly in recent years. Exemplary extended reality environments include at least some virtual elements that replace or augment the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / extended reality environment. Exemplary virtual elements include virtual objects such as digital images, video, text, icons, and control elements such as buttons and other graphics. Summary of the Invention
[0004] Some methods and interfaces for navigating and editing content are cumbersome, inefficient, and limited. For example, systems for scrolling content, adding and editing text, and performing actions with a cursor are complex, tedious, and error-prone, creating significant cognitive burdens for users and detracting from the experience of using virtual / augmented reality environments. In addition, these methods take longer than necessary, thereby wasting computer system energy. This latter consideration is particularly important in battery-operated devices.
[0005] Thus, there is a need for a computer system having improved methods and interfaces for scrolling, creating, editing, and navigating content that are more efficient and intuitive for users. Such methods and interfaces, optionally complement or replace conventional methods for performing such operations. Such methods and interfaces reduce the number, extent, and / or type of inputs from a user by helping the user understand the connection between the input provided and the device response to that input, thereby producing a more efficient human-machine interface.
[0006] The above-mentioned drawbacks and other problems associated with user interfaces of computer systems are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generating components, the output devices including one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing a plurality of functions. In some embodiments, a user interacts with the GUI through stylus and / or finger contacts and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body as captured by cameras and other movement sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through the interactions optionally include image editing, drawing, presenting, word processing, creating spreadsheets, playing games, making phone calls, video conferencing, emailing, instant messaging, training support, digital photography, digital videography, web browsing, playing digital music, note taking, and / or playing digital videos, and executable instructions to perform those functions are optionally contained in a transient and / or non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.
[0007] As described above, there is a need for electronic devices with improved methods and interfaces for interacting with content. Such methods and interfaces can complement or replace conventional methods for interacting with content. Such methods and interfaces reduce the number, extent, and / or type of input from a user, creating a more efficient human-machine interface. For battery-operated computing devices, such methods and interfaces conserve power and increase the time between battery charges.
[0008] In some embodiments, the computer system scrolls scrollable content in response to various user inputs. In some embodiments, the computer system enters text into a text entry field in response to voice input. In some embodiments, the computer system facilitates interaction with a soft keyboard. In some embodiments, the computer system facilitates interaction with a cursor. In some embodiments, the computer system facilitates deletion of text from a text entry field. In some embodiments, the computer system facilitates interaction with a hardware input device.
[0009] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification, and claims. Furthermore, it should be noted that the language used in this specification has been selected solely for the purposes of readability and explanation, and not to define or limit the subject matter of the present invention. [Brief explanation of the drawings]
[0010] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:
[0011] [Figure 1] FIG. 1 is a block diagram illustrating an operating environment for a computer system for providing an XR experience, according to some embodiments.
[0012] [Figure 2] FIG. 1 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience, according to some embodiments.
[0013] [Figure 3] FIG. 1 is a block diagram illustrating display generation components of a computer system configured to provide a user with visual components of an XR experience, according to some embodiments.
[0014] [Figure 4] FIG. 1 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input, according to some embodiments.
[0015] [Figure 5]FIG. 1 is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input, according to some embodiments.
[0016] [Figure 6] FIG. 1 is a flow diagram illustrating a glint-assisted gaze tracking pipeline, according to some embodiments.
[0017] [Figure 7A] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. [Figure 7B] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. [Figure 7C] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. [Figure 7D] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. [Figure 7E] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. [Figure 7F] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. [Figure 7G] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. [Figure 7H] 1 illustrates an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments.
[0018] [Figure 8A]FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8B] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8C] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8D] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8E] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8F] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8G] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8H] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8I] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8J] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8K] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. [Figure 8L] FIG. 2 is a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments.
[0019] [Figure 9A]1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9B] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9C] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9D] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9E] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9F] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9G] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9H] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9I] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9J] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9K] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9L] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9M]1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments. [Figure 9N] 1 illustrates an example technique for entering text into a text entry field in response to voice input, according to some embodiments.
[0020] [Figure 10A] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10B] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10C] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10D] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10E] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10F] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10G] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10H] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10I] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10J] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10K] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10L]FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10M] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10N] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10O] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10P] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10Q] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments. [Figure 10R] FIG. 2 is a flow diagram of a method for entering text into a text entry field according to various embodiments.
[0021] [Figure 11A] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11B] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11C] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11D] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11E] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11F] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11G] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11H] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11I] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11J] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11K] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11L] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11M] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11N] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 11O] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments.
[0022] [Figure 12A] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12B] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12C] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12D] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12E] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12F]FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12G] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12H] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12I] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12J] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12K] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12L] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12M] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12N] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12O] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 12P] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments.
[0023] [Figure 13A] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 13B] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 13C] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 13D] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 13E] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments.
[0024] [Figure 14A] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14B] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14C] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14D] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14E] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14F] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14G] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14H] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14I] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 14J] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments.
[0025] [Figure 15A] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 15B]1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 15C] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 15D] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 15E] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 15F] 1 illustrates an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments.
[0026] [Figure 16A] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16B] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16C] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16D] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16E] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16F] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16G] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16H] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16I] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16J] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. [Figure 16K] FIG. 1 is a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments.
[0027] [Figure 17A] 1 illustrates an exemplary technique for facilitating interaction with a cursor, according to some embodiments. [Figure 17B] 1 illustrates an exemplary technique for facilitating interaction with a cursor, according to some embodiments. [Figure 17C] 1 illustrates an exemplary technique for facilitating interaction with a cursor, according to some embodiments. [Figure 17D] 1 illustrates an exemplary technique for facilitating interaction with a cursor, according to some embodiments. [Figure 17E] 1 illustrates an exemplary technique for facilitating interaction with a cursor, according to some embodiments. [Figure 17F] 1 illustrates an exemplary technique for facilitating interaction with a cursor, according to some embodiments.
[0028] [Figure 18A] FIG. 1 is a flow diagram of a method for facilitating interaction with a cursor, according to some embodiments. [Figure 18B] FIG. 1 is a flow diagram of a method for facilitating interaction with a cursor, according to some embodiments. [Figure 18C] FIG. 1 is a flow diagram of a method for facilitating interaction with a cursor, according to some embodiments. [Figure 18D] FIG. 1 is a flow diagram of a method for facilitating interaction with a cursor, according to some embodiments. [Figure 18E] FIG. 1 is a flow diagram of a method for facilitating interaction with a cursor, according to some embodiments.
[0029] [Figure 19A]1 illustrates an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 19B] 1 illustrates an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 19C] 1 illustrates an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 19D] 1 illustrates an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 19E] 1 illustrates an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 19F] 1 illustrates an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 19G] 1 illustrates an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments.
[0030] [Figure 20A] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20B] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20C] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20D] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20E]FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20F] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20G] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20H] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20I] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20J] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20K] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20L] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. [Figure 20M] FIG. 1 is a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments.
[0031] [Figure 21A] 1 illustrates an exemplary technique for modifying text contained in a text entry field, according to some embodiments. [Figure 21B] 1 illustrates an exemplary technique for modifying text contained in a text entry field, according to some embodiments. [Figure 21C]1 illustrates an exemplary technique for modifying text contained in a text entry field, according to some embodiments. [Figure 21D] 1 illustrates an exemplary technique for modifying text contained in a text entry field, according to some embodiments. [Figure 21E] 1 illustrates an exemplary technique for modifying text contained in a text entry field, according to some embodiments. [Figure 21F] 1 illustrates an exemplary technique for modifying text contained in a text entry field, according to some embodiments. [Figure 21G] 1 illustrates an exemplary technique for modifying text contained in a text entry field, according to some embodiments.
[0032] [Figure 22A] FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments. [Figure 22B] FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments. [Figure 22C] FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments. [Figure 22D] FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments. [Figure 22E] FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments. [Figure 22F] FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments. [Figure 22G] FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments. [Figure 22H]FIG. 1 is a flow diagram of a method for modifying text contained in a text entry field according to some embodiments.
[0033] [Figure 23A] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23B] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23C] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23D] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23E] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23F] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23G] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23H] 1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 23I]1 illustrates an exemplary technique for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments.
[0034] [Figure 24A] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24B] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24C] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24D] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24E] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24F] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24G] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24H] FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. [Figure 24I]FIG. 1 is a flow diagram of a method for updating user interface elements according to the status of hardware input devices in communication with a computer system, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0035] The present disclosure relates to a user interface that provides an extended reality (XR) experience to a user, according to some embodiments.
[0036] The systems, methods, and GUIs described herein improve user interface interaction with virtual / augmented reality environments in several ways.
[0037] In some embodiments, the computer system scrolls the content in response to various user inputs, such as gaze-based user inputs and gesture-based user inputs (e.g., air gesture inputs, described in more detail below). In some embodiments, the computer system presents scrollable content including a first region of scrollable content and a second region of scrollable content. In response to detecting the user's attention directed to the second region of the scrollable content, the computer system optionally scrolls the scrollable content to advance content displayed in the second region toward the first region. In some embodiments, the computer system scrolls the content in response to detecting air gesture inputs, including pinch-and-drag gestures, while the user's attention is directed to the content.
[0038] In some embodiments, the computer system enters text into a text entry field in response to speech input according to some embodiments. In response to detecting a user's attention directed at the text entry field, the computer system optionally initiates a process for accepting dictation input directed at the text entry field. The computer system optionally presents feedback (e.g., visual, audio) in response to speech input directed at the text entry field, and displays a text representation of the speech input in the text entry field.
[0039] In some embodiments, a computer system facilitates interaction with a soft keyboard. The computer system optionally displays an object (e.g., a user interface, a window, or another container) including a text entry field that is more than a threshold distance from a user's viewpoint in a three-dimensional environment. In response to input directed at the text entry field, the computer system displays a soft keyboard. In some embodiments, the computer system displays the soft keyboard within the threshold distance of the user.
[0040] In some embodiments, the computer system facilitates interaction with a soft keyboard. The computer system optionally displays the soft keyboard without displaying one or more cursors for interacting with the soft keyboard. In some embodiments, the computer system detects user input directed to one or more keys of the soft keyboard provided by a discrete part of the user (e.g., the user's hand(s)). The computer system optionally displays movement of the one or more keys away from the discrete part of the user toward a surface of the keyboard, and performs one or more actions associated with the one or more keys of the keyboard in response to the user input directed to the one or more keys of the keyboard.
[0041] In some embodiments, the computer system facilitates interaction with a soft keyboard. The computer system optionally displays the soft keyboard along with one or more cursors for interacting with the soft keyboard. The computer system optionally moves the cursor in response to detecting movement of one or more respective parts of the user (e.g., hand(s). In some embodiments, in response to detecting input provided by one or more respective parts of the user corresponding to making a selection with the one or more cursors, the computer system activates one or more keys of the soft keyboard corresponding to the one or more cursors.
[0042] In some embodiments, the computer system facilitates interaction with a cursor. The computer system optionally displays a cursor in a distinct region of the three-dimensional environment. In some embodiments, the computer system updates the position of the cursor according to movement of a distinct part of the user (e.g., a hand) and the user's attention. While the user's attention is directed to a distinct region of the three-dimensional environment and the cursor is displayed in the distinct region of the three-dimensional environment, the computer system moves the cursor within the distinct region in accordance with movement of the distinct part of the user. In some embodiments, in response to detecting coordinated movement of the distinct part of the user and a shift of the user's attention from the distinct region to another location in the three-dimensional environment, the computer system displays the cursor in a new region in accordance with the attention and movement of the distinct part of the user.
[0043] In some embodiments, the computer system facilitates text entry in response to speech input. The computer system optionally displays a dictation user interface element at least partially overlaid on the text entry field to enable dictation of text into the text entry field. In some embodiments, the computer system enters text into the text entry field in response to a confirmation input confirming that the text in the dictation user interface element should be entered into the text entry field. In some embodiments, the computer system refrains from entering text into the text entry field unless and until the confirmation input is received.
[0044] In some embodiments, a computer system facilitates deleting text from a text entry field. The computer system optionally displays a user interface element associated with the soft keyboard, including a text entry field including a copy of text included in a second text entry field in a user interface of an application having a current focus of the soft keyboard. In some embodiments, in response to detecting a user's attention directed to a portion of the text entry field included in the user interface element, the computer system displays an option for deleting one or more characters from the text entry field. In response to detecting selection of the option and / or selection of a portion of the text entry field included in the user interface element, the computer system deletes one or more characters from the text.
[0045] In some embodiments, a computer system facilitates interaction with a hardware input device. The computer system optionally displays user interface elements having a predetermined spatial relationship to a hardware input device that is within a field of view of the computer system and in communication with the computer system. In some embodiments, the user interface elements include a text entry field that includes a representation of text included in a second text entry field of a user interface of an application that has a current focus of the hardware input device, an option for displaying a software input element, a dictation option, and an option for inserting suggested text into the text entry field.
[0046] FIGS. 1-6 illustrate an exemplary computer system for providing an XR experience to a user. FIGS. 7A-7H illustrate an exemplary technique for scrolling scrollable content in response to various user inputs, according to some embodiments. FIGS. 8A-8L are a flow diagram of a method for scrolling scrollable content in response to various user inputs, according to various embodiments. The user interfaces of FIGS. 7A-7H are used to illustrate the process of FIGS. 8A-8L. FIGS. 9A-9N illustrate an exemplary technique for entering text into a text entry field in response to voice input, according to some embodiments. FIGS. 10A-10R are a flow diagram of a method for entering text into a text entry field, according to various embodiments. The user interfaces of FIGS. 9A-9N are used to illustrate the process of FIGS. 10A-10R. FIGS. 11A-11O illustrate an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. FIGS. 12A-12P are a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. The user interfaces of FIGS. 11A-11O are used to illustrate the process of FIGS. 12A-12P. FIGS. 13A-13E illustrate an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. FIGS. 14A-14J are a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. The user interfaces of FIGS. 13A-13E are used to illustrate the process of FIGS. 14A-14J. FIGS. 15A-15F illustrate an exemplary technique for facilitating interaction with a soft keyboard, according to some embodiments. FIGS. 16A-16K are a flow diagram of a method for facilitating interaction with a soft keyboard, according to some embodiments. The user interfaces of FIGS. 15A-15F are used to illustrate the process of FIGS. 16A-16K. FIGS. 17A-17F illustrate an exemplary technique for facilitating interaction with a cursor, according to some embodiments. FIGS. 18A-18E are a flow diagram of a method for facilitating interaction with a cursor, according to some embodiments.The user interfaces of Figures 17A-17F are used to illustrate the processes of Figures 18A-18E. Figures 19A-19G illustrate an exemplary technique for entering text into a text entry field in response to receiving speech input, according to some embodiments. Figures 20A-20M are a flow diagram of a method for entering text into a text entry field in response to receiving speech input, according to some embodiments. The user interfaces of Figures 19A-19G are used to illustrate the processes of Figures 20A-20M. Figures 21A-21G illustrate an exemplary technique for modifying text contained in a text entry field, according to some embodiments. Figures 22A-22H are a flow diagram of a method for modifying text contained in a text entry field, according to some embodiments. The user interfaces of Figures 21A-21G are used to illustrate the processes of Figures 22A-22H. Figures 23A-23I illustrate an exemplary technique for updating user interface elements according to the status of a hardware input device in communication with a computer system, according to some embodiments. 24A-24I are flow diagrams of a method for updating user interface elements according to the status of a hardware input device in communication with a computer system, according to some embodiments. The user interfaces of FIGS. 23A-23I are used to illustrate the processes of FIGS. 24A-24I.
[0047] The processes described below enhance device usability and make user-device interfaces more efficient (e.g., by helping users provide appropriate inputs and reducing user errors when operating / interacting with the device) through various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an action, providing additional control options without cluttering the user interface with additional controls, performing an action without requiring further user input when a set of conditions is met, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and improve device battery life by allowing users to use the device more quickly and efficiently. Saving battery power, and therefore weight, improves device ergonomics. These techniques also enable real-time communication and the use of fewer and / or less accurate sensors, resulting in more compact, lighter, and less expensive devices, and allowing devices to be used in a variety of lighting conditions. These techniques reduce energy use and thereby reduce the heat given off by the device, which is particularly important for wearable devices where a device that is well within the operating parameters for the device components may become uncomfortable for the user to wear if it is generating too much heat.
[0048] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps recited in the claim are repeated in a particular order until the conditions are satisfied and then no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition recited in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions that perform a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.
[0049] 1, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, and / or a touchscreen), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, and / or a velocity sensor), and optionally one or more peripheral devices 195 (e.g., a consumer electronics device and / or a wearable device). In some embodiments, one or more of input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with display generation component 120 (e.g., within a head-mounted or handheld device).
[0050] When describing an XR experience, various terms are used to individually refer to several related, but distinct, environments that a user senses and / or can interact with (e.g., using inputs detected by computer system 101 that cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to computer system 101 generating the XR experience). The following is a subset of these terms:
[0051] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through their senses, such as sight, touch, hearing, taste, and smell.
[0052] Extended reality: In contrast, an extended reality (XR) environment refers to a wholly or partially simulated environment that people sense and / or interact with through electronic systems. In XR, a subset of a person's body movements or representations thereof are tracked, and one or more properties of one or more virtual objects simulated within the XR environment are adjusted accordingly to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and adjust the graphical content and sound field presented to the person accordingly, in a manner similar to how such views and sounds change in a physical environment. In some circumstances (e.g., for accessibility reasons), adjustments to property(ies) of virtual object(s) in the XR environment may be made in response to representations of body movements (e.g., voice commands). A person may sense and / or interact with an XR object using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. In another example, audio objects may enable audio transparency that selectively incorporates ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may sense and / or interact with only audio objects.
[0053] Examples of XR include virtual reality and mixed reality.
[0054] Virtual Reality: A virtual reality (VR) environment refers to a simulated environment designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through a simulation of the person's presence in the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.
[0055] Mixed reality: A mixed reality (MR) environment refers to a simulated environment designed to incorporate sensory input from or representations of a physical environment in addition to including computer-generated sensory input (e.g., virtual objects), as opposed to a VR environment designed to be based entirely on computer-generated sensory input. On a virtual continuum, a mixed reality environment is anywhere between, but not including, a complete physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Some electronic systems for presenting MR environments may also track location and / or orientation relative to the physical environment to allow virtual objects to interact with real objects (i.e., physical items from the physical environment or representations thereof). For example, the system may take into account movement so that a virtual tree appears stationary relative to the physical ground.
[0056] Examples of mixed reality include extended reality and augmented virtuality.
[0057] Extended Reality: An extended reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person uses the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment that are representations of the physical environment. The system composites the images or videos with the virtual objects and presents the composite on the opaque display. The person uses the system to indirectly view the physical environment through the images or videos of the physical environment and perceive the virtual objects superimposed on the physical environment. As used herein, video of a physical environment shown on an opaque display is referred to as “pass-through video,” meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Alternatively, the system may include a projection system that projects virtual objects, e.g., as holograms, into the physical environment or onto a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. An extended reality environment also refers to a simulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system may distort one or more sensor images to impose a selected perspective (e.g., viewpoint) other than the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be distorted by graphically modifying (e.g., enlarging) a portion thereof, such that the modified portion becomes a non-photorealistic, altered version that represents the originally captured image.As a further example, the representation of the physical environment may be altered by graphically removing or obscuring portions of it.
[0058] Augmented Virtuality: An augmented virtuality (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, while people with faces are realistically recreated from images taken of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects may adopt shadows that match the position of the sun in the physical environment.
[0059] Perspective-Locked Virtual Object: A virtual object is perspective-locked when the computer system displays the virtual object in the same location and / or position within the user's perspective, even as the user's perspective shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's perspective is locked to the forward-facing orientation of the user's head (e.g., the user's perspective is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's perspective remains fixed even as the user's line of sight moves without moving the user's head. In embodiments in which the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's perspective is the extended reality view being presented to the user on the display generating component of the computer system. For example, a perspective-locked virtual object displayed in the upper left corner of the user's perspective when the user's perspective is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper left corner of the user's perspective even if the user's perspective changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which a viewpoint-locked virtual object is displayed in a user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."
[0060] Environment-Locked Virtual Object: A virtual object is environment-locked (or "world-locked") when a computer system displays the virtual object at a location and / or position within a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) locations and / or objects within a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint shifts, the locations and / or objects within the environment relative to the user's viewpoint change, resulting in the environment-locked virtual object appearing at a different location and / or position within the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered within the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes more left-leaning within the user's viewpoint (e.g., the position of the tree within the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear more left-leaning within the user's viewpoint. In other words, the location and / or position at which the environment-locked virtual object appears within the user's viewpoint depends on the position and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object in the physical environment) to determine a position at which to display an environment-locked virtual object in the user's viewpoint. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object) or can be locked to a moving portion of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body that moves independent of the user's viewpoint, such as the user's hand, wrist, arm, or leg), so that the virtual object moves as the viewpoint or part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0061] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed-following behavior, which reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed-following behavior, the computer system intentionally delays movement of the virtual object when it detects movement of a reference point that the virtual object is following (e.g., a part of the environment, the viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 and 300 cm from the viewpoint). For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits delayed-following behavior, the device ignores small amounts of movement of the reference point (e.g., ignores movement of the reference point that is less than a threshold amount of movement, such as movement between 0 and 5 degrees or movement between 0 and 50 cm). For example, when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a “delayed following” threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object maintaining a substantially fixed position relative to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / backward relative to the position of the reference point).
[0062] Hardware: There are many different types of electronic systems that allow a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display rather than an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed toward a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser-scanned light source, or any combination of these technologies. The medium may be a light guide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. A projection-based system may employ retinal projection technology that projects graphical images onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or onto physical surfaces. In some embodiments, the controller 110 is configured to manage and coordinate the XR experience for the user.In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Controller 110 is described in more detail below with reference to FIG. 2. In some embodiments, controller 110 is a computing device that is local or remote to scene 105 (e.g., the physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server or a central server) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, and / or a touchscreen) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, and / or IEEE 802.3x). In another example, the controller 110 is contained within the housing (e.g., physical housing) of one or more of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the foregoing.
[0063] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of the XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with reference to FIG. 3. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.
[0064] According to some embodiments, the display generation component 120 provides an XR experience to the user while the user is virtually and / or physically present in the scene 105.
[0065] In some embodiments, the display generating component is worn on a part of the user's body (e.g., the user's head or the user's hand). Thus, display generating component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, display generating component 120 surrounds the user's field of view. In some embodiments, display generating component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, where the user holds the device with a display pointed toward the user's field of view and a camera pointed toward scene 105. In some embodiments, the handheld device is optionally located within a housing worn on the user's head. In some embodiments, the handheld device is optionally located on a support (e.g., a tripod) in front of the user. In some embodiments, display generating component 120 is an XR chamber, housing, or room configured to present XR content without the user wearing or holding display generating component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with XR content triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD in which the interactions occur in the space in front of the HMD and the XR content responses are displayed via the HMD. Similarly, a user interface showing interactions with XR content triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)) may be implemented similarly to an HMD in which the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)).
[0066] While relevant features of operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the exemplary embodiments disclosed herein.
[0067] 2 is a block diagram of an example controller 110, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more pertinent aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0068] In some embodiments, one or more communication buses 204 include circuitry that interconnects and controls communications between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0069] Memory 220 includes high-speed random-access memory, such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240:
[0070] Operating system 230 includes instructions for handling various basic system services and performing hardware-dependent tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To that end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, an adjustment unit 246, and a data transmission unit 248.
[0071] 1 , and optionally one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data acquisition unit 241 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0072] In some embodiments, tracking unit 242 is configured to map scene 105 and track the position / location of at least display generating component 120 relative to scene 105 of FIG. 1 , and optionally relative to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, tracking unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, tracking unit 242 includes hand tracking unit 244 and / or eye tracking unit 243. In some embodiments, hand tracking unit 244 is configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 , relative to display generating component 120, and / or relative to a coordinate system defined relative to the user's hand. Hand tracking unit 244 is described in more detail below with respect to FIG. 4. In some embodiments, eye tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or the user (e.g., the user's hands)), or relative to XR content displayed via display generation component 120. Eye tracking unit 243 is described in more detail below with respect to FIG. 5.
[0073] In some embodiments, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120 and, optionally, by one or more of output devices 155 and / or peripheral devices 195. To that end, in various embodiments, coordination unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0074] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data and / or location data) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0075] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 can be located within separate computing devices.
[0076] Furthermore, Figure 2 is intended more to illustrate the functionality of various features that may be present in particular embodiments, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 2 can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary depending on implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0077] 3 is a block diagram of an example of a display generation component 120, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that, for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward-facing and / or outward-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0078] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, and / or a blood glucose sensor), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.
[0079] In some embodiments, the one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emissive element display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to diffractive, reflective, polarized, holographic, and / or waveguide displays. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content.
[0080] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as eye-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as hand-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene as the user would view it if the display generating component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and / or the like.
[0081] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR presentation module 340:
[0082] The operating system 330 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. To that end, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0083] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, and / or location data) from at least the controller 110 of Figure 1. To that end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0084] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To that end, in various embodiments, the XR presentation unit 344 includes its instructions and / or logic, as well as heuristics and metadata therefor.
[0085] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an extended reality) based on the media content data. To that end, in various embodiments, the XR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0086] In some embodiments, data transmission unit 348 is configured to transmit data (e.g., presentation data and / or location data) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0087] Although the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 may be located in separate computing devices.
[0088] Furthermore, Figure 3 is intended more to illustrate the functionality of various features that may be present in particular implementations, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 3 can be implemented within a single module, and various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0089] 4 is a schematic diagram of an example embodiment of a hand tracking device 140. In some embodiments, hand tracking device 140 (FIG. 1) is controlled by hand tracking unit 244 (FIG. 2) to track the location / position and / or movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 (e.g., relative to a portion of the physical environment surrounding the user, relative to display generating components 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to the user's hand). In some embodiments, hand tracking device 140 is part of display generating components 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, hand tracking device 140 is separate from display generating components 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0090] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images with sufficient resolution to allow for differentiation of the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either zoom capabilities or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.
[0091] In some embodiments, image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to controller 110, which extracts high-level information from the map data. This high-level information is typically provided via an application program interface (API) to an application running on the controller, which drives display generation component 120 accordingly. For example, a user can interact with software running on controller 110 by moving their hand 406 and changing the posture of their hand.
[0092] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the pattern's spots. This approach is advantageous in that it does not require the user to hold or wear any type of beacon, sensor, or other marker. This provides depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In this disclosure, the image sensor 404 is assumed to define a set of orthogonal x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) can use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurement, based on single or multiple cameras or other types of sensors.
[0093] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps including the user's hand while the user moves the hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software matches these descriptors with patch descriptors stored in the database 408, based on a previous learning process, to estimate the pose of the hand in each frame. The pose typically includes the 3D locations of the user's wrist joints and fingertips.
[0094] The software can also analyze hand and / or finger trajectories across multiple frames in a sequence to identify gestures. The pose estimation functionality described herein may be interleaved with motion tracking functionality, whereby patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover pose changes that occur across the remaining frames. Pose, motion, and gesture information is provided to an application program running on controller 110 via the API described above. This program can, for example, move and modify an image presented on display generation component 120 or perform other functions in response to the pose and / or gesture information.
[0095] In some embodiments, the gesture includes an air gesture, which is detected without (or independent of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to another of the user's hands, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of the user's body part (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving a predetermined speed or amount of rotation of the user's body part).
[0096] In some embodiments, input gestures used in various examples and embodiments described herein include air gestures performed by movement of a user's finger(s) relative to other finger(s) or part(s) of the user's hand to interact with an XR environment (e.g., a virtual or mixed reality environment), according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture that includes movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes rotation of a part of the user's body at a predetermined speed or amount).
[0097] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides a computer system with information about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is detected attention (e.g., gaze) to a user interface element in combination with (e.g., simultaneous with) movement of the user's finger(s) and / or hand to perform pinch and / or tap input, as described in more detail below.
[0098] In some embodiments, an input gesture directed at a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly at a user interface object in response to performing an input gesture with the user's hand at a position corresponding to the user interface object's position in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, an input gesture is performed indirectly at a user interface object in response to detecting the user's attention (e.g., gaze) to the user interface object while performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in the three-dimensional environment. For example, in the case of a direct input gesture, a user can direct the user's input at a user interface object by initiating the gesture at or near a position corresponding to the user interface object's displayed position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from an optional outer edge or optional central portion). For indirect input gestures, a user can direct their input to a user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user initiates an input gesture (e.g., at any position detectable by the computer system) (e.g., at a position that does not correspond to the displayed position of the user interface object).
[0099] In some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment, according to some embodiments. For example, pinch inputs and tap inputs, as described below, are performed as air gestures.
[0100] In some embodiments, the pinch input is part of an air gesture, including one or more of a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture, which is an air gesture, includes moving two or more fingers of a hand to contact each other, optionally with a short break (e.g., within 0-1 second) after contact. A long pinch gesture, which is an air gesture, includes moving two or more fingers of a hand to contact each other for at least a threshold time (e.g., at least 1 second) before detecting a break in contact. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., when two or more fingers are touching), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected immediately in succession (e.g., within a predetermined period of time) of each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaking contact between two or more fingers), and performs a second pinch input within a predetermined period of time (e.g., within 1 second or 2 seconds) after releasing the first pinch input.
[0101] In some embodiments, a pinch-and-drag gesture that is an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., followed by) a drag input that changes the position of a user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, a user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers apart) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers together and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by a user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from a first position to a second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of a user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using a first hand of the user and a second pinch input performed using the other hand (e.g., a second of the user's hands) in conjunction with performing the pinch input using the first hand. In some embodiments, a movement between a user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).
[0102] In some embodiments, a tap input (e.g., directed toward a user interface element) performed as an air gesture includes movement(s) of a user's finger(s) toward the user interface element, movement of a user's hand toward a user interface element, optionally with the user's finger(s) extended toward the user interface element, a downward movement of a user's finger (e.g., mimicking a mouse click action or a tap on a touchscreen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture, moving the finger or hand away from the user's viewpoint and / or toward the object that is the target of the tap input followed by an end of the movement. In some embodiments, an end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward the object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand).
[0103] In some embodiments, the user's attention is determined to be directed to a portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the device determines that the user's attention is directed to the portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, and / or requiring the gaze to be directed to the portion of the three-dimensional environment, and if one of the additional conditions is not met, the device determines that the user's attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).
[0104] In some embodiments, detection of a ready configuration of a user or a portion of a user is detected by a computer system, and detection of a ready configuration of the hands is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs performed with the hands (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand geometry (e.g., a pre-pinch geometry with the thumb and one or more fingers extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap geometry with one or more fingers extended and the palm facing away from the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., above the user's waist, moved toward an area in front of the user below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface is responsive to attentional (e.g., gaze) input.
[0105] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or alternatively may be provided on a tangible, non-transitory medium, such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively, or additionally, some or all of the described functionality of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). While the controller 110 is shown in FIG. 4 as, by way of example, a separate unit from the image sensor 404, some or all of the processing functionality of the controller may be implemented by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device), or otherwise associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device) or using any other suitable computerized device, such as a game console or media player. The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that is controlled by the sensor output.
[0106] FIG. 4 also includes a schematic diagram of a depth map 410 captured by the image sensor 404, according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. A pixel 412 corresponding to the hand 406 is segmented from the background and wrist in this map. The intensity of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from the image sensor 404, with increasing gray levels as depth increases. The controller 110 processes these depth values to identify and segment components of the image (i.e., groups of adjacent pixels) that have characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.
[0107] 4 also schematically illustrates a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to some embodiments. In FIG. 4, the hand skeleton 414 is overlaid on a hand background 416 that was segmented from the original depth map. In some embodiments, key feature points on the hand (e.g., knuckles, fingertips, center of the palm, and / or the end of the hand where it connects to the wrist), and optionally the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these key feature points over multiple image frames are used by the controller 110 to determine hand gestures performed by the hand or the current state of the hand, according to some embodiments.
[0108] FIG. 5 shows an exemplary embodiment of eye tracking device 130 ( FIG. 1 ). In some embodiments, eye tracking device 130 is controlled by eye tracking unit 243 ( FIG. 2 ) to track the position and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, eye tracking device 130 is integrated with display generation component 120. For example, in some embodiments, if display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed in a wearable frame, the head-mounted device includes both components for generating XR content for viewing by the user and components for tracking the user's gaze relative to the XR content. In some embodiments, eye tracking device 130 is separate from display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, eye tracking device 130 is optionally a device separate from the handheld device or the XR chamber. In some embodiments, eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, head-mounted eye tracking device 130 is optionally used in conjunction with head-mounted or non-head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally used in combination with head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally part of non-head-mounted display generating components.
[0109] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display that allows the user to view the physical environment directly and display virtual objects on the transparent or translucent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, allowing an individual using the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0110] As shown in FIG. 5 , in some embodiments, eye tracking device 130 (e.g., gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) camera or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR or NIR light from the light source directly from the eyes, or alternatively, may be aimed at a “hot” mirror positioned between the user's eyes and a display panel that reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. Eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate eye tracking information, and communicates the eye tracking information to controller 110. In some embodiments, the user's eyes are tracked separately by their respective eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by a separate eye-tracking camera and lighting source.
[0111] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the eye tracking device's parameters for the particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility prior to delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automatic or manual calibration process. The user-specific calibration process may include estimation of a particular user's eye parameters, such as pupil location, fovea location, optical axis, visual axis, and / or eye spacing. According to some embodiments, once the device-specific and user-specific parameters of the eye tracking device 130 are determined, images captured by the eye tracking camera may be processed using glint-assisted methods to determine the user's current visual axis and gaze point relative to the display.
[0112] As shown in FIG. 5, eye tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and a gaze tracking system including at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking occurs and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 may be positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, and / or a projector) and may be directed at a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (e.g., as shown at the top of FIG. 5), or may be directed at the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of FIG. 5).
[0113] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, such as in processing frames 562 for display. Controller 110 optionally estimates the user's viewpoint on display 510 based on gaze tracking input 542 obtained from eye tracking camera 540, using a glint-assisted method or other suitable method. The viewpoint estimated from gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0114] Some possible use cases of the user's current gaze direction are described below, but are not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content with higher resolution in a central visual area determined from the user's current gaze direction than in a peripheral area. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display particular virtual content within a view based at least in part on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 can orient an external camera to capture the physical environment of the XR experience and focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface within the environment the user is currently viewing on the display 510. As another exemplary use case, eyepiece 520 may be a focusable lens, and eye-tracking information is used by the controller to adjust the focus of eyepiece 520 so that the virtual object the user is currently looking at has the proper binocular coordination to match the convergence of the user's eyes 592. Controller 110 can utilize the eye-tracking information to orient and focus eyepiece 520 so that close objects the user is looking at appear at the correct distance.
[0115] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye tracking camera (e.g., eye tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) attached to the wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.
[0116] In some embodiments, the display 510 emits light in the visible light range and not in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the location and angle of the eye tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0117] Embodiments of an eye tracking system such as that shown in FIG. 5 may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.
[0118] FIG. 6 illustrates a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the glint-assisted gaze tracking system tracks the pupil contour and glint in the current frame using prior information from the previous frame when analyzing the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.
[0119] As shown in FIG. 6, an eye-tracking camera can capture left and right images of a user's left and right eyes. The captured images are then input into an eye-tracking pipeline for processing beginning at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60-120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0120] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. If the tracking status is no at 610, the image is analyzed to detect the user's pupil and glint in the image, as shown at 620. If the pupil and glint are successfully detected at 630, the method proceeds to element 640. If not, the method returns to element 610 to process the next image of the user's eyes.
[0121] At 640, proceeding from element 610, the current frame is analyzed to track pupils and glints based in part on previous information from the previous frame. At 640, proceeding from element 630, a tracking state is initialized based on the detected pupils and glints in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results can be checked to determine whether a sufficient number of glints are successfully tracked or detected in the current frame to perform pupil and gaze estimation. At 650, if the results are not reliable, the tracking state is set to no at element 660 and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes) and the pupil and glint information is passed to element 680 to estimate the user's gaze point.
[0122] 6 is intended to serve as an example of eye-tracking technology that may be used in particular implementations. As will be recognized by those skilled in the art, other eye-tracking technologies, now existing or developed in the future, may be used in place of or in combination with the glint-assisted eye-tracking technology described herein in computer system 101 to provide a user with an XR experience according to various embodiments.
[0123] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, e.g., a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.
[0124] Accordingly, the description herein describes several embodiments of three-dimensional environments (e.g., XR environments) that include representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table present in a physical environment that is captured and displayed within the three-dimensional environment (e.g., actively via a camera and display of the computer system, or passively via a transparent or translucent display of the computer system). As described above, the three-dimensional environment is optionally a mixed reality system based on a physical environment, where the three-dimensional environment is captured by one or more sensors of the computer system and displayed via a display generation component. As a mixed reality system, the computer system can optionally selectively display portions and / or objects of the physical environment such that each portion and / or object of the physical environment appears to exist within the three-dimensional environment displayed by the computer system. Similarly, the computer system can optionally display virtual objects in the three-dimensional environment such that each portion and / or object of the physical environment appears to exist within the real world (e.g., the physical environment) by placing the virtual objects at respective locations within the three-dimensional environment that have corresponding locations in the real world. For example, the computer system optionally displays the vase so that it appears as if the real vase were placed on a table in the physical environment, hi some embodiments, distinct locations in the three-dimensional environment have corresponding locations in the physical environment.Thus, when a computer system is described as displaying a virtual object at a location distinct from a physical object (e.g., at or near the location of a user's hand, or on or near a physical table, etc.), the computer system displays the virtual object at a particular location in the three-dimensional environment so that the virtual object appears to be at or near the physical object in the physical world (e.g., the virtual object is displayed at a location in the three-dimensional environment that corresponds to the location in the physical environment where the virtual object would be displayed if the virtual object were a real object at that particular location).
[0125] In some embodiments, real-world objects present in the physical environment (e.g., and / or visible via display generation components) that are displayed in the three-dimensional environment can interact with virtual objects that exist only in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on the table, where the table is a view (or representation) of the physical table in the physical environment and the vase is a virtual object.
[0126] Similarly, a user can optionally use one or more hands to interact with virtual objects in the three-dimensional environment as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the computer system optionally capture one or more of the user's hands and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, due to the transparency / translucency of the user interface, or the projection of the user interface onto a transparent / translucent surface, or the portion of the display generating components displaying the projection of the user interface to the user's eyes or field of view of the user's eyes, the user's hands are visible through the display generating components by the ability to see the physical environment through the user interface. Thus, in some embodiments, the user's hands are displayed at discrete locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, the computer system can update the display of the representation of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.
[0127] In some of the embodiments described below, the computer system is optionally capable of determining an “effective” distance between a physical object in the physical world and a virtual object in the three-dimensional environment, for example, to determine whether the physical object is directly interacting with the virtual object (e.g., whether a hand is touching, grasping, and / or holding the virtual object, or whether it is within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the fingers of a hand pressing a virtual button, a user's hand grasping a virtual vase, two fingers of a user's hand pinching / holding together an application's user interface, and any other types of interactions described herein. For example, when determining whether and / or how a user is interacting with a virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the target virtual object in the three-dimensional environment. For example, one or more hands of a user are positioned at particular positions in the physical world, which the computer system optionally captures and displays at particular corresponding positions in the three-dimensional environment (e.g., positions in the three-dimensional environment at which the hands are displayed, if the hands are virtual rather than physical hands). The positions of the hands in the three-dimensional environment are optionally compared to positions of target virtual objects in the three-dimensional environment to determine a distance between the user's one or more hands and the virtual objects. In some embodiments, the computer system optionally determines the distance between a physical object and a virtual object by comparing positions in the physical world (e.g., as opposed to comparing positions in the three-dimensional environment).For example, when determining the distance between one or more of a user's hands and a virtual object, the computer system optionally determines the corresponding location in the physical world of the virtual object (e.g., the position where the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and the user's one or more hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the above-mentioned techniques to map the location of the physical object to the three-dimensional environment and / or to map the location of the virtual object to the physical environment.
[0128] In some embodiments, the same or similar techniques are used to determine where and what a user's gaze is directed at and / or where and what a physical stylus held by the user is directed at. For example, if a user's gaze is directed at a particular position in the physical environment, the computer system optionally determines a corresponding position in the three-dimensional environment (e.g., a virtual position of the gaze), and if a virtual object is located at that corresponding virtual position, the computer system optionally determines that the user's gaze is directed at that virtual object. Similarly, the computer system can optionally determine where the physical stylus is pointing in the physical environment based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines a corresponding virtual position in the three-dimensional environment that corresponds to the location in the physical environment where the stylus is pointing, and optionally determines that the stylus is pointing to the corresponding virtual position in the three-dimensional environment.
[0129] Similarly, embodiments described herein may refer to the location of a user (e.g., a user of a computer system) and / or the location of the computer system within a three-dimensional environment. In some embodiments, a user of a computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the location of the user. In some embodiments, the location of the computer system and / or user within the physical environment corresponds to a distinct location within the three-dimensional environment. For example, if a user stands at a location facing a distinct portion of the physical environment visible by the display generating components, the location of the computer system is a location within the physical environment (and its corresponding location within the three-dimensional environment) at which the user views or sees objects within the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the objects are displayed by the display generating components of the computer system within the three-dimensional environment. Similarly, if the virtual objects displayed in the three-dimensional environment were physical objects in the physical environment (e.g., the virtual objects are located in the same physical environment location and have the same physical environment size and orientation as in the three-dimensional environment), the location of the computer system and / or user is the position at which the user would see the virtual objects in the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other and to real-world objects) as they were displayed by the display generation components of the computer system in the three-dimensional environment.
[0130] In this disclosure, various input methods are described with respect to interaction with a computer system. Where one example is provided using one input device or input method and another example is provided using a different input device or input method, it should be understood that each example may be compatible with, and optionally utilize, the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. Where one example is provided using one output device or output method and another example is provided using a different output device or output method, it should be understood that each example may be compatible with, and optionally utilize, the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. Where one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with, and optionally utilize, the method described with respect to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User Interface and Related Processes
[0131] Attention is now directed to embodiments of user interfaces ("UI") and associated processes that may be implemented on a computer system, such as a portable multifunction device or head-mounted device, in communication with display generating components and one or more input devices.
[0132] 7A-7H illustrate exemplary techniques for scrolling scrollable content in response to various user inputs, according to some embodiments. The user interfaces of Figures 7A-7H are used to illustrate processes described below, including the processes of Figures 8A-8L.
[0133] 7A illustrates computer system 101 displaying a three-dimensional environment 701 from a user's perspective via a display generating component (e.g., display generating component 120 of FIG. 1 ). As described above with reference to FIGS. 1-6 , computer system 101 optionally includes a display generating component (e.g., a touchscreen) and multiple image sensors 314 (e.g., image sensors 314 of FIG. 3 ). Image sensors 314 optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that computer system 101 can use to capture one or more images of a user or a portion of a user (e.g., one or more of the user's hands) while the user is interacting with computer system 101. In some embodiments, the user interfaces shown and described below may also be realized on a head-mounted display that includes display generating components that display the user interface or three-dimensional environment to the user, and sensors (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face) for detecting movements of the physical environment and / or the user's hands, such as movements that are interpreted by the computer system as gestures, such as air gestures. It should be understood that in some embodiments, one or more techniques described herein are applied to two-dimensional environments without departing from the scope of this disclosure.
[0134] 7A , computer system 101 presents scrollable content 702 via display generation component 120. In some embodiments, scrollable content 702 includes text content 707 and additional content 705. For example, scrollable content 702 is an article, text content 707 is the text of the article, and additional content 705 is embedded advertisements and / or one or more links to related articles. In some embodiments, scrollable content includes first scrolling region 704 and second scrolling region 706. As described in more detail below, computer system 101 scrolls scrollable content 702 in response to detecting a user's gaze directed toward first scrolling region 704 or second scrolling region 706 without detecting a ready state of the user's hand. In some embodiments, detecting a ready state of the user's hand includes detecting a ready state associated with an air gesture, as described in more detail above. In some embodiments, in response to detecting a user's gaze directed toward an area of scrollable content 702 between scroll regions 704 and 706, the computer system maintains the display of scrollable content 702 without scrolling the scrollable content.
[0135] 7A , in some embodiments, scrolling regions 704 and 706 are adjacent to the boundaries of scrollable content 702. For example, scrollable content 702 is vertically scrollable, such that first scrolling region 704 is at the top of scrollable content 702 and second scrolling region 706 is at the bottom of scrollable content 702. As shown in FIG. 7A , first scrolling region 704 at the top of scrollable content 702 is smaller than second scrolling region 706 at the bottom of scrollable content 702. In some embodiments, if scrollable content 702 were horizontally scrollable, scrollable content 702 would include a left scrolling region and a right scrolling region (e.g., instead of or in addition to an upper scrolling region, such as first scrolling region 704, and a lower scrolling region, such as second scrolling region 706).
[0136] As shown in Figure 7A, computer system 101 detects a user's gaze 713a directed toward second scrolling region 706. In some embodiments, in response to detecting a user's gaze 713a directed toward second scrolling region 706, computer system 101 scrolls scrollable content 702 down, as shown in Figure 7B.
[0137] 7B illustrates how computer system 101 scrolls scrollable content 702 in response to detecting a user's gaze 713a directed toward second scrolling region 706 in FIG. 7A. As shown in FIG. 7B, in response to detecting a user's gaze 713a in FIG. 7A directed toward second scrolling region 706 at the bottom of scrollable content 702, computer system 101 scrolls scrollable content 702 down (e.g., moves scrollable content 702 up to reveal additional scrollable content 702 at the bottom of scrollable content 702). In some embodiments, if the user's gaze was directed toward first scrolling region 704 at the top of scrollable content 702, computer system 101 scrolls scrollable content 702 up (e.g., moves scrollable content 702 down to reveal additional scrollable content 702 at the top of scrollable content 702).
[0138] In some embodiments, the scrolling acceleration and / or speed is different when scrolling up (e.g., in response to detecting a user's gaze directed toward the first scrolling region 704) than when scrolling down (e.g., in response to detecting a user's gaze directed toward the second scrolling region 706). In some embodiments, the scrolling acceleration and / or speed is faster when scrolling up (e.g., in response to detecting a user's gaze directed toward the first scrolling region 704) than when scrolling down (e.g., in response to detecting a user's gaze directed toward the second scrolling region 706). In some embodiments, the scrolling acceleration and / or speed is slower when scrolling up (e.g., in response to detecting a user's gaze directed toward the first scrolling region 704) than when scrolling down (e.g., in response to detecting a user's gaze directed toward the second scrolling region 706).
[0139] In some embodiments, in response to detecting a transition of the user's gaze 713a from not being directed toward one of the scrolling regions 704 or 706 to being directed toward one of the scrolling regions 704 or 706, the computer system 101 gradually increases the scrolling speed of the scrollable content 702 from not scrolling to scrolling at an individual scrolling speed. As described above, the individual scrolling speed is based on which of the two scrolling regions 704 or 706 the user's gaze is directed toward. In some embodiments, the individual scrolling speed is based on the distance from an edge of the scrollable content 702 within the scrolling region 704 or 706 at which the user's gaze is detected. For example, in response to detecting the user's gaze 713a within the second scrolling region 706 at the position shown in FIG. 7A , the computer system 101 scrolls the scrollable content 702 at a first speed. In Figure 7B, computer system 101 detects a user's gaze 713b directed to a different location within second scrolling region 706 that is closer to an edge (e.g., a bottom edge) of scrollable content 702 compared to the location of user's gaze 713a shown in Figure 7A. In some embodiments, in response to detecting user's gaze 713b at a position within second scrolling region 706 shown in Figure 7B, computer system 101 scrolls scrollable content 702 at a faster rate than the rate of scrolling in response to gaze 713a within second scrolling region 706 as shown in Figure 7A.
[0140] Figure 7C illustrates computer system 101 scrolling scrollable content 702 in response to a user's gaze 713b directed to a position within second scroll region 706 shown in Figure 7B. Because user's gaze 713b in Figure 7B is closer to the boundary (e.g., bottom edge) of scrollable content 702 than the location of user's gaze 713a in Figure 7A, the amount of scrolling illustrated in Figure 7C is greater than the amount of scrolling illustrated in Figure 7B.
[0141] In some embodiments, computer system 101 stops scrolling of scrollable content 702 in response to detecting a user's gaze directed toward a portion of scrollable content 702 outside scrolling region 704 or 706, or in response to detecting a ready state of the user's hand while the user's gaze is directed toward one of scrolling regions 704 or 706. For example, FIG. 7C shows a user's gaze 713d directed toward a portion of scrollable content 702 that is not included in first scrolling region 704 or second scrolling region 706. FIG. 7C also shows a user's hand 703a in a ready state (e.g., "hand state A") while the user's gaze 713c is directed toward second scrolling region 706 of scrollable content 702. In response to detecting user's gaze 713d shown in FIG. 7C or the ready state of user's gaze 713c and hand 703a shown in FIG. 7C, computer system 101 stops scrolling of the scrollable content, as shown in FIG. 7D.
[0142] 7D illustrates computer system 101 maintaining display of scrollable content 702 without scrolling scrollable content 702 in response to one of the inputs described above with respect to FIG. 7C. In some embodiments, when stopping scrolling of scrollable content 702, computer system 101 gradually decelerates the scrolling of scrollable content 702 until scrolling is stopped.
[0143] 7D also shows computer system 101 detecting an input provided by a user's hand 703b to scroll scrollable content 702. In some embodiments, the input to scroll scrollable content 702 includes detecting a user's gaze 713e and hand movement (e.g., air gesture, touch input, or other hand input) 703b directed toward the scrollable content 702 while hand 703b is in a pinch hand shape with the thumb touching or within a threshold distance (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, or 1 centimeter) of touching another finger of hand 703b (“Hand State C”). 7D , computer system 101 detects that hand 703b is moving upward while in a pinch hand shape while the user's gaze 713e is directed toward scrollable content 702, and in response, scrolls scrollable content 702 downward (e.g., by moving scrollable content 702 upward to reveal additional scrollable content 702 below scrollable content 702), as shown in FIG. 7E . Although FIG. 7D shows user's gaze 713e directed toward a portion of scrollable content 702 that is not within scroll regions 704 or 706, in some embodiments, the computer system scrolls scrollable content 702 in response to input that includes movement of hand 703b and the user's gaze directed toward one of scroll regions 704 or 706 of scrollable content 702.
[0144] 7E illustrates how computer system 101 updates the display of scrollable content 702 by scrolling scrollable content 702 in response to the input illustrated in FIG. 7D, as described above. In FIG. 7E, computer system 101 detects an input to scroll scrollable content 702 up, provided by user's hand 703c while user's gaze 713f is directed at scrollable content 702. As illustrated in FIG. 7E, computer system 101 detects hand 703c moving down while in a pinch hand shape (e.g., "hand state C") while user's gaze 713f is directed at scrollable content 702. In response to the scrolling input illustrated in FIG. 7E, computer system 101 scrolls scrollable content 702 up (e.g., by moving scrollable content 702 down to reveal additional scrollable content 702 at the top of scrollable content 702), as illustrated in FIG. 7F. Although FIG. 7E shows a user's gaze 713f directed toward a portion of the scrollable content 702 that is not within scroll region 704 or 706, in some embodiments, the computer system scrolls the scrollable content 702 in response to inputs including movement of the hand 703c and the user's gaze directed toward one of the scroll regions 704 or 706 of the scrollable content 702.
[0145] Figure 7F illustrates how computer system 101 updates the display of scrollable content 702 by scrolling scrollable content 702 in response to the input shown in Figure 7E, as described above. In some embodiments, computer system 101 scrolls scrollable content 702 down in response to the scroll input provided by the user's hand by a greater amount than the amount computer system 101 scrolls scrollable content 702 up in response to the scroll input provided by the user's hand for the same amount of hand movement (e.g., air gesture, touch input, or other hand input) in the opposite direction. For example, the amount of movement of hand (e.g., air gesture, touch input, or other hand input) 703b shown in Figure 7D is the same as the amount of movement of hand (e.g., air gesture, touch input, or other hand input) 703c in Figure 7E, but the amount of scrolling of scrollable content 702 in Figure 7E in response to the input in Figure 7D is greater than the amount of scrolling of scrollable content 702 in Figure 7F in response to the input in Figure 7E. In some embodiments, the “amount” of hand movement (e.g., air gesture, touch input, or other hand input) includes the amount of distance, duration, and / or velocity of the hand movement (e.g., air gesture, touch input, or other hand input) while in a pinch configuration and the user's gaze is directed toward the scrollable content 702 to provide scrolling input directed toward the scrollable content 702.
[0146] 7D or 7E , computer system 101 increases the speed of scrolling as the hand moves further from the location where the pinch hand shape was initiated. For example, in response to detecting a first amount of hand movement (e.g., air gesture, touch input, or other hand input) from the hand location when the pinch hand shape was initiated, computer system 101 scrolls scrollable content 702 at a first rate, and optionally continues scrolling at the first rate while the hand remains at the updated location after the first amount of hand movement. In this example, in response to detecting a second amount of hand movement (e.g., air gesture, touch input, or other hand input) greater than the first amount of hand movement from the hand location when the pinch hand shape was initiated, computer system 101 scrolls scrollable content 702 at a second rate greater than the first rate, and optionally continues scrolling at the second rate while the hand remains at the updated location after the second amount of hand movement.
[0147] In some embodiments, computer system 101 scrolls scrollable content 702 in response to detecting hand movement (e.g., air gesture, touch input, or other manual input) in a pinch hand configuration while the user's gaze is directed at the scrollable content 702, according to a determination that the hand movement (e.g., air gesture, touch input, or other manual input) while the hand is in the pinch hand configuration meets one or more criteria. In some embodiments, if the amount of movement (e.g., speed, distance, and / or duration of the movement) is less than a predetermined threshold amount, computer system 101 maintains the display of scrollable content 702 without scrolling the scrollable content 702. Exemplary thresholds are provided below with reference to method 800 and FIGS. 8A-8L. In some embodiments, if the hand movement (e.g., air gesture, touch input, or other manual input) in the pinch configuration is downward and exceeds a threshold speed, computer system 101 maintains the display of scrollable content 702 without scrolling the scrollable content 702. Exemplary threshold speeds are provided below with reference to method 800 and Figures 8A-8L.
[0148] In some embodiments, computer system 101 selects one or more selectable user interface elements displayed via display generation component 120 in response to detecting a user's gaze directed toward the selectable user interface elements while detecting a pinch gesture performed with the user's hand. In some embodiments, the one or more selectable user interface elements are selectable options, representations of content items, application icons, user interface containers (e.g., windows), hyperlinks, etc. Exemplary actions performed in response to selection of these elements include navigating a user interface, presenting an item of content, saving or opening a file or document, initiating communication with another computer system, changing settings on the computer system, updating the current input focus, etc.
[0149] 7G illustrates computer system 101 presenting text content 707 of scrollable content 702 without displaying additional content 705 of scrollable content 702 in a reader mode of computer system 101. The examples illustrated in FIGS. 7A-7F above are examples of computer system 101 presenting scrollable content 702 in a browse mode, including text content 707 and additional content 705. In some embodiments, computer system 101 transitions between displaying the content in a reader mode and displaying the content in a browse mode in response to one or more user inputs.
[0150] In some embodiments, while the computer system 101 is displaying text content 707 of the scrollable content in reader mode, as shown in FIG. 7G , the computer system 101 is configured to scroll the text content 707 according to a user's gaze being directed toward the first scrolling region 708 or the second scrolling region 710, in a manner similar to that described above with reference to FIGS. 7A-7D for the browsing mode. In some embodiments, the computer system 101 is also configured to scroll the text content 707 line by line in response to detecting that the user is reading the text content 707. In some situations, when people read text, after finishing a line of text, they direct their gaze toward the beginning of the next line by moving their gaze along the line from the end of the just-read line to the beginning of the just-read line before looking at the next line. In FIG. 7G , the computer system 101 detects the user's gaze 713h moving from the end of a line of the text content 707 toward the beginning of the line. In response to detecting the movement of gaze 713h shown in Figure 7G, computer system 101 scrolls text content 707 (e.g., by one line) as shown in Figure 7H. In some embodiments, computer system 101 scrolls text content 707 in response to the movement of gaze 713h shown in Figure 7G regardless of whether the user's hand is detected in a ready state.
[0151] Figure 7H illustrates computer system 101 displaying text content 707 after scrolling text content 707 in accordance with the movement of user's gaze 713h shown in Figure 7G. As shown in Figure 7H, in some embodiments, computer system 101 scrolls text content 707 by one line of text content 707 in response to the movement of gaze 713h shown in Figure 7G.
[0152] In some embodiments, computer system 101 displays word definition 712 in response to detecting a user's gaze directed at the word for at least a predetermined threshold time. Exemplary time thresholds are provided below with reference to method 800 and FIGS. 8A-8L. For example, in FIG. 7H, computer system 101 detects a user's gaze 713i directed at the word for the time threshold and, in response, displays word definition 712 overlaid on textual content 707. In some embodiments, computer system 101 similarly displays word definitions while displaying scrollable content 702 including textual content 707 and additional content 705 in the browsing mode shown in FIGS. 7A-7F. Additional description regarding FIGS. 7A-7H is provided below with reference to method 800 described with reference to FIGS. 7A-7H.
[0153] 8A-8L are flow diagrams of methods for scrolling scrollable content in response to various user inputs, according to various embodiments. In some embodiments, method 800 is performed on a computer system (e.g., computer system 101 of FIG. 1) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4). In some embodiments, method 800 is governed by instructions stored on a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1A). Some operations of method 800 are, optionally, combined, and / or the order of some operations is, optionally, changed.
[0154] In some embodiments, method 800 is performed on a computer system (e.g., 101) in communication with a display generation component and one or more input devices (e.g., 314), such as in FIG. 7A (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display (optionally a touchscreen display) integrated with the computer system, an external display such as a monitor, projector, television, or a hardware component (optionally integrated or external) for projecting a user interface or making the user interface visible to one or more users. In some embodiments, the one or more input devices include a computer system or component capable of receiving user input (e.g., capturing and / or detecting user input) and transmitting information associated with the user input to the computer system. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the computer system), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye-tracking device, and / or a motion sensor (e.g., a hand-tracking device and / or a hand motion sensor). In some embodiments, the computer system communicates with a hand-tracking device (e.g., one or more cameras, depth sensors, proximity sensors, and / or touch sensors (e.g., a touchscreen or trackpad)). In some embodiments, the hand-tracking device is a wearable device such as a smart glove. In some embodiments, the hand-tracking device is a handheld input device such as a remote control or a stylus.
[0155] In some embodiments, such as in FIG. 7A , a computer system (e.g., 101) displays (802a) a user interface (e.g., 702) including scrollable content (e.g., 705 or 707) via a display generation component. In some embodiments, the scrollable content includes text and / or images. In some embodiments, the scrollable content exceeds the size of the scrollable user interface element in which the scrollable content is displayed. In some embodiments, in response to a request to scroll the scrollable content, the computer system ceases displaying a first portion of the scrollable content and begins displaying a second portion of the content, optionally while maintaining display of a third portion of the content within the scrollable user interface element. In some embodiments, the scrollable content is displayed within a three-dimensional environment. In some embodiments, the three-dimensional environment includes virtual objects, such as application windows, operating system elements, representations of other users, and / or representations of content items and physical objects in the computer system's physical environment. In some embodiments, representations of physical objects are displayed in the three-dimensional environment via a display generation component (e.g., virtual pass-through or video pass-through). In some embodiments, the representation of the physical object is a view (e.g., true pass-through or actual pass-through) of the physical object within the computer system's physical environment that is visible through a transparent portion of the display generating component. In some embodiments, the computer system displays the three-dimensional environment from the user's perspective at a location within the three-dimensional environment that corresponds to the computer system's physical location within the computer system's physical environment. In some embodiments, the three-dimensional environment is generated, displayed, or otherwise made viewable by a device (e.g., a computer-generated reality (CGR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment).
[0156] In some embodiments, such as FIG. 7A, a computer system (e.g., 101) detects (802b) a user's gaze (e.g., 713a) directed toward scrollable content (e.g., 705 or 707) via one or more input devices (e.g., eye tracking device 314).
[0157] In some embodiments, such as FIG. 7C , in response to detecting a user's gaze (e.g., 713d) directed toward the scrollable content (802c), in accordance with determining that the user's gaze (e.g., 713d) is directed toward a first region of the scrollable content (e.g., 707), the computer system (e.g., 101) maintains display of the scrollable content (e.g., 707) without scrolling the scrollable content (e.g., 707) (e.g., 802d). In some embodiments, the first region of the scrollable content is spaced apart from one or more directions in which the scrollable content is scrollable. For example, if the scrollable content is vertically scrollable, the first region of the scrollable content is a region of the scrollable content between a top and a bottom of the scrollable content. As another example, if the scrollable content is horizontally scrollable, the first region of the scrollable content is a region of the scrollable content between a left portion and a right portion of the scrollable content. In some embodiments, the computer system detects a user's gaze directed at the scrollable content, but the computer system does not detect additional input (e.g., via one or more input devices other than the eye tracking device) corresponding to a request to scroll the content.
[0158] In some embodiments, such as FIG. 7B , in response to detecting (802c) a user's gaze (e.g., 713b) directed toward the scrollable content (e.g., 707), the computer system (e.g., 101) scrolls (802e) the scrollable content (e.g., 707) according to the user's gaze (e.g., 713b) in accordance with a determination that the user's gaze (e.g., 713b) is directed toward a second region (e.g., 706) that is different from the first region of the scrollable content (e.g., 707) and that a distinct part of the user (e.g., a hand or head) meets the respective criteria. In some embodiments, a distinct part of the user meets the respective criteria when the distinct part of the user is in a predefined pose relative to the user's torso or another reference point (e.g., in the three-dimensional environment). For example, the user's hand meets the respective criteria when it is at the user's side, in the user's lap, or otherwise not elevated (e.g., outside a predefined region of the three-dimensional environment with a distinct spatial orientation relative to the user's torso).
[0159] In some embodiments, such as FIG. 7C , in response to detecting (802c) a user's gaze (e.g., 713c) directed toward the scrollable content (e.g., 707), and in accordance with determining that the user's gaze (e.g., 713c) is directed toward the second region (e.g., 706) and that a distinct portion of the user (e.g., 703a) does not meet the respective criteria, the computer system (e.g., 101) maintains display of the scrollable content (e.g., 707) without scrolling the scrollable content (e.g., 707) (802f). In some embodiments, the second region faces one or more directions in which the scrollable content is scrollable. For example, if the scrollable content is vertically scrollable, the second region of the scrollable content is a top region or a bottom region of the scrollable content. As another example, if the scrollable content is horizontally scrollable, the second region of the scrollable content is a left region or a right region of the scrollable content. In some embodiments, the computer system scrolls the scrollable content to reveal a portion of the scrollable content that was not displayed when the user's gaze was (e.g., initially) detected, and displays that portion of the scrollable content in a second region or in a region proximate to the second region. In some embodiments, in response to detecting the user's gaze directed toward a first region of the scrollable content, the computer system scrolls the content in a first direction to reveal a new portion of the content at a location in the first region or in a location proximate to the first region. In some embodiments, as described in more detail below, in response to detecting the user's gaze directed toward a second region of the scrollable content, the computer system scrolls the content in a second direction to reveal a new portion of the content at a location in the second region or in a location proximate to the second region.
[0160] Scrolling scrollable content according to a user's gaze improves user interaction with a computer system by providing an efficient way to navigate scrollable content and reducing the number of inputs required to perform actions (e.g., scrolling according to gaze instead of scrolling according to inputs in addition to or instead of gaze detection).
[0161] In some embodiments, such as FIG. 7B , each criterion includes a criterion that is met when a distinct part of the user (e.g., 703 a) is not detected in a predefined pose (e.g., the user's hands are not in a ready state and / or the user's hands are not visible) (804). In some embodiments, detecting the predefined pose includes detecting a distinct part of the user in a ready state. In some embodiments, the criterion is met when the distinct part of the user is in a rest pose and / or a pose that does not indicate an intention to interact with the computer system. For example, the distinct part of the user is the user's hand, and the criterion is met when the hand is in the user's lap, at the user's side, not within the field of view of the hand tracking device, or possibly not lifted and / or not in a ready state. In some embodiments, while scrolling the scrollable content according to the user's gaze, in response to detecting a distinct part of the user in a predefined pose while the user continues to look at the second region (e.g., detecting a ready state), the computer system stops scrolling the scrollable content.
[0162] Displaying scrollable content without scrolling the scrollable content in response to detecting individual portions of the user in a pose other than a default pose improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0163] In some embodiments, while displaying a user interface including scrollable content (e.g., 707) (806a), a computer system (e.g., 101) detects input via one or more input devices directed at a respective user interface element (e.g., a user interface element within the scrollable content), where detecting the input includes detecting a user's gaze directed at a respective user interface element, such as detecting gaze 713e of FIG. 7D directed at a selectable user interface element and detecting hand 703b performing a respective gesture, and detecting the user performing a respective gesture with a respective part of the user (806b). In some embodiments, the input is an air gesture. In some embodiments, detecting the user performing a respective gesture with a respective part of the user includes detecting the user performing a gesture (e.g., a pinch gesture or a tap gesture) with a hand included in the air gesture input. In some embodiments, the respective part of the user does not meet the respective criteria when the computer system detects the respective gesture. In some embodiments, the input corresponds to a request to select a respective user interface element.
[0164] In some embodiments, while displaying a user interface including scrollable content (806a), in response to detecting input directed to a respective user interface element, the computer system (e.g., 101) performs an action associated with the respective user interface element (806c). In some embodiments, the action associated with the respective user interface element is an action performed in response to detecting selection of the respective user interface element. For example, in response to detecting input directed to an option for navigating to the respective user interface, the computer system presents the respective user interface. As another example, in response to detecting input directed to an option for playing or pausing the content item, the computer system plays or pauses the content item.
[0165] Performing actions associated with individual user interface elements in response to detecting input directed at the individual user interface elements, including detecting the user's gaze and individual gestures with individual parts of the user, enhances user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0166] In some embodiments, such as in FIG. 7A , the second region of the scrollable content (e.g., 706) includes an edge of the scrollable content (e.g., 707) (808). In some embodiments, the second region includes and / or is located proximate to a top edge, bottom edge, left edge, or right edge of the scrollable content. In some embodiments, the second region includes and / or is located at an edge that corresponds to the direction in which the scrollable content is scrollable. For example, the second region includes or is proximate to a top edge or bottom edge of the scrollable content in a vertical direction, or the second region includes or is proximate to a left edge or right edge of the scrollable content in a horizontal direction. Including an edge of the scrollable content in the second region improves user interaction with the computer system by providing additional control options without cluttering the user interface.
[0167] 7A-7B, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a first direction in accordance with a determination that the user's gaze is directed toward a second region (e.g., 706). For example, the computer system scrolls the scrollable content downward in accordance with a determination that the user's gaze is directed toward a region along the bottom of the scrollable content. As another example, the computer system scrolls the scrollable content upward in accordance with a determination that the user's gaze is directed toward a region along the top of the scrollable content.
[0168] In some embodiments, while displaying a user interface (e.g., 702) including scrollable content (e.g., 707) via a display generation component (e.g., 120), in response to detecting a user's gaze directed toward the scrollable content (e.g., 707), the user's gaze is directed toward a third region (e.g., region 704 of FIG. 7B) of the scrollable content, and in accordance with a determination that the third region (e.g., 704) is different from the second region (e.g., 706) and the first region, and individual portions of the user satisfy the respective criteria, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a second direction different from the first direction, such as in FIG. 7F, in accordance with the user's gaze, and the second region (e.g., 706) and the third region (e.g., 704) have different sizes (810b). In some embodiments, the second direction is opposite the first direction, and the third region is disposed along an edge of the scrollable content opposite the edge of the scrollable content along which the second region is disposed. In some embodiments, the second region and the third region have the same size (e.g., width, length, and / or height) along the first direction and different sizes (e.g., width, length, and / or height) along the second direction. For example, the second region and the third region have the same width and different heights.
[0169] Scrolling the scrollable content in different directions depending on whether the user's gaze is directed toward a second region or a third region enhances user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0170] 7A , the second portion (e.g., 706) of the scrollable content (e.g., 707) is located at the bottom of the scrollable content (e.g., 707) and has a first size (e.g., height, width, or length) (812a). In some embodiments, in response to detecting a user's gaze directed toward the second region while a distinct portion of the user meets the respective criteria, the computer system scrolls the scrollable content down.
[0171] In some embodiments, such as in FIG. 7A , a third portion (e.g., 704) of the scrollable content (e.g., 707) is located at the top of the scrollable content (e.g., 707) and has a second size (e.g., height, width, or length) that is smaller than the first size (812b). In some embodiments, in response to detecting a user's gaze directed toward the third region while the user's individual portion meets the respective criteria, the computer system scrolls the scrollable content up. In some embodiments, the height of the third region is smaller than the height of the second region. In some embodiments, the widths of the second region and the third region are the same. In some embodiments, the widths of the second region and the third region are different.
[0172] Providing a third region at the top of the scrollable content that is smaller than the second region of the scrollable content at the bottom of the scrollable content enhances user interaction with the computer system by providing the user with additional control options without cluttering the user interface.
[0173] In some embodiments, scrolling the scrollable content (e.g., 707) according to the user's gaze includes scrolling the scrollable content (e.g., 707) at a first speed (814b) according to the user's gaze, as in FIG. 7B, according to a determination that the user's gaze (e.g., 713a) is directed to a location that is a first distance from the respective position of the scrollable content (e.g., 707), as in FIG. 7A. In some embodiments, the respective position of the scrollable content is a boundary of the second region and / or the start / end of the scrollable content. In some embodiments, the boundary of the second region of the scrollable content is the boundary of the second region or is proximate to the boundary of the second region. For example, if the second region is along the bottom of the scrollable content, the boundary is the bottom region of the scrollable content.
[0174] In some embodiments, scrolling the scrollable content (e.g., 707) according to the user's gaze includes scrolling the scrollable content (e.g., 707) at a second speed (814c) different from the first speed according to the user's gaze, as in FIG. 7C, according to a determination that the user's gaze (e.g., 713b) is directed to a location at a second distance different from the first distance from the individual position of the scrollable content (e.g., 707), as in FIG. 7B. In some embodiments, the scrolling speed increases the closer the gaze is to a boundary of the scrollable content. In some embodiments, the speed of scrolling changes as the user's gaze moves within the second region of the scrollable content. For example, the scrolling speed gradually increases as the user's gaze moves toward the individual position of the scrollable content.
[0175] Scrolling scrollable content at different speeds depending on the distance between the user's gaze and individual positions of the scrollable content improves user interaction with a computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0176] In some embodiments, while the user's gaze (e.g., 713b) is directed toward a second region (e.g., 706) of the scrollable content (e.g., 707), and while individual portions of the user meet the respective criteria, and while scrolling the scrollable content (e.g., 707) according to the user's gaze as in FIG. 7B, the computer system (e.g., 101) detects (816a) via one or more input devices the user's gaze (e.g., 713d) directed away from the second region of the scrollable content as in FIG. 7C. In some embodiments, the computer system detects the user's gaze directed toward the first region of the scrollable content. In some embodiments, the computer system detects the user's gaze directed toward a region of the three-dimensional environment that does not contain scrollable content. In some embodiments, the computer system detects the user's gaze away from the three-dimensional environment or closes their eyes for more than a threshold time associated with blinking (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds).
[0177] In some embodiments, in response to detecting a user's gaze (e.g., 713d) directed away from a second region (e.g., 706) of the scrollable content (e.g., 707), as in FIG. 7C , the computer system (e.g., 101) reduces (816b) the speed at which the scrollable content is scrolling until scrolling of the scrollable content (e.g., 707) stops, as in FIG. 7D . In some embodiments, the computer system stops scrolling the scrollable content in response to detecting a user's gaze directed away from the second region of the scrollable content by slowing down the speed of scrolling with simulated inertia until scrolling stops. In some embodiments, while slowing down the scrolling speed of the scrollable content and continuing to scroll the scrollable content, in response to detecting a user's gaze directed toward the second region of the scrollable content while a distinct portion of the user meets the respective criteria, the computer system accelerates the scrolling speed of the scrollable content. In some embodiments, in this situation, the computer system increases the scrolling speed until the scrolling speed reaches a predetermined speed (e.g., a speed associated with the location within the second region where the user is looking, as described above).
[0178] Slowing down the scrolling of the scrollable content until scrolling stops in response to detecting the user's gaze directed away from the second region of the scrollable content improves user interaction with the computer system by providing the user with improved visual feedback (e.g., indicating to the user that scrolling will stop if the user continues to look away from the second region).
[0179] 7A , in response to detecting a user's gaze (e.g., 713a) directed toward the scrollable content (e.g., 707), and in accordance with a determination that the user's gaze (e.g., 713a) is directed toward the second region (e.g., 706) and a distinct portion of the user meets the respective criteria, scrolling the scrollable content (e.g., 707) according to the user's gaze includes gradually increasing a speed at which the scrollable content (e.g., 707) is scrolled while the user's gaze (e.g., 713a) is directed toward the second region (e.g., 706) and a distinct portion of the user meets the respective criteria (818). In some embodiments, the computer system gradually increases the scrolling speed until the scrolling speed reaches a predetermined speed (e.g., a speed associated with the location within the second region where the user is viewing, as described above). In some embodiments, the computer system gradually decreases the scrolling speed to zero in response to the user directing their gaze from the second region to the first region, as described above. In some embodiments, the computer system gradually changes the scrolling speed in response to the user updating their gaze to a location at a different distance from the edge of the content in the second region.
[0180] Gradually increasing the scrolling speed of the scrollable content in response to detecting the user's gaze directed toward the second region of the scrollable content improves user interaction with the computer system by providing the user with improved visual feedback (e.g., indicating to the user that scrolling will continue if the user continues to look at the second region).
[0181] In some embodiments, while displaying a user interface (e.g., 702) including scrollable content (e.g., 707) (e.g., without scrolling the scrollable content) (820a), the computer system (e.g., 101) detects, via one or more input devices (e.g., hand tracking devices), a discrete portion of a user performing a discrete gesture including moving the user's hand (e.g., 703b) while the user's hand is in a pinch hand shape, such as that of FIG. 7D , and the discrete portion of the user fails to meet a respective criterion while performing the discrete gesture (820b). In some embodiments, the discrete gesture includes detecting that the user is making a pinch shape with their hand (e.g., a hand shape in which the thumb is within or touching a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, or 1 centimeter) of another finger of the hand) and moving their hand while maintaining the pinch shape. In some embodiments, in response to detecting that the user has stopped performing a pinch gesture with their hand, the computer system stops scrolling the scrollable content according to further movement of the hand detected while the hand is not in the pinch shape (e.g., an air gesture, touch input, or other hand input).
[0182] In some embodiments, as in FIG. 7E , in response to detecting that a distinct portion of a user (e.g., 703b) has performed a distinct gesture while displaying a user interface (e.g., 702) including scrollable content (e.g., 707) (e.g., without scrolling the scrollable content), the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) according to the user's hand movement (e.g., air gesture, touch input, or other hand input) (e.g., 703c) in accordance with a determination that one or more criteria are met (820c), while displaying a user interface (e.g., 702) including scrollable content (e.g., 707) (e.g., without scrolling the scrollable content), the computer system scrolls the scrollable content (e.g., 707) according to the user's hand movement (e.g., air gesture, touch input, or other hand input) while the hand is in a pinch shape. For example, the computer system scrolls the content in the same direction as the hand movement while in the pinch shape, by an amount corresponding to the amount of movement (e.g., speed, duration, and / or distance). In some embodiments, while scrolling the scrollable content according to the air gesture input, the computer system does not scroll the scrollable content according to the line of sight. For example, in response to detecting a user's gaze directed toward a second region of the scrollable content while detecting an air gesture input (e.g., corresponding to a request to scroll the scrollable content, corresponding to a different request regarding the scrollable content, or corresponding to a request independent of the scrollable content), the computer system ceases scrolling the scrollable content in accordance with the gaze being directed toward the second region of the scrollable content.
[0183] Scrolling scrollable content according to the user's hand movements (e.g., air gestures, touch input, or other manual input) while the user's hands are in a pinch hand configuration enhances user interaction with a computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0184] In some embodiments, such as FIG. 7D, the movements (eg, speed, distance, and / or duration) of distinct portions (eg, 703b) of the user have distinct magnitudes (822a).
[0185] In some embodiments, in accordance with determining that the movement of the user's individual portion (e.g., 703b) is in a first direction, such as in FIG. 7D, the computer system (e.g., 101) scrolls (822b) the scrollable content (e.g., 707) in a second direction by a first amount in response to detecting that the user's individual portion (e.g., 703b) has performed a individual gesture, such as in FIG. 7E. In some embodiments, the second direction in which the computer system scrolls the scrollable content corresponds to the first direction of movement of the user's individual portion. In some embodiments, the second direction and the first direction are the same direction (e.g., moving the user's individual portion up to scroll up, or moving the user's individual portion down to scroll down). In some embodiments, the second direction and the first direction are opposite directions (e.g., moving the user's individual portion up to scroll down, or moving the user's individual portion down to scroll up). In some embodiments, the first amount corresponds to an individual magnitude. If the individual magnitude is larger, the first amount is larger, and if the individual magnitude is smaller, the first amount is smaller.
[0186] In some embodiments, in accordance with determining that the movement of the user's individual portion (e.g., 703c) is in a third direction different from the first direction, such as in FIG. 7E, the computer system (e.g., 101) scrolls the scrollable content (e.g., 70) by a second amount different from the first amount in a fourth direction in response to detecting that the user's individual portion (e.g., 703c) has performed a individual gesture, the fourth direction being different from the second direction (822c), such as in FIG. 7F. In some embodiments, the fourth direction in which the computer system scrolls the scrollable content corresponds to the third direction of movement of the user's individual portion. In some embodiments, the fourth direction and the third direction are the same direction (e.g., moving the user's individual portion up to scroll up, or moving the user's individual portion down to scroll down). In some embodiments, the fourth direction and the third direction are opposite directions (e.g., moving the user's individual portion up to scroll down, or moving the user's individual portion down to scroll up). In some embodiments, the second amount corresponds to the individual magnitude. If the individual magnitude is larger, the second amount is larger, and if the individual magnitude is smaller, the second amount is smaller. In some embodiments, in response to detecting downward movement of the user's individual portion having the individual magnitude, the computer system scrolls the scrollable content by an amount that is less than the amount that the computer system scrolls the scrollable content in response to detecting upward movement of the user's individual portion having the same individual magnitude.
[0187] Scrolling scrollable content by different amounts in response to movement of separate portions of the user having separate magnitudes in different directions enhances user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0188] 7E , the movement of the user's hand (e.g., air gesture, touch input, or other hand input) (e.g., 703c) includes moving the hand (e.g., air gesture, touch input, or other hand input) (e.g., 703c) from a first location to a second location, and the user's hand (e.g., 703c) maintains a pinch hand shape while moving from the first location to the second location (824a). In some embodiments, the first location is the location of the user's individual parts when the user's individual parts first make the pinch hand shape, such as when the thumb and index finger of the user's hand touch together.
[0189] 7E, scrolling the scrollable content (e.g., 707) in response to detecting a distinct portion of the user (e.g., 703c) performing a distinct gesture includes scrolling the scrollable content (e.g., 707) at a first rate (824c) in accordance with determining that the distance between the first location and the second location is a first distance (824b). In some embodiments, the computer system continues to scroll the scrollable content at the first rate while continuing to detect a predetermined portion of the user at a second location that is the first distance from the first location.
[0190] 7E , scrolling the scrollable content (e.g., 707) in response to detecting that a distinct portion of the user (e.g., 703c) has performed a distinct gesture includes scrolling the scrollable content (e.g., 707) at a second speed greater than the first speed in accordance with a determination that the distance between the first location and the second location is a second distance greater than the first distance (824d). In some embodiments, the computer system continues scrolling the scrollable content at the second speed while continuing to detect a predetermined portion of the user at a second location that is the second distance from the first location. In some embodiments, when the user's hand moves while in the pinch shape, the computer system changes the scrolling speed of the scrollable content according to the distance between the current location of the user's hand and the first location of the user's hand.
[0191] Scrolling scrollable content at a rate that depends on the distance between a first location of a user's hand and a second location of the user's hand improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0192] In some embodiments, the one or more criteria include a criterion that is met when the user's hand (e.g., 703b) moves at least a threshold amount (e.g., velocity (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, 1, 2, 3, or 5 centimeters / second), distance (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 centimeters), and / or duration (e.g., 0.05, 0.1, 0.2, 0.3, 0.5, or 1 second)), such as in FIG. 7D , while maintaining the pinch hand shape (826a).
[0193] 7C , in response to detecting that a distinct portion (e.g., 703a) of the user has performed a distinct gesture, the computer system (e.g., 101) maintains display of the scrollable content (e.g., 707) without scrolling the scrollable content (e.g., 707) according to a determination that the user's hand movement (e.g., air gesture, touch input, or other hand input) (e.g., 703a) does not meet one or more criteria (826b). In some embodiments, if the hand movement (e.g., air gesture, touch input, or other hand input) while the hands are in the pinch shape is an amount less than a threshold, the computer system ceases scrolling the scrollable content according to the hand movement (e.g., air gesture, touch input, or other hand input) while the hands are in the pinch shape.
[0194] Maintaining the display of scrollable content without scrolling the scrollable content in response to detecting a respective portion of a user performing a respective gesture that does not meet one or more criteria because the hand movement (e.g., air gesture, touch input, or other hand input) is less than a threshold amount improves user interaction with the computer system by reducing user errors when interacting with the computer system.
[0195] In some embodiments, one or more criteria are not met 828a when the velocity of the movement of the user's hand (e.g., air gesture, touch input, or other manual input) (e.g., hand 703a in FIG. 7C ) is greater than a threshold velocity (e.g., 1, 2, 3, 5, 10, 15, 30, or 50 centimeters per second) and the direction of the movement of the user's hand (e.g., air gesture, touch input, or other manual input) (e.g., 703a) is downward. In some embodiments, the threshold velocity is associated with the velocity at which the user drops their hand without intending to continue scrolling the scrollable content.
[0196] 7C , in response to detecting that distinct portions of the user have performed distinct gestures, in accordance with a determination that one or more criteria are not met, the computer system (e.g., 101) maintains display of the scrollable content (e.g., 707) without scrolling the scrollable content (e.g., 707) (828b). In some embodiments, the computer system scrolls the scrollable content at a speed less than a threshold speed according to a portion of a downward hand movement (e.g., an air gesture, touch input, or other hand input). For example, if the hand movement (e.g., an air gesture, touch input, or other hand input) includes a first portion of downward movement below a threshold speed and a second portion of downward movement above the threshold speed, the computer system scrolls the scrollable content according to the first portion of downward movement without further scrolling the scrollable content according to the second portion of downward movement.
[0197] Maintaining the display of scrollable content without scrolling the scrollable content in response to detecting a discrete portion of a user performing a discrete gesture that does not meet one or more criteria because the hand movement (e.g., an air gesture, touch input, or other manual input) is downward at a rate that exceeds a threshold rate improves user interaction with the computer system by reducing user errors when interacting with the computer system.
[0198] In some embodiments, in response to detecting a user's gaze (e.g., 713a) directed toward the scrollable content (e.g., 707), and in accordance with a determination that the user's gaze (e.g., 713a) is directed toward a second region (e.g., 706) of the scrollable content (e.g., 707) and that distinct portions of the user meet the respective criteria, the computer system (e.g., 101) scrolls (830a) the scrollable content (e.g., 707) in a first direction according to the user's gaze (e.g., 713a), as in FIG. 7A . In some embodiments, the first direction of scrolling corresponds to the location of the second region of the scrollable content within the scrollable content. For example, if the second region is at the bottom of the scrollable content, the computer system scrolls the content down (e.g., revealing additional content below the scrollable content).
[0199] In some embodiments, while displaying a user interface (e.g., 120) including scrollable content (e.g., 707) via a display generating component (e.g., 120), as in FIG. 7F, in response to detecting a user's gaze directed toward the scrollable content, the computer system (e.g., 101) scrolls the scrollable content (e.g., 707) in a second direction opposite to the first direction according to the user's gaze in accordance with a determination that the user's gaze is directed toward a third region (e.g., region 704 in FIG. 7A) of the scrollable content (e.g., 707), that the third region (e.g., 704) is different from the second region (e.g., 706), and that a distinct portion of the user's gaze meets the respective criteria (830b). In some embodiments, the second direction of scrolling corresponds to the location of the third region of the scrollable content within the scrollable content. For example, if the third region is at the top of the scrollable content, the computer system scrolls the content up (e.g., revealing additional content at the top of the scrollable content). In some embodiments, scrolling the scrollable content in response to detecting the user's gaze directed toward a third region of the content includes scrolling the content along an axis different from the axis along which the computer system scrolls the scrollable content in response to detecting the user's gaze directed toward a second region of the scrollable content. For example, the computer system scrolls the scrollable content vertically in response to detecting the user's gaze directed toward a region along the top or bottom of the content, and scrolls the scrollable content horizontally (e.g., while a distinct portion of the user's gaze meets one or more criteria) in response to detecting the user's gaze directed toward a region along the left or right of the scrollable content.
[0200] Scrolling scrollable content in different directions depending on the area the user's gaze is directed at enhances user interaction with a computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0201] In some embodiments, in response to detecting a user's gaze (e.g., 713a) directed toward the scrollable content (e.g., 707), as in FIG. 7F, scrolling the scrollable content (e.g., 707) in a first direction, such as in FIG. 7B, in accordance with the user's gaze, in accordance with a determination that the user's gaze (e.g., 713a) is directed toward a second region (e.g., 706) of the scrollable content (e.g., 707) and that individual portions of the user satisfy the respective criteria, includes scrolling the scrollable content (e.g., 707) at a first acceleration (832a). In some embodiments, the first direction of scrolling corresponds to the location of the second region of the scrollable content within the scrollable content. For example, if the second region is at the bottom of the scrollable content, the computer system scrolls the content down (e.g., revealing additional content at the bottom of the scrollable content). In some embodiments, the first acceleration is an acceleration at which the computer system begins scrolling the scrollable content in response to detecting a user's gaze directed toward the second region of the scrollable content in accordance with a determination that individual portions of the user satisfy the respective criteria. In some embodiments, the computer system scrolls the scrollable content at a first rate in response to detecting a user's gaze directed toward a second region of the scrollable content while a distinct portion of the user meets the respective criteria.
[0202] In some embodiments, in response to detecting a user's gaze directed toward the scrollable content (e.g., 707), scrolling the scrollable content in a second direction, such as in FIG. 7F , in accordance with a determination that the user's gaze is directed toward a third region (e.g., region 704 in FIG. 7A ) of the scrollable content (e.g., 707) and that a distinct portion of the user meets the respective criteria includes scrolling the scrollable content (e.g., 707) at a second acceleration that is different (e.g., greater than or less than) the first acceleration (832b). In some embodiments, the second direction of scrolling corresponds to the location of the third region of the scrollable content within the scrollable content. For example, if the third region is at the top of the scrollable content, the computer system scrolls the content up (e.g., revealing additional content at the top of the scrollable content). In some embodiments, the second acceleration is an acceleration at which the computer system begins scrolling the scrollable content in response to detecting a user's gaze directed toward the third region of the scrollable content in accordance with a determination that a distinct portion of the user meets the respective criteria. In some embodiments, the computer system scrolls the scrollable content at a second speed different from the first speed mentioned above in response to detecting the user's gaze directed toward a third region of the scrollable content while a distinct portion of the user meets the respective criteria.
[0203] Scrolling scrollable content at different accelerations when the user's gaze is directed to different regions of the scrollable content improves user interaction with a computer system by providing additional control options without cluttering the user interface with displayed controls.
[0204] In some embodiments, such as in FIG. 7A , the scrollable content includes textual content (e.g., 707) and other content (e.g., 705) (e.g., images, interactive content, and / or interactive user interface elements) (834a). In some embodiments, the other content includes additional textual content not included in the textual content of the scrollable content. For example, an article includes textual content including the text of the article and other content including advertisements, including the textual content of the advertisements. In some embodiments, the other content includes multimedia and / or interactive content, such as selectable options for navigating a user interface including the scrollable content (e.g., links to other content). In some embodiments, the computer system displays the scrollable content including the textual content and the other content in a first mode (e.g., a browsing mode) and displays the textual content without the other content in a second mode (e.g., a reader mode). In some embodiments, the computer system transitions between displaying the scrollable content in a first mode and displaying the text content of the scrollable content in a second mode in response to one or more user inputs (e.g., selection of one or more user interface elements, speech input, and / or a predefined gesture performed by a body part of the user) corresponding to a request to change presentation modes.
[0205] 7G, while displaying textual content (e.g., 707) of the scrollable content without displaying other content of the scrollable content (834b), the computer system (e.g., 101) detects (834c) a movement of the user's gaze (e.g., 713h) via one or more input devices. In some embodiments, the movement of the user's gaze corresponds to the user reading the textual content of the scrollable content.
[0206] In some embodiments, in response to detecting (834d) movement of a user's gaze (e.g., 713h) while displaying text content (e.g., 707) of the scrollable content without displaying other content of the scrollable content (834b), the computer system (e.g., 101) scrolls (834e) the text content (e.g., 707) such as in FIG. 7H in accordance with a determination that the movement of the user's gaze (e.g., 713h) satisfies one or more criteria, including criteria satisfied based on movement of the user's gaze (e.g., 713h) relative to a line of text within the text content (e.g., 707) such as in FIG. 7G. In some embodiments, the one or more criteria are associated with the user finishing reading a line of the text content. In some embodiments, the computer system can detect, based on the detected user's eye movement, whether the user is only looking at a first portion of the text or whether the user is reading the first portion of the text item. The computer system optionally compares one or more captured images of the user's eyes to determine whether the user's eye movement matches movements consistent with reading. In some embodiments, people tend to move their gaze from the end of a line they have finished reading to the front of that line, or to the front of the next line after finishing reading a line of text. In some embodiments, the one or more criteria include a criterion that is met when the user's gaze moves from the end of the line toward the beginning of the line or the beginning of the next line. In some embodiments, in response to detecting a movement of the user's gaze corresponding to the user finishing reading a line of text, the computer system scrolls the text content. In some embodiments, the computer system scrolls the text content by one line to display the next line at a location in the three-dimensional environment where the line of text the user just read was displayed while the user was reading the line of text. For example, the computer system scrolls the text vertically to display a separate line of text at the height where the line of text the user previously read was previously displayed.As another example, the electronic device scrolls the text horizontally to display individual lines of text in horizontal locations where lines of text previously read by the user were previously displayed. Scrolling the text content optionally includes updating the location of lines of text previously read by the user (e.g., moving a first portion of the text vertically or horizontally to make space for a second portion of the text) or ceasing to display lines of text previously read by the user. In some embodiments, the computer system scrolls the text content in response to data collected by the eye tracking device without receiving additional input from another input device in communication with the computer system (e.g., air gesture input or input detected via a hardware input device).
[0207] In some embodiments, as in FIG. 7G , in response to detecting a movement of the user's gaze (e.g., 713h) while displaying text content (e.g., 707) of the scrollable content without displaying other content of the scrollable content (834b), the computer system (e.g., 101) maintains display of the text content (e.g., 707) without scrolling the text content (834f) in accordance with a determination that the movement of the user's gaze (e.g., 713h) does not satisfy one or more criteria. In some embodiments, the user's gaze does not satisfy one or more criteria while the user is reading an individual line of the scrollable content (e.g., toward the beginning or center of the line). In some embodiments, the user's gaze does not satisfy one or more criteria when the user's gaze reaches the end of a line of the text content without moving toward the beginning of the line of the text content. For example, the user reads a line of text content and then directs their gaze to another part of the three-dimensional environment that is different from the beginning of the line of the text content or the beginning of the next line of the text content.
[0208] Scrolling textual content according to a determination that a user's gaze movement meets one or more criteria enhances user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0209] In some embodiments, as in FIG. 7G , scrolling the text content (e.g., 707) in response to detecting a user's gaze movement (e.g., 713h) that meets one or more criteria is independent of whether the user's individual portion is detected in a default pose (836). In some embodiments, the user's individual portion is in a default pose when the user's hand is in a ready state. In some embodiments, while the computer system is displaying text content of the scrollable content without displaying additional content of the scrollable content (e.g., in reader mode), the computer system scrolls the text content according to the user's gaze, regardless of the pose and / or location of the user's hand. In some embodiments, the computer system scrolls the text content in response to detecting a user's gaze movement that meets one or more criteria while the user's individual portion meets the respective criteria. In some embodiments, the computer system scrolls the text content in response to detecting a user's gaze movement that meets one or more criteria while the user's individual portion does not meet the respective criteria.
[0210] Scrolling textual content according to a determination that a user's gaze movement meets one or more criteria, regardless of whether individual portions of the user are in a predefined pose, improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0211] In some embodiments, while displaying text content of the scrollable content without other content of the scrollable content (838a) (e.g., while displaying the text content in the reader mode described above), the computer system (e.g., 101) detects (838b) a user's gaze directed at the text content via one or more input devices.
[0212] In some embodiments, while displaying text content of the scrollable content without other content of the scrollable content (838a) (e.g., while displaying the text content in the reader mode described above), in response to detecting a user's gaze directed toward the text content (838c), in accordance with a determination that the user's gaze is directed toward a first region of the text content and that the user's gaze movement does not meet one or more criteria, the computer system (e.g., 101) maintains display of the text content without scrolling the text content (838d). In some embodiments, the first region of the text content is away from one or more directions in which the text content is scrollable. For example, if the text content is vertically scrollable, the first region of the text content is a region of the text content between a top and a bottom of the text content. As another example, if the text content is horizontally scrollable, the first region of the text content is a region of the text content between a left portion and a right portion of the text content. In some embodiments, the first region of the text content is similar to the first region of the scrollable content described above. In some embodiments, the computer system maintains display of the text content without scrolling the text content in response to detecting a user's gaze directed toward a first region of the text content while a movement of the user's gaze does not correspond to the user reading the text content.
[0213] In some embodiments, as in FIG. 7G , while displaying text content (e.g., 707) of the scrollable content without other content of the scrollable content (e.g., while displaying the text content in the reader mode described above) (838a), in response to detecting a user's gaze directed toward the text content (838c), the computer system (e.g., 101) scrolls the text content in accordance with the user's gaze (838e) in accordance with a determination that the user's gaze is directed toward a second region of the text content (e.g., 710) that is different from the first region of the text content, that individual parts of the user (e.g., hands or head) meet the respective criteria (e.g., the user's hands are not in a ready state), and that the user's gaze movement does not meet one or more criteria. In some embodiments, scrolling the text content according to the user's gaze in accordance with a determination that the user's gaze is directed toward the second region of the text content and that the user's individual portions satisfy the respective criteria has one or more characteristics in common with the above-described technique for scrolling the scrollable content in response to detecting the user's gaze directed toward the second region of the scrollable content while the user's individual portions satisfy the respective criteria. In some embodiments, the computer system scrolls the text content in response to detecting the user's gaze directed toward the second region of the text content, the user's gaze movement corresponding to the user reading the text content. In some embodiments, the computer system scrolls the text content in response to detecting the user's gaze directed toward the second region of the text content while the user's gaze movement does not correspond to the user reading the text content. In some embodiments, the second region of the text content is similar to the second region of the scrollable content described above.
[0214] In some embodiments, while displaying text content (e.g., 707) of the scrollable content without other content of the scrollable content (838a) (e.g., while displaying the text content in the reader mode described above), in response to detecting a user's gaze (e.g., 713h) directed toward the text content (838c), as in FIG. 7G, and in accordance with a determination that the user's gaze (e.g., 713h) is directed toward a first region of the text content and that movement of the user's gaze (e.g., 713h) meets one or more criteria, the computer system (e.g., 101) scrolls the text content (838f), as in FIG. 7H. In some embodiments, the computer system scrolls the text content in accordance with the user's gaze being directed toward a second region of the text content and scrolls the text content in accordance with movement of the user's gaze relative to the lines of the text content as described above, while the computer system displays the text content of the scrollable content without other content of the scrollable content as described above. In some embodiments, there are at least two ways to scroll text content based on gaze.
[0215] Scrolling text content according to the user's gaze enhances user interaction with a computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0216] In some embodiments, while displaying scrollable content (e.g., 707), such as in FIG. 7H , in response to detecting a user's gaze (e.g., 713i) directed toward the scrollable content (e.g., 707), the computer system (e.g., 101) displays (840), via the display generation component (e.g., 120), a definition (e.g., 712) of the word included in the scrollable content (e.g., 707) in accordance with a determination that the user's gaze (e.g., 713i) is directed toward a word included in the first region of the scrollable content (e.g., 707) for at least a threshold time (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds). In some embodiments, the word definition is displayed as an overlay on the scrollable content. In accordance with a determination that the user's gaze is directed toward a word included in the first region of the scrollable content for less than the threshold time, the computer system ceases displaying the word definition. In accordance with determining that the user's gaze is not directed at a word included in the first region of the scrollable content, the computer system ceases to display the definition of the word.
[0217] Displaying word definitions according to the user's gaze enhances user interaction with a computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0218] In some embodiments, aspects / operations of methods 1000, 1200, 1400, 1600, 1800, 2000, 2200, and / or 2400 may be interchanged, substituted, and / or added between these methods. For example, the computer system optionally scrolls content generated via speech input according to method 1000 according to one or more steps of method 800. For example, the computer system optionally scrolls content generated via a soft keyboard according to methods 1200, 1400, and / or 1600 according to one or more steps of method 800. For brevity, those details will not be repeated here.
[0219] 9A-9N illustrate exemplary techniques for entering text into a text entry field in response to voice input, according to some embodiments. The user interfaces of Figures 9A-9N are used to illustrate processes described below, including the processes of Figures 10A-10R.
[0220] 9A illustrates computer system 101 displaying a three-dimensional environment 901 from a user's perspective via a display generating component (e.g., display generating component 120 of FIG. 1). As described above with reference to FIGS. 1-6, computer system 101 optionally includes a display generating component (e.g., a touch screen) and multiple image sensors 314 (e.g., image sensors 314 of FIG. 3). Image sensors 314 optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that computer system 101 can use to capture one or more images of a user or a portion of a user (e.g., one or more of the user's hands) while the user interacts with computer system 101. In some embodiments, the user interfaces shown and described below may also be realized on a head-mounted display that includes display generating components that display the user interface or three-dimensional environment to the user, and sensors (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face) to detect movements of the physical environment and / or the user's hands, such as movements that are interpreted by the computer system as gestures, such as air gestures. It should be understood that in some embodiments, one or more techniques described herein are applied to two-dimensional environments (e.g., on a touch-sensitive display or other display) without departing from the scope of this disclosure.
[0221] 9A illustrates computer system 101 displaying web browsing user interface 902 via display generation component 120. In some embodiments, web browsing user interface 902 includes an indication 904 of the URL of a website currently presented by the web browser. For example, in FIG. 9A , web browsing user interface 902 includes a web search website that includes a text entry field 906 to which input specifying one or more search terms is directed and a selectable option 908 that, when selected, causes computer system 101 to perform a search using the one or more search terms provided in text entry field 906. In some embodiments, computer system 101 is configured to detect input for entering text into text entry field 906 via a soft keyboard, via a hardware keyboard, or via dictation according to one or more steps of methods 1200, 1400, and 1600, as described herein.
[0222] In some embodiments, computer system 101 initiates the process of accepting dictation input directed at text entry field 906 in response to detecting a user's attention, including a user's gaze 913a directed at text entry field 906, via one or more input devices (e.g., image sensor 314). In some embodiments, computer system 101 initiates the process of accepting dictation input in response to detecting a user's attention directed at text entry field 906, without or regardless of detecting additional input, such as input provided using an air gesture or a hardware input device. In some embodiments, in response to detecting a user's gaze 913a directed at text entry field 906, computer system 101 gradually expands the text entry field. For example, as shown in FIGS. 9A-9B , in response to a user's gaze 913a being directed at text entry field 906, computer system 101 gradually increases the width of text entry field 906 while user's gaze 913a is directed at text entry field 906. In some embodiments, when the user's gaze 913a is directed toward the text entry field 906 for a threshold time, the computer system 101 stops expanding the text entry field 906 and begins the process of accepting speech input directed toward the text entry field. An exemplary time threshold is provided below in the description of method 1000 with reference to Figures 10A-10R.
[0223] Figure 9B illustrates web browsing user interface 902 updated in response to computer system 101 detecting a user's gaze 913a directed at text entry field 906 for the above-referenced threshold time. As shown in Figure 9B, computer system 101 displays text entry field 906 with a width greater than the width of the text entry field in Figure 9A when computer system 101 first detects that user's gaze 913a is directed at text entry field 906. Figure 9B also illustrates computer system 101 generating audio output 910a indicating that computer system 101 is configured to accept speech input for dictating text directed at text entry field 906 in response to user's gaze 913a being directed at text entry field 906 for the threshold time. The computer system 101 also highlights placeholder text 914 displayed in the text entry field 906 in response to the user's gaze 913a being directed at the text entry field 906 for a threshold time before the computer system 101 detects the user's gaze 913a directed at the text entry field 906.
[0224] 9B shows computer system 101 displaying cursor 912 in response to a user's gaze 913a being directed at text entry field 906 for a threshold time, in some embodiments, computer system 101 does not display cursor 912 unless and until the user provides speech input that dictates text to be entered into text entry field 906. In some embodiments, in response to detecting a user's gaze 913a directed at text entry field 906 for at least the threshold time (e.g., without or regardless of detecting additional input, such as an air gesture or input detected via a hardware input device), computer system 101 displays an additional visual indication that computer system 101 is configured to enter dictated text provided via speech input into text entry field 906, in a manner similar to how computer system 101 displays microphone icon 930 in FIGS. 9G and 9H below.
[0225] In Figure 9B, computer system 101 receives speech input 916a from the user while continuing to detect the user's gaze 913a directed at text entry field 906. In response to the input shown in Figure 9B, computer system 101 displays a text representation of speech input 916a in text entry field 906 for entering the text of the speech input into text entry field 906, as shown in Figure 9C.
[0226] 9C illustrates computer system 101 displaying a text representation 920 of the speech input shown in FIG. 9B in text entry field 906 in response to the input shown in FIG. 9B. In some embodiments, computer system 101 initiates a process of accepting dictation input for entering text into text entry field 906 in response to detecting a user's gaze, without detecting the speech input or regardless of detecting the speech input, as described above with reference to FIGS. 9A-9B. In some embodiments, computer system 101 is configured to accept dictation input for entering text into text entry field 906, but computer system 101 enters text into text entry field 906 in response to the speech input, as shown in FIGS. 9B-9C, without detecting air gesture input and / or input detected via a hardware input device or regardless of detecting the speech input. In some embodiments, while detecting the speech input, computer system 101 generates a glow effect 918 around text entry field 906 that changes over time based on the volume of the received speech input. For example, computer system 101 modifies the size, translucency, color, darkness, or another visual characteristic of glow effect 918 according to the volume of the speech input while the speech input is being received by computer system 101. In some embodiments, computer system 101 displays cursor 912 in text entry field 906 while the speech input is being received.
[0227] In Figure 9C, computer system 101 detects subsequent speech input 916b while user's gaze 913b is no longer directed at text entry field 906. Although Figure 9C shows user's gaze 913b as being directed at an area of the web browsing user interface that does not include text entry field 906, in some embodiments, the user's gaze is directed away from web browsing user interface 902, e.g., toward a portion of display generating component 120 different from the portion of display generating component 120 that includes text entry field 906, or is directed away from display generating component 120. In some embodiments, computer system 101 detects subsequent speech input 916b while the user closes their eyes for longer than a time threshold associated with the user's blink. Exemplary time thresholds are provided below in the description of method 1000 with reference to Figures 10A-10R.
[0228] In some embodiments, in response to detecting follow-up speech input 916b while the user's gaze 913b is not directed toward the text entry field 906, computer system 101 inputs a text-based representation of the follow-up speech input 916b, as described below with reference to FIG. 9D. In some embodiments, in response to detecting follow-up speech input 916b while the user's gaze 913b is not directed toward the text entry field 906, computer system 101 maintains display of the text representation 920 of the previously entered text without displaying the text representation of the follow-up speech input 916b, as described below with reference to FIG. 9D. In some embodiments, in response to detecting follow-up speech input 916b while the user's gaze 913b is not directed toward the text entry field 906, computer system 101 removes (e.g., some or all) the text from the text entry field 906 and stops accepting dictation input directed toward the text entry field 906, as described below with reference to FIG. 9E.
[0229] In some embodiments, in response to detecting a follow-up speech input 916b while the user's gaze 913b is not directed at the text entry field 906, the computer system 101 displays a text representation of the follow-up speech input 916b in the text entry field, as shown in FIG. 9D, if the computer system 101 has already begun accepting dictation input, or ceases displaying the text representation of the follow-up speech input 916b, as shown in FIG. 9E, if the computer system 101 has not already accepted dictation input. In some embodiments, the computer system 101 removes (e.g., some or all) the text from the text entry field 906 and ceases displaying the text representation of the follow-up speech input 916b in the text entry field, as shown in FIG. 9E, because the text entry field 906 is a search text entry field. In some embodiments, a search text entry field is included in a first type of text entry field, which also includes messaging text entry fields and web browser address fields. In some embodiments, if the text entry field is a long-form text entry field, such as the text entry fields shown in FIGS. 9F-9H, and / or requires input in addition to detecting a user's attention directed at the text entry field to accept speech input to provide text in the text entry field, computer system 101 continues to display previously dictated text but does not display a continued text representation of the speech input detected while the user's gaze was not directed at the text entry field.
[0230] 9D illustrates computer system 101 updating text entry field 906 in response to the continuation of the speech input shown in FIG. 9C , according to some embodiments. As described above, in some embodiments, in response to the continuation of the speech input shown in FIG. 9C , computer system 101 maintains the display of text representation 920 of the speech input (e.g., the word "Lorem") in text entry field 906. In some embodiments, computer system 101 also displays the continuation of the speech input (e.g., the word "Ipsum") shown in FIG. 9C . FIG. 9D includes a dashed box around the text representation of the continuation of the speech input 916b (e.g., the word "Ipsum") shown in FIG. 9C because, in some embodiments, computer system 101 ceases to display the text representation of the continuation of the speech input 916b shown in FIG. 9C , as described above. It should be appreciated that in some embodiments, computer system 101 displays the text representation of continued speech input 916b without displaying a dashed box around the text representation of continued speech input 916b. In some embodiments, computer system 101 ceases displaying the text representation of continued speech input 916b and ceases displaying the dashed box. As noted above, in some embodiments, computer system 101 displays the text representation of continued speech input 916b in text entry field 906 of FIG. 9D even though the user's gaze was not directed at the text entry field 906 while continued speech input 916b was detected because dictation had already begun when the continuation of the speech input of FIG. 9C was received.
[0231] 9D, computer system 101 displays glow effect 918 having updated visual characteristics around text entry field 906 in accordance with the change in volume level of speech input 916b shown in FIG. 9C. In some embodiments, computer system 101 displays glow effect 918 when computer system 101 displays the text representation of the continuation of the speech input, and does not display glow effect 918 when computer system 101 ceases displaying the text representation of the continuation of the speech input.
[0232] In some embodiments, while displaying text 920 in text entry field 906, computer system 101 detects speech input 916c corresponding to a command associated with text entry field 906. For example, because text entry field 906 is a search field, speech input 916c includes the word "search." Other examples of speech commands and their associated text entry fields are provided below in the description of method 1000 with reference to FIGS. 10A-10R. In some embodiments, if speech input 916c corresponding to a command is received while gaze 913c is directed at text entry field 906, computer system 101 performs an operation corresponding to text entry field 906, such as performing a search on the search term(s) included in the text entry field when the command was received. In some embodiments, if speech input 916c corresponding to a command is received while gaze 913b is not directed at text entry field 906, computer system 101 refrains from performing the operation corresponding to text entry field 906. In some embodiments, the computer system 101 performs an action corresponding to the text entry field 906, such as performing a search on the search term(s) contained in the text entry field when the command is received, regardless of whether the speech input 916c corresponding to the command is received while the gaze 913c is directed at the text entry field 906 or while the gaze 913b is not directed at the text entry field 906.
[0233] FIG. 9E illustrates computer system 101 updating a text entry field in response to a continuation of the speech input shown in FIG. 9C , according to some embodiments. As described above, in some embodiments, computer system 101 removes text corresponding to the speech input shown in FIG. 9B from text entry field 906 in response to a continuation of the speech input detected while the user's gaze is not directed at text entry field 906 shown in FIG. 9C . In some embodiments, as described above with reference to FIG. 9C and below in the description of method 1000 with reference to FIGS. 10A-10R , computer system 101 removes text corresponding to the speech input shown in FIG. 9B from text entry field 906 because the text entry field is a website search field or another text entry field of the same type as a search field. In some embodiments, by removing the text corresponding to the speech input shown in FIG. 9B from text entry field 906, computer system 101 cancels the dictation input. In some embodiments, the computer system updates the appearance of text entry field 906 to indicate that the dictation entry has been canceled, such as by removing the text from the text entry field or by returning text entry field 906 to the appearance of text entry field 906 in FIG. 9A (e.g., by reducing the width of text entry field 906).
[0234] 9F-9H illustrate computer system 101 displaying word processing user interface 922 including a text entry field 926, a save option 924a, an undo option 924b, a font option 924c, and an option 924d for ceasing display of word processing user interface 922. In some embodiments, text entry field 926 of word processing user interface 922 is a long-form text entry field. In some embodiments, computer system 101 initiates a process for accepting dictation input directed to long-form text entry field 926 in response to additional input to initiate dictation in text entry field 926, such as input detected via an air gesture or a hardware input device. Exemplary inputs are described below in the description of method 1000 with reference to FIGS. 10A-10R. In some embodiments, the computer system initiates dictation in response to detecting a user's attention directed to the text entry field 926, including detecting the user's gaze 913d directed toward the text entry field 926 for a threshold time, without or regardless of receiving additional input, such as input detected via an air gesture or hardware input device. An exemplary threshold time is described below in the description of method 1000 with reference to FIGS. 10A-10R. In some embodiments, before dictation begins, the computer system 101 displays a cursor 928 in the text entry field 926 indicating a location where text will be inserted in response to input provided via a soft keyboard and / or hardware keyboard according to methods 1200, 1400, and / or 1600. Once dictation begins, the computer system 101 ceases displaying the cursor 928, as described with reference to FIGS. 9G-9H.
[0235] FIG. 9G illustrates how computer system 101 updates word processing user interface 922 in response to initiation of dictation. In some embodiments, dictation is initiated based on detecting a user's gaze directed toward text entry field 926, as shown in FIG. 9F, without or regardless of detecting additional input, such as input detected via an air gesture or hardware input device. In some embodiments, dictation is initiated in response to additional input, as described below in the description of method 1000 with reference to FIGS. 10A-10R. As shown in FIG. 9G, when dictation is initiated, computer system 101 generates audio output 910b, which may be the same as or different from audio output 910a described above with reference to FIG. 9B. FIG. 9G also illustrates that computer system 101, in response to detecting speech input provided by the user, displays a microphone icon 930 at a location within text entry field 926 where the dictated text will be inserted. In some embodiments, microphone icon 930 is displayed at a location within the text entry field where the user's gaze was directed when dictation was initiated. Thus, in some embodiments, if the user were to look at a different location within text entry field 926, computer system 101 would display microphone icon 930 at that location instead of the location shown in Figure 9G. Once dictation begins and computer system 101 displays microphone icon 930 at the location to be inserted, as shown in Figure 9G, computer system 101 ceases displaying cursor 928, as shown in Figure 9F. In some embodiments, instead of displaying microphone icon 930 as shown in Figure 9G, computer system 101 displays a different visual indication at the location within text entry field 926 where the dictated text will be inserted.
[0236] 9G, the computer system detects voice input 916d provided by the user while the user's gaze 913d is directed at the text entry field 926. In some embodiments, in response to receiving the voice input 926d, the computer system 101 displays text corresponding to the voice input 926d in the text entry field 926, as shown in FIG.
[0237] 9H shows computer system 101 displaying text 932 corresponding to the voice input shown in FIG. 9G in the text entry field. In some embodiments, while displaying text 932 corresponding to the voice input, the computer system continues to display microphone icon 930 (e.g., if dictation is still active). As shown in FIG. 9H, microphone icon 930 is displayed after text 932 corresponding to the voice input because text corresponding to additional voice input is displayed after text 932 corresponding to the voice input. In some embodiments, if the user's gaze is directed away from text entry field 926 while the user continues to dictate text, computer system 101 maintains the display of text 932 corresponding to the voice input and, optionally, enters text corresponding to subsequent voice entry input detected while the user's gaze was directed away from text entry field 926, as described above in the description of method 1000 with reference to FIGS. 10A-10R and in more detail below, because text entry field 926 is a long-form text entry field.
[0238] 9I-9N illustrate examples of computer system 101 entering text into text entry field 906 in response to voice input. In FIG. 9I, computer system 101 displays web browsing user interface 902 including text entry field 906 in which user input specifying a website address and / or search terms for a web search is accepted. For example, in response to detecting a user entering text into text entry field 906 and subsequently detecting an input to conduct a web search using the text (e.g., selecting a search option, performing a search gesture, and / or a search voice command), computer system 101 initiates a web search for content on the internet that corresponds to the text. As shown in FIG. 9I, text entry field 906 includes placeholder text 934. In some embodiments, placeholder text 934 is displayed in a color that animates a changing hue, darkness, and / or saturation over time in a predetermined pattern or according to the changing audio level of detected sounds (e.g., speech, music, and / or other noises in the environment of computer system 101). In some embodiments, computer system 101 displays placeholder text 934 in text entry field 906 before receiving input to enter text into the text entry field. In some embodiments, computer system 101 displays placeholder text 934 in text entry field 906 in response to receiving one or more inputs corresponding to a request to delete existing text from the text entry field, such as the URL of website A currently displayed in internet browsing user interface 902.
[0239] 9I, text entry field 906 is displayed with a background that does not change color according to the changing audio levels of detected audio (e.g., ambient noise or speech) and is displayed without any visible glowing around the edges of text entry field 906. In some embodiments, this appearance of text entry field 906 shown in FIG. 9I indicates that the computer system, in response to receiving speech input, does not enter text corresponding to the speech input into text entry field 906. For example, if the user speaks one or more words while computer system 101 is displaying text entry field 906 as shown in FIG. 9I, computer system 101 maintains the display of placeholder text 934 in text entry field 906.
[0240] FIG. 9I illustrates a dictation icon 936 included in the text entry field 906. In some embodiments, the computer system 101 displays the dictation icon 936 in the text entry field 906 in response to detecting a user's attention directed to the text entry field 906, as described above. As shown in FIG. 9I, the computer system 101 detects a user's attention 913e directed to the dictation icon 936. In response to detecting a user's attention directed to the dictation icon 936 in FIG. 9I, the computer system updates the appearance of the text entry field 906 and enters text corresponding to the speech input in response to detecting the speech input, as described with reference to at least FIG. 9J.
[0241] FIG. 9J illustrates computer system 101 displaying text entry field 906 with an updated appearance in response to detecting user attention 913e directed to dictation icon 936, as described above with reference to FIG. 9I. In some embodiments, as shown in FIG. 9I, computer system 101 updates text entry field 906 to include dictation icon 938 in a location within text entry field 906 that is different from the location shown in FIG. 9I. In some embodiments, as shown in FIG. 9J, computer system 101 updates text entry field 906 to appear with a background that changes color according to the changing audio levels of detected audio, including voice input 916e. In some embodiments, as shown in FIG. 9J, computer system 101 updates text entry field 906 to appear with a glowing outline 942a that changes color, intensity, and / or radius according to the changing audio levels of detected audio, including voice input 916e. In some embodiments, as shown in FIG. 9J, the computer system 101 updates the text entry field 906 to include an insertion marker 944a that changes color and / or has a glowing effect that changes color, intensity, and / or radius according to the changing audio level of the detected audio, including the voice input 916e.
[0242] In some embodiments, while displaying text entry field 906 with the appearance shown in Figure 9J, computer system 101 receives speech input 916e while user's attention 913e is directed to text entry field 906. In some embodiments, in response to receiving speech input 916e while displaying text entry field 906 with the appearance shown in Figure 9J and while user's attention 913e is directed to text entry field 906, computer system 101 enters text corresponding to the speech input into text entry field 906, as shown in Figure 9K. In some embodiments, if user's attention 913e is not directed to text entry field 906 while computer system 101 detects speech input 916e, computer system 101 refrains from entering text corresponding to speech input 916e into the text entry field. In some embodiments, in response to detecting the user's attention being directed away from the text entry field 906, the computer system 101 ceases displaying the text entry field 906 with the appearance shown in FIG. 9J and displays the text entry field 906 with the appearance shown in FIG. 9I.
[0243] 9K and 9L illustrate that computer system 101, in response to speech input 916e described above with reference to FIG. 9J, enters text corresponding to speech input 916e of FIG. 9J. In some embodiments, computer system 101 animates entering the text character by character, as shown in FIG. 9K and 9L. As shown in FIG. 9K, while entering the text corresponding to the speech input, computer system 101 updates the background color of text entry field 906, the glow effect 942b around text entry field 906, and / or the color and / or glow of insertion marker 912 according to the detected audio level (e.g., of speech input 916e). FIG. 9K illustrates that while entering the text corresponding to speech input 916e, computer system 101 displays a first portion 946a of text corresponding to speech input 916e in a first color and a second portion 948a of text corresponding to speech input 916e in a second color and / or glow effect. For example, when computer system 101 displays an additional character corresponding to speech input 916e, computer system 101 displays the character with a color and / or glow effect that varies according to the detected audio level, and then transitions to a solid first color. In some embodiments, the color of the background of text entry field 906, the glow effect 942b around text entry field 906, the color and / or glow of insertion marker 912, and the color of second portion 948a vary coordinately according to the detected audio level.
[0244] Figure 9L illustrates the continued entry of text in response to speech input 916e shown in Figure 9J. As shown in Figure 9L, as the detected audio level continues to change (e.g., the electronic device continues to detect speech input 916e), computer system 101 updates the background color of text entry field 906, the glow effect 942c around text entry field 906, and / or the color and / or glow of insertion marker 912 according to the detected audio level (e.g., of speech input 916e). As shown in Figure 9L, as computer system 101 adds characters to the entered text, portion 946b of text displayed in a solid color includes the additional characters and displays characters 948b in a color and / or glow that corresponds to the audio level before the characters were displayed in a solid color as they were added to text entry field 906.
[0245] In some embodiments, when computer system 101 no longer detects speech input 916e, computer system 101 displays text entry field 906 with the text corresponding to the speech input with the appearance shown in FIG. 9I. For example, computer system 101 displays text entry field 906 with a solid background color that remains the same regardless of the detected audio level, ceases displaying a glowing effect around text entry field 906, and ceases displaying insertion marker 912 in text entry field 906. In some embodiments, while displaying the text corresponding to speech input 916e in the text entry field, computer system 101 receives input corresponding to a request to conduct an Internet search based on the text in the text entry field. In some embodiments, in response to the input, computer system 101 displays search results related to the text in the text entry field (e.g., the text corresponding to speech input 916e).
[0246] In some embodiments, computer system 101 enters text into text entry field 906 in response to one or more typed text entry inputs, but computer system 101 displays text entry field 906 with the appearance shown in FIG. 91 instead of the appearance shown in FIGS. 9J-9L. For example, FIGS. 9M-9N show examples of computer system 101 entering text into text entry field 906 in response to input received using soft keyboard 950. In some embodiments, computer system 101 similarly enters text into text entry field 906 in response to input received using a hardware keyboard. In some embodiments, computer system 101 enters text into text entry field 906 in response to input directed at a soft keyboard according to one or more steps of method(s) 1200, 1400, 1600, and / or 2200. In some embodiments, computer system 101 enters text into text entry field 906 in response to input directed at a hardware keyboard according to one or more steps of method 2400.
[0247] In FIG. 9M, computer system 101 simultaneously displays text entry field 906 along with soft keyboard 950. In some embodiments, soft keyboard 950 is displayed with option 954 that, when selected, causes computer system 101 to enter text into text entry field 906 in response to speech input. In some embodiments, as shown in FIG. 9M, computer system 101 displays text entry field 906 with a background color that does not change according to the detected audio level without a glow effect. In some embodiments, text entry field 906 does not include a dictation icon. While FIG. 9M shows text entry field 906 displayed without an insertion marker, in some embodiments, text entry field 906 includes an insertion marker. In FIG. 9M, computer system 101 receives input directed at soft keyboard 950 with hand 903. In response to the input shown in FIG. 9M, computer system 101 enters text corresponding to the input directed at the soft keyboard, as shown in FIG. 9N.
[0248] 9N illustrates computer system 101 displaying text entry field 906 with text 952 corresponding to the input shown in FIG. 9M. In some embodiments, computer system 101 displays text 952 in a color that does not change over time and / or in response to detected audio levels as the computer system enters text 952 and / or after computer system 952 enters the text. In some embodiments, during and after entering text 952 in text entry field 906, computer system 101 displays text entry field 906 with a background color that does not change in response to detected audio levels without a glow effect. In some embodiments, while displaying text 952 in text entry field 906, computer system 101 receives input corresponding to a request to conduct an Internet search based on text 952 in text entry field 906 and, in response to the input, displays search results that correspond to text 952.
[0249] Additional description regarding FIGS. 9A-9N is provided below with reference to the method 1000 described with respect to FIGS. 9A-9N.
[0250] 10A-10R are flow diagrams of methods for entering text into a text entry field, according to various embodiments. In some embodiments, method 1000 is performed on a computer system (e.g., computer system 101 of FIG. 1) that includes a display generating component (e.g., display generating component 120 of FIGS. 1, 3, and 4) (e.g., a heads-up display, a display, a touchscreen, and / or a projector) and one or more input devices. In some embodiments, method 1000 is governed by instructions stored on a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1A). Some operations in method 1000 are, optionally, combined, and / or the order of some operations is, optionally, changed.
[0251] 9A , method 1000 is performed in a computer system (e.g., 101) in communication with a display generation component (e.g., 120) and one or more input devices. In some embodiments, the computer system is the same as or similar to the computer system described above with reference to method 800. In some embodiments, the one or more input devices are the same as or similar to the one or more input devices described above with reference to method 800. In some embodiments, the display generation component is the same as or similar to the display generation component described above with reference to method 800.
[0252] In some embodiments, a computer system (e.g., 101), via a display generation component (e.g., 120), displays (1002a) a text entry field (e.g., 906), such as in FIG. 9A . In some embodiments, the text entry field is displayed in a three-dimensional environment that is the same as or similar to the three-dimensional environment described above with reference to method 800. In some embodiments, the text entry field is an interactive user interface element that accepts text input. In some embodiments, the three-dimensional environment includes selectable options that, when selected, cause the computer system to perform an action on text that was (e.g., previously) entered into the text entry field. For example, the text entry field is a web address bar, a search box, a field that accepts a filename, a message field, or a word processor, and the selectable options are, respectively, navigation options, search options, save or load options, options for sending a message, or an option for saving the entered text as a document. In some embodiments, the text entry field has one or more of the characteristics of a text entry field described below with reference to methods 1200, 1400, and / or 1600.
[0253] In some embodiments, while displaying (1002b) a text entry field, such as in FIG. 9A , via a display generation component (e.g., 120), the computer system (e.g., 101) detects (1002c) a first speech input (e.g., 916a) from a user, such as in FIG. 9B , via one or more input devices (e.g., a microphone). In some embodiments, receiving the first speech input includes detecting the user speaking a word, a number, a letter, and / or a special character (e.g., a non-letter symbol included in written text). In some embodiments, while detecting the user's gaze directed at the text entry field and the first speech input, the computer system does not detect any additional input (e.g., via one or more input devices other than an eye tracking device and / or a microphone) corresponding to a request to enter text into the text entry field.
[0254] In some embodiments, in response to detecting (1002d), via the display generation component (e.g., 120), while displaying the text entry field (1002b), a first speech input (e.g., 916a) from the user, as in FIG. 9A, and pursuant to determining that the user's attention (including, e.g., gaze 913a) was directed to the text entry field (e.g., 906) such as in FIG. 9B when the first speech input (e.g., 916a) was received (e.g., the user's gaze or a proxy of the user's gaze was maintained for a threshold period prior to detecting the first speech input, as described in more detail below), the computer system displays (1002e), via the display generation component (e.g., 120), a text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) such as in FIG. 9C. In some embodiments, the text representation of the first speech input is a written representation of the words and / or characters spoken by the user. In some embodiments, before receiving the first speech input, the computer system presents distinct text in the text entry field, and displaying the font-based text representation of the first speech input includes replacing the distinct text with the text representation of the first speech input. For example, the distinct text indicates a purpose of the text entry field (e.g., "Message" or similar text in a messaging text entry field, "Search" or "Enter search terms here" in a search text entry field) or includes text associated with a previous or current function of an application associated with the text entry field (e.g., a website URL presented in a web browser when the first speech input was received). In some embodiments, the font-based text representation of the first speech input in the text entry field is added to the distinct text, such as adding text to a document in a word processing application.
[0255] In some embodiments, via the display generation component (e.g., 120), while displaying the text entry field (1002b), in response to detecting a first speech input (e.g., 916a) from the user (1002d), as in FIG. 9A, or in response to determining that the user's attention (e.g., 913b) is not directed to the text entry field (e.g., 902) when the first speech input (e.g., 916b) from the user is received (e.g., the user's gaze is not directed to the text entry field or the gaze is maintained for less than a threshold period, described in more detail below, and prior to detecting the first speech input), as in FIG. 9C, the computer system (e.g., 101) ceases displaying (e.g., 1002f) the text representation of the first speech input in the text entry field (e.g., 906), as in FIG. 9E. In some embodiments, ceasing to display the text representation of the first speech input in the text entry field includes maintaining display of the discrete text displayed in the text entry field while the first speech input was detected.
[0256] Displaying a textual representation of a first speech input in a text entry field as described above enhances user interaction with a computer system by providing an additional control technique (e.g., speech input) without cluttering the user interface with additional displayed controls.
[0257] 9B , determining that a user's attention (e.g., 913a) is directed to the text entry field (e.g., 906) when a first speech input from a user (e.g., 916a) is received includes determining that the user's gaze (e.g., 913a) is directed to the text entry field (e.g., 906) for at least a time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds) (1004a). In some embodiments, the computer system uses eye tracking devices included in one or more input devices to determine the location of the user's gaze. In some embodiments, determining that the user's attention is not directed to the text entry field includes determining that the user's gaze is not directed to the text entry field or that the user's gaze is directed to the text entry field for less than a time threshold. Displaying a text representation of a first speech input in a text entry field based on detecting a user's gaze directed at the text entry field for a time threshold improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0258] In some embodiments, such as FIG. 9A , detecting that a user's attention (e.g., 913a) is directed to a text entry field (e.g., 906) includes detecting that the user's gaze (e.g., 913a) is directed to the text entry field (e.g., 906) for longer than a time threshold (e.g., 0.1, 0.2, 0, 3, 0.5, 1, 2, or 3 seconds) (1006a). In some embodiments, such as FIG. 9B , in response to detecting the user's gaze (e.g., 913a) directed to the text entry field (e.g., 906) while displaying the text entry field (e.g., 906), the computer system (e.g., 101) presents an indication (e.g., 910a and / or 914) of the duration the user's gaze was directed to the text entry field (1006b). In some embodiments, the computer system modifies the indication of the duration the user's gaze was directed to the text entry field as the user's gaze continues to be directed to the text entry field. In some embodiments, the computer system presents an indication in response to detecting a user's gaze directed toward the text entry field for a time threshold. In some embodiments, the indication is a visual indication displayed via a display generation component. In some embodiments, the indication is an audio indication presented via one or more audio output devices in communication with the computer system. In some embodiments, the visual indication is a gradual expansion (e.g., horizontally) of the text entry field. In some embodiments, the visual indication is a progress bar. In some embodiments, the visual indication is a gradual change in the color of the text entry field and / or the outline of the text entry field.
[0259] Presenting an indication of the duration that a user's gaze has been directed at a text entry field improves user interaction with a computer system by providing enhanced feedback to the user.
[0260] 9B , in response to detecting (1008a) a user's gaze (e.g., 913a) directed toward the text entry field (e.g., 906) while displaying the text entry field (e.g., 906), and following a determination of a duration (e.g., meeting or exceeding a time threshold) that the user's gaze (e.g., 913a) was directed toward the text entry field (e.g., 906), the computer system (e.g., 101) presents (1008b) a second indication (e.g., 910a and / or 914) indicating that a first speech input (e.g., 916a) is directed toward the text entry field (e.g., 906). In some embodiments, presenting the second indication indicating that the first speech input is directed toward the text entry field includes expanding the text entry field. For example, the computer system increases the width of the text entry field. In some embodiments, presenting the second indication that the first speech input is directed to the text entry field includes initiating the display of a visual indication (e.g., an icon or image, such as an image of a microphone or a speech bubble). In some embodiments, the second indication that the first speech input is directed to the text entry field is displayed at an insertion location in the text within the text entry field where the text of the first speech input is entered in response to the first speech input. In some embodiments, the second indication that the first speech input is directed to the text entry field is an audio indication presented via one or more audio output devices in communication with the computer system.
[0261] 9A , in response to detecting a user's gaze (e.g., 913a) directed toward the text entry field (e.g., 906) (1008a) while displaying the text entry field (e.g., 906), the computer system (e.g., 101) ceases presenting the second indication (1008c) in accordance with a determination that the user's gaze (e.g., 913a) was directed toward the text entry field (e.g., 906) for a duration less than a time threshold. In some embodiments, in response to detecting that the user's gaze was directed toward the text entry field, the computer system maintains display of the visual indication regardless of whether the user's gaze was directed toward the text entry field for the time threshold. In some embodiments, in response to detecting that the user's gaze was directed toward the text entry field for the time threshold, the computer system ceases displaying the visual indication that the user's gaze is directed toward the text entry field.
[0262] Presenting a second indication that the first speech input is directed to the text entry field in response to the user's gaze being directed to the text entry field for a time threshold improves user interaction with the computer system by providing enhanced feedback to the user.
[0263] In some embodiments, in response to detecting (1010a) a first speech input (e.g., 916a) from a user while displaying the text entry field (e.g., 906), as in FIG. 9B , and pursuant to determining that the user's attention (e.g., 913a) is directed to the text entry field (e.g., 906), the computer system (e.g., 101), via the display generation component (e.g., 120), displays a text cursor (e.g., 912) within the text entry field (e.g., 906), and a text representation (e.g., 920) of the first speech input is inserted (1010b) into the text entry field (e.g., 906) at the location of the text cursor (e.g., 912) within the text entry field (e.g., 906), as in FIG. 9C . In some embodiments, the computer system does not display the text cursor in the text entry field unless and until it detects a first speech input from the user while the user's attention is directed to the text entry field. In some embodiments, the text cursor is an insertion marker. In some embodiments, after displaying the text representation of the first speech input in the text entry field, in accordance with a determination that the user's attention is still directed to the text entry field, the computer system maintains the display of the text cursor at an updated location in the text entry field (e.g., at the end of the text representation of the first speech input). In some embodiments, the text cursor is a visual indication displayed via a display generation component that indicates the location in the text entry field where text will be entered in response to input corresponding to a request to enter text in the text entry field (e.g., dictation input, soft keyboard input according to methods 1200, 1400, and / or 1600, or hardware keyboard input).In some embodiments, the computer system updates the position of the text cursor within the text entry field while entering the individual pieces of text into the text entry field in response to an input corresponding to the request to enter text to indicate that subsequent text entered in response to the subsequent input corresponding to the request to enter text into the text entry field will be entered after the individual pieces of text.
[0264] In some embodiments, in response to detecting (1010a) a first speech input (e.g., 916a) from a user while displaying a text entry field (e.g., 906), such as in FIG. 9B, and determining that the user's attention (e.g., 913b) is not directed to the text entry field (e.g., 906), such as in FIG. 9C, the computer system (e.g., 101) ceases to display (1010c) a text cursor within the text entry field (e.g., 906), such as in FIG. 9A.
[0265] Displaying a text cursor in a text entry field enhances user interaction with a computer system by providing the user with enhanced visual feedback.
[0266] In some embodiments, as in FIG. 9A, detecting that a user's attention (e.g., 913a) is directed to a text entry field (e.g., 906) includes detecting that the user's gaze (e.g., 913a) is directed to the text entry field (e.g., 906) for longer than a time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds) (1012a).
[0267] In some embodiments, as in FIG. 9A , while the user's attention (e.g., 913a) is directed away from the text entry field (e.g., 906), the computer system (e.g., 101), via the display generation component (e.g., 120), displays (1012b) the text entry field (e.g., 906) with visual characteristics (e.g., color, opacity, line style, and / or size) having a first value.
[0268] In some embodiments, while displaying a text entry field (e.g., 906) having a visual characteristic with a first value via a display generation component (e.g., 120), as in FIG. 9A, a computer system (e.g., 101) detects (1012c) a user's gaze (e.g., 913a) directed toward the text entry field (e.g., 913a) as in FIG. 9A via one or more input devices (e.g., 314).
[0269] 9B , in response to detecting a user's gaze (e.g., 913a) directed at the text entry field (e.g., 906), the computer system (e.g., 101) gradually modifies (1012d), via the display generation component (e.g., 120), the display of the text entry field (e.g., 906) having the visual characteristic with a first value, so as to display, via the display generation component (e.g., 120), the text entry field (e.g., 906) having the visual characteristic with a second value different from the first value, according to the duration that the user's gaze (e.g., 913a) is directed at the text entry field (e.g., 906). In some embodiments, the value of the visual characteristic changes over time as the user's gaze remains directed at the text entry field. For example, the visual characteristic may be color, size, border, or brightness, and the computer system may display the text entry field in a first color, size, border, or brightness while the user's gaze is not directed at the text entry field, gradually change the color, size, border, or brightness of the text entry field while the user's gaze remains directed at the text entry field, and transition to displaying the text entry field in a second color, size, border, or brightness in response to detecting the user's gaze directed at the text entry field over a time threshold.
[0270] Gradually modifying the values of visual characteristics of a text entry field in response to detecting a user's gaze directed at the text entry field improves user interaction with a computer system by providing the user with enhanced visual feedback.
[0271] In some embodiments, while displaying the text entry field (e.g., 906) via the display generation component (e.g., 120), in response to detecting a first speech input (e.g., 916b) from the user in accordance with a determination that the user's attention (e.g., 913a) is directed to the text entry field (e.g., 906) when a first speech input (e.g., 916a) from the user is received, the computer system (e.g., 101) displays (1014) via the display generation component (e.g., 120) the text entry field (e.g., 906) with visual characteristics (e.g., size, color, opacity, outline style, and / or visual effects such as glow or shadow) having discrete values that change over time in accordance with changes in the characteristics (e.g., volume, tone, and / or frequency) of the first speech input (e.g., 916a) over time, as shown in Figure 9B-9C. In some embodiments, the visual characteristic is a glow effect displayed around the text entry field. In some embodiments, the intensity (e.g., color darkness, brightness, saturation, thickness, and / or opacity) of the glow (and / or other visual characteristics) varies over time according to the audio level of the first speech input.
[0272] Displaying a text entry field with visual characteristics having individual values that change over time according to characteristics of the first speech input improves user interaction with the computer system by providing the user with enhanced visual feedback.
[0273] In some embodiments, after detecting a first speech input (e.g., 916a) from a user via one or more input devices while displaying a text entry field (e.g., 906) via a display generation component (e.g., 120) (1016a), as in FIG. 9B, the computer system (e.g., 101) detects a second speech input (e.g., 916b), a continuation of the first speech input from the user, via one or more input devices (e.g., 314) (1016b) while the user's attention (e.g., 913b) is not directed to the text entry field, as in FIG. 9C. In some embodiments, the start of the second speech input from the user is detected within a time threshold (e.g., 0.5, 1, 2, 3, 4, or 5 seconds) for detecting the end of the first speech input. For example, the computer system detects that the user has not spoken for less than the time threshold between the first speech input and the second speech input. In some embodiments, the user's attention is directed to an area of the three-dimensional environment other than the text entry field. In some embodiments, the user's attention is not directed to the three-dimensional environment. In some embodiments, the user's eyes are closed for longer than a time threshold associated with blinking (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds).
[0274] In some embodiments, after detecting a first speech input (e.g., 916a) from a user via one or more input devices while displaying the text entry field (e.g., 906) (1016a), via the display generation component (e.g., 120), as in FIG. 9B; in response to detecting a second speech input (e.g., 916b) from the user while the user's attention (e.g., 913b) was not directed to the text entry field (e.g., 906) (1016c), as in FIG. 9C; or in response to determining that the user's attention (e.g., 913a) was directed to the text entry field (e.g., 906) when the first speech input (e.g., 916a) from the user was received, the computer system (e.g., 101) displays a text representation (e.g., 920) of the second speech input in the text entry field (e.g., 906) via the display generation component (e.g., 120) (1016d), as in FIG. 9D. In some embodiments, the computer system displays the text representation of the first speech input in the text entry field while the user is providing the second speech input. In some embodiments, the computer system displays the text representation of the second speech input in the text entry field simultaneously with the text representation of the first speech input. In some embodiments, the computer system initiates a process of presenting the text representation of the speech input in the text entry field in response to detecting the user's attention directed to the text entry field, and continues entering the text representation of the additional speech input even if additional speech input is detected while the user's attention is no longer directed to the text entry field.
[0275] In some embodiments, after detecting a first speech input (e.g., 916a) from a user via one or more input devices, as in FIG. 9B, while displaying the text entry field (e.g., 906) via the display generation component (e.g., 120) (1016a), in response to detecting a second speech input (e.g., 916b) from the user while the user's attention (e.g., 913b) is not directed to the text entry field (e.g., 906) (1016c), as in FIG. 9C, or in response to determining that the user's attention was not directed to the text entry field (e.g., 906) when the first speech input was received, the computer system (e.g., 101) ceases displaying the text representation of the second speech input in the text entry field (e.g., 906) via the display generation component (e.g., 120) (1016e), as in FIG. 9E. In some embodiments, because the user's attention was not directed to the text entry field when the first speech input was received, the computer system ceases displaying the text representation of the first speech input in the text entry field and displays the text entry field without the text representation of the first speech input while the second speech input is received (e.g., regardless of where the user is looking while the computer system detects the second speech input). In some embodiments, the computer system does not begin the process of entering the text representation of the speech input into the text entry field unless and until the computer system detects the user's attention being directed to the text entry field.
[0276] Displaying a text representation of the second speech input in a text entry field enhances user interaction with the computer system by performing an action when a set of conditions is met without requiring further user input.
[0277] In some embodiments, as in Figure 9C, while displaying (1018a) a text representation (e.g., 920) of the first speech input within the text entry field (e.g., 906) in response to detecting a first speech input from the user while the user's attention was directed away from the text entry field, the computer system (e.g., 101) receives (1018b) a second speech input (e.g., 916b) via one or more input devices (e.g., 120) that is a continuation of the first speech input from the user while the user's attention (e.g., 906) was directed away from the text entry field. In some embodiments, the second speech input received while the user's attention was directed away from the text entry field is similar to the second speech input received while the user's attention was directed away from the text entry field, as described above.
[0278] In some embodiments, as in FIG. 9C , while displaying (1018a) a text representation (e.g., 920) of the first speech input within a text entry field (e.g., 906) in response to detecting a first speech input from a user while the user's attention is directed to the text entry field:
[0279] 9D , in response to receiving the second speech input (e.g., 916b in FIG. 9C ), the computer system (e.g., 101) displays (1018c) a second text representation of the first speech input (e.g., 920) via the display generation component (e.g., 120). In some embodiments, the computer system continues to input a text representation of the user utterance after inputting the text representation of the first speech input in response to detecting the user's attention directed to the text entry field while providing the first speech input, as described above.
[0280] Displaying a textual representation of the continuation of a first speech input in a text entry field enhances user interaction with a computer system by performing an action when a set of conditions is met without requiring further user input.
[0281] 9C , while displaying a text representation (e.g., 920) of a first speech input in a text entry field (e.g., 906) via a display generation component (e.g., after detecting the first speech input while the user's attention is directed to the text entry field), the computer system (e.g., 101) detects (1020a) via one or more input devices (e.g., 314) a second speech input (e.g., 916b) that is a continuation of the first speech input from the user while the user's attention (e.g., 913b) was not directed to the text entry field (e.g., 906). In some embodiments, the second speech input from the user detected while the user's attention was not directed to the text entry field is similar to the second speech input from the user detected while the user's attention was not directed to the text entry field described above.
[0282] 9E , in response to detecting a second speech input (e.g., 916b) from the user while the user's attention (e.g., 913b) is not directed to the text entry field (e.g., 906), the computer system (e.g., 101), via the display generation component (e.g., 120), ceases displaying (1020b) the text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906). In some embodiments, in response to detecting that the user's attention has been directed away from the text entry field (e.g., detecting that the user has looked away from the text entry field), the computer system deletes the text in the text entry field that was previously entered via dictation. In some embodiments, in response to detecting that the user's attention has been directed away from the text entry field (e.g., detecting that the user has looked away from the text entry field), the computer system deletes the text in the text entry field that was entered via dictation without performing an action associated with the text entry field (e.g., searching for the search term entered in the text entry field, sending the message entered in the text entry field, and / or navigating to the website entered in the text entry field).
[0283] Ceasing to display the text representation of the first speech input in the text entry field in response to detecting a second speech input while the user's attention is directed away from the text entry field improves user interaction with the computer system by reducing the number of inputs required to perform an action (e.g., removing the text representation of the first speech input from the text entry field).
[0284] 9B , in response to detecting a first speech input (e.g., 916a) from a user, in accordance with a determination that the user's attention (e.g., 913a) was directed to the text entry field (e.g., 906) when the first speech input (e.g., 916a) was received, the computer system (e.g., 101), via the display generation component (e.g., 120), displays (1022a) the text entry field (e.g., 906) with visual characteristics (e.g., color, size, opacity, text style such as font style, text size, and / or text highlighting, and / or border style) having a first value. In some embodiments, the computer system displays the text entry field with visual characteristics having the first value while detecting speech input while the user's attention is directed to the text entry field. In some embodiments, the computer system displays a text representation of the speech input with highlighting in response to (e.g., and during) detecting speech input while the user's attention is directed to the text entry field. In some embodiments, the computer system performs the display.
[0285] 9C , while displaying, via a display generating component (e.g., 120), a text entry field (e.g., 906) having a visual characteristic with a first value, a computer system (e.g., 101) detects (1022b) via one or more input devices (e.g., 314) that a user's attention (e.g., 913b) is not directed to the text entry field (e.g., 906). In some embodiments, the user's attention is directed to an area of the three-dimensional environment other than the text entry field. In some embodiments, the user's attention is directed away from the three-dimensional environment (e.g., away from the display generating component). In some embodiments, the user closes their eyes for more than a time threshold associated with blinking (e.g., 0.5, 1, 2, 3, or 5 seconds).
[0286] In some embodiments, as in FIG. 9A , in response to detecting that the user's attention is not directed to the text entry field (e.g., 906), the computer system (e.g., 101), via the display generation component (e.g., 120), displays (1022c) the text entry field (e.g., 906) with a visual characteristic having a distinct value that changes over time until it reaches a second value different from the first value. In some embodiments, the value of the visual characteristic gradually changes over time until it reaches the second value in response to detecting that the user's attention is not directed to the text entry field. For example, a highlight on the text included in the text entry field gradually fades. Transitioning to displaying the text entry field with the visual characteristic having the second value in response to detecting that the user's attention is not directed to the text entry field improves user interaction with the computer system by providing the user with improved visual feedback.
[0287] In some embodiments, as in FIG. 9D , in response to detecting a first speech input from a user (1024a) while displaying a text representation (e.g., 920) of the first speech input in a text entry field (e.g., 906), the computer system (e.g., 101) detects a second speech input (e.g., 916c) (1024b) via one or more input devices (e.g., 314).
[0288] In some embodiments, in response to detecting a first speech input from a user (1024a) while displaying the text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906), as in Figure 9D, and in response to detecting a second speech input (e.g., 916c), as in Figure 9D, and in response to determining that the second speech input (e.g., 916c) corresponds to a request to perform an action on the text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) and one or more criteria are met (e.g., including criteria that are met when the user's attention is directed to the text entry field), the computer system (e.g., 101) performs an action on the text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) (1024d). In some embodiments, the second speech input is or includes a predetermined utterance associated with the action. For example, the text entry field may be a message composition field, the second speech input may be "send," "send that," etc., and the action may be to send a message that includes the text representation of the first speech input. As another example, the text entry field may be a search field, the second speech input may be "search," "go," etc., and the action may be to perform a search that includes the text representation of the first speech input as a search term.
[0289] 9D , in response to detecting a first speech input from a user (1024a) while displaying the text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906), in response to detecting a second speech input (e.g., 916c) and determining that the second speech input does not correspond to a request to perform an action on the text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) or one or more criteria are not met, the computer system (e.g., 101) refrains from performing an action on the text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) as of (1024e). In some embodiments, the second speech input does not include a predetermined speech associated with the action. In some embodiments, in response to a second speech input that does not correspond to a request to perform an action (e.g., instead of or in addition to the text representation of the first speech input), the computer displays a text representation of the second speech input in the text entry field.
[0290] Performing an action on a text representation of a first speech input in a text entry field in response to a second speech input enhances user interaction with a computer system by providing additional control without cluttering the user interface with additional displayed controls.
[0291] 9D , in accordance with a determination that the text entry field (e.g., 906) is a first type text entry field, a determination that the second speech input (e.g., 916c) corresponds to a request to perform an action on a text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) is based on one or more first criteria (1026a). In some embodiments, the one or more first criteria include a criterion that is met when the second speech input includes the first utterance. For example, if the text entry field is a search field, the one or more first criteria include a criterion that is met when the second speech input includes “search,” “go,” etc. In some embodiments, in accordance with a determination that the second speech input corresponds to an action associated with a second type text entry field that is different from the first type text entry field, the computer system refrains from performing the action on the text representation of the first speech input in the text entry field.
[0292] In some embodiments, in accordance with a determination that the text entry field is a second type of text entry field different from the first type of text entry field (e.g., a text entry field different from text entry field 906 of FIG. 9D ), a determination that the second speech input (e.g., 916c) corresponds to a request to perform an action on the text representation of the first speech input in the text entry field (e.g., 920) is based on one or more second criteria different from the one or more first criteria (1026b). In some embodiments, the one or more second criteria include a criterion that is met when the second speech input includes the second utterance. For example, if the text entry field is a messaging field, the one or more first criteria include a criterion that is met when the second speech input includes “send,” “send it,” etc. In some embodiments, in accordance with a determination that the second speech input corresponds to an action associated with a first type of text entry field different from the second type of text entry field, the computer system refrains from performing the action on the text representation of the first speech input in the text entry field.
[0293] Evaluating the second speech input according to different criteria depending on the type of text entry field enhances user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0294] In some embodiments, the one or more criteria include a criterion that is met when the user's gaze (e.g., 913c) is directed toward the text entry field (e.g., 906) (e.g., for at least a time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 seconds)) while the computer system (e.g., 101) detects the second speech input (e.g., 916c) (1028), as in FIG. 9D . In some embodiments, pursuant to a determination that the user's gaze is not directed toward the text entry field while the computer system detects the second input, the computer system refrains from performing the action on the text representation of the first speech input in the text entry field, regardless of whether the second speech input satisfies one or more additional criteria for determining that the second speech input corresponds to a request to perform an action on the text representation of the first speech input in the text entry field, such as the first speech input including default speech associated with the action.
[0295] Determining that the second speech input corresponds to a request to perform an action on the text representation of the first speech input in the text entry field based on the user's gaze being directed toward the text entry field while detecting the second speech input improves user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0296] In some embodiments, such as in FIG. 9A , prior to detecting a first speech input (e.g., 916a) from a user, as in FIG. 9B , the computer system (e.g., 101) displays (1030a) personalized text in a text entry field (e.g., 906) via a display generation component (e.g., 120). In some embodiments, the personalized text was previously entered in response to a second speech input similar to the first speech input described above and according to the same or similar conditions as those described above. In some embodiments, the personalized text was previously entered by the user via a different input modality, such as using a soft keyboard according to one or more of methods 1200, 1400, or 1600 described below, or using a hardware keyboard. In some embodiments, the personalized text is placeholder text that is automatically displayed by the computer system without receiving input corresponding to a request to enter placeholder text in the text entry field.
[0297] In some embodiments, in response to detecting a first speech input (e.g., 916a) from a user as in FIG. 9B , in accordance with a determination that the user's attention (e.g., 913a) is directed to the text entry field (e.g., 906), the computer system (e.g., 101), via the display generation component (e.g., 120), ceases displaying the individual text in the text entry field (e.g., 906) and displays (1030b) a text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) as in FIG. 9C . In some embodiments, in response to detecting the first speech input from the user while the user's attention is directed to the text entry field, the computer system replaces the individual text with the text representation of the first speech input in the text entry field. In some embodiments, the computer system ceases displaying the individual text in the text entry field in response to the first speech input without detecting any additional input corresponding to a request to cease displaying the individual text in the text entry field. Ceasing the display of the separate text and displaying a text representation of the first speech input in the text entry field in response to detecting the first speech input enhances user interaction with the computer system by performing an action when a set of conditions is met without requiring further user input.
[0298] In some embodiments, prior to detecting a first speech input from a user, as in FIG. 9F , a computer system (e.g., 101), via a display generation component (e.g., 120), displays (1032a) an individual piece of text and a cursor (e.g., 928) at a first location within a text entry field (e.g., 926). In some embodiments, the computer system displays the cursor in accordance with a determination that the individual piece of text can be edited. In some embodiments, the computer system displays the cursor in accordance with a determination that the individual piece of text can be edited in a manner other than replacing the entire individual piece of text (e.g., adding text or deleting a portion of the individual piece of text without detecting the entire individual piece of text). In some embodiments, in accordance with a determination that the individual piece of text cannot be edited, the computer system ceases displaying the cursor prior to detecting the first speech input (optionally in response to detecting a first speech input while the user's attention was directed to the text entry field).
[0299] 9G , in response to detecting a first speech input (e.g., 916d) from a user, in accordance with a determination (1032a) that the user's attention (e.g., 913d) is directed to the text entry field (e.g., 926), the computer system (e.g., 101), via the display generation component (e.g., 120), maintains (1032b) the display of the discrete text in the text entry field (e.g., 926). In some embodiments, in accordance with a determination that the discrete text is not editable, in response to detecting the first speech input while the user's attention is directed to the text entry field, the computer system ceases displaying the discrete text in the text entry field and displays a cursor or visual indication, described in more detail below.
[0300] In some embodiments, as in FIG. 9G, in response to detecting a first speech input (e.g., 916d) from the user and following a determination (1032a) that the user's attention (e.g., 913d) is directed to the text entry field (e.g., 926), the computer system (e.g., 101), via the display generation component (e.g., 120), ceases displaying (1032d) a cursor in the text entry field (e.g., 926).
[0301] 9G , in response to detecting a first speech input (e.g., 916d) from the user and following a determination (1032a) that the user's attention (e.g., 913d) is directed to the text entry field (e.g., 926), the computer system (e.g., 101), via the display generation component (e.g., 120), displays a visual indication (e.g., 930) at a second location (e.g., the same as the first location or different from the first location) within the text entry field (e.g., 926), and a textual representation of the first speech input is added (1032e) to the respective text at the second location within the text entry field (e.g., 926). In some embodiments, the visual indication is different from a cursor. In some embodiments, the visual indication is the same as a cursor. In some embodiments, after entering the text representation of the first speech input into the text entry field, the computer system displays a visual indication adjacent to (e.g., immediately after) the text representation of the first speech input. In some embodiments, the visual indication is an image of a microphone or a speech bubble or a person speaking. In some embodiments, in response to detecting the user's attention being directed away from the text entry field without detecting a continuation of the first speech input, the computer system ceases displaying the visual indication and begins displaying a cursor (e.g., at a location in the text entry field corresponding to the text representation of the first speech input).
[0302] Displaying a visual indication at a location within the text entry field where a text representation of the first speech input is added in response to detecting the first speech input from the user improves user interaction with the computer system by providing the user with enhanced visual feedback.
[0303] In some embodiments, as in Figures 9G-9H, in accordance with a determination that the user's gaze (e.g., 913d) is directed toward a first portion of text in the text entry field (e.g., 926) while a first speech input (e.g., 916d) from the user is detected, a second location at which the text representation of the speech is added to the individual text is proximate (e.g., near or adjacent) to the first portion of text (1034a).
[0304] In some embodiments, in accordance with a determination that the user's gaze is directed to a second portion of text in a text entry field (e.g., 926) that is different from the first location in the text entry field while the first speech input from the user is detected (e.g., the user's gaze 913d in FIG. 9G is at a location other than the location shown in FIG. 9G), the second location at which the text representation of the speech is added to the individual text is proximate to (e.g., near or adjacent to) the second portion of text (1034b). In some embodiments, the computer system displays a visual indication at a location within the text entry field where the user is looking while the user's attention is directed to the text entry field. In some embodiments, the computer system updates the position of the position of the visual indication in accordance with the user's gaze moving from one location within the text entry field to another location within the text entry field before the user provides the first speech input. In some embodiments, once the computer system displays the visual indication at the second location, the computer system maintains the display of the visual indication at the second location, even if the user's gaze moves from the second location, until the computer system ceases displaying the visual indication in accordance with one or more criteria being met (e.g., the user moves their attention away from the text entry field, or the user provides input to a user interface element other than the text entry field, or the user provides input to stop entering text into the text entry field based on the first speech input).
[0305] Displaying visual indications and entering text at a location based on the user's gaze enhances user interaction with a computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0306] In some embodiments, such as in FIG. 9C , while displaying (1036a) a text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) via a display generation component (e.g., 120) in response to detecting the first speech input from the user in accordance with a determination that the user's attention was directed to the text entry field (e.g., 906) when the first speech input from the user was received, the computer system (e.g., 101) detects (1036b) the user's attention (e.g., 913b) that is not directed to the text entry field (e.g., 906) via one or more input devices (e.g., 314) while detecting (1036b) a second speech input (e.g., 916b) that is a continuation of the first speech input from the user via one or more input devices, as in FIG. 9C . In some embodiments, the second speech input that is a continuation of the first speech input detected while the user's attention is not directed to the text entry field is similar to the second speech input that is a continuation of the first speech input detected while the user's attention is not directed to the text entry field, as described in more detail above.
[0307] In some embodiments, such as in FIG. 9C , while in response to detecting the first speech input from the user, displaying (1036a) a text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) via the display generation component (e.g., 120) in accordance with a determination that the user's attention is directed to the text entry field (e.g., 906) when the first speech input from the user is received, in response to detecting (913b) the user's attention not being directed to the text entry field (e.g., 926), as in FIG. 9G , in response to a determination that the text entry field (e.g., 906) is a first type of text entry field, the computer system (e.g., 101) displays (1036d) a continuation of the first speech input (e.g., 920) in the text entry field (e.g., 906) via the display generation component (e.g., 120). In some embodiments, the computer system maintains display of the text representation of the first speech input in the text entry field while simultaneously displaying the text representation of the second speech input. In some embodiments, the computer system ceases displaying the text representation of the first speech input in the text entry field and replaces it with the text representation of the second speech input in the text entry field. In some embodiments, the first type of text entry field is a long-form text entry field, such as a notes field, a word processing application field, or an email composition field, in which the computer system requires input in addition to detecting a user's gaze directed toward the text entry field to begin dictation. In some embodiments, in addition to detecting a user's gaze directed toward the text entry field, the input is a dictation input, a discrete gesture performed with a part of the user's body, and / or a selection of a user interface element associated with a discrete speech input (e.g., "Hey voice assistant, start dictation").
[0308] In some embodiments, such as in FIG. 9C , while displaying (1036a) a text representation (e.g., 920) of the first speech input in the text entry field (e.g., 906) via the display generation component (e.g., 120) in response to detecting the first speech input from the user in accordance with a determination that the user's attention was directed to the text entry field (e.g., 906) when the first speech input from the user was received, the computer system (e.g., 101) ceases displaying (1036e) a text representation of the second speech input in the text entry field via the display generation component (e.g., 120) in response to detecting the user's attention (e.g., 913b) not directed to the text entry field (e.g., 906) in accordance with a determination that the text entry field (e.g., 906) is a second type of text entry field different from the first type of text entry field, as in FIG. 9E . In some embodiments, the computer system ceases displaying the text representation of the first speech input. In some embodiments, the computer system maintains a display of a text representation of the first speech input. In some embodiments, the second type of text entry field is a shorthand text entry field in which the computer system initiates dictation in response to the user's attention being directed to the text entry field without detecting additional input to initiate dictation, such as a message field, a message or notification quick reply field, a search field, a web browser search, browse, or address field.
[0309] Selectively displaying a textual representation of a second speech input in a text entry field in response to detecting a second speech input while the user's attention is directed away from the text entry field enhances user interaction with the computer system by providing additional control options without cluttering the user interface with additional displayed controls.
[0310] In some embodiments, in response to detecting the user's attention (e.g., 913b) not being directed to the text entry field (e.g., 906), as in Figure 9C, the computer system (e.g., 101) maintains display of the text representation of the first speech input in the text entry field (e.g., 926) in accordance with a determination that the text entry field (e.g., 906) is a first type of text entry field, as in Figure 9F, and ceases displaying (1038) the text representation of the first speech input in the text entry field (e.g., 906) in accordance with a determination that the text entry field (e.g., 926) is a second type of text entry field, as in Figure 9E. In some embodiments, the computer system displays a text representation of a second speech input that is a continuation of the first speech input simultaneously with the text representation of the first speech input in the text entry field in accordance with a determination that the text entry field is a first type of text entry field. In some embodiments, in response to determining that the text entry field is a first type text entry field, the computer system displays a text representation of the continuation of the speech input in the text entry field in response to the continuation of the speech input, even if the user's attention is directed away from the text entry field while the continuation of the speech input is detected. In some embodiments, in response to determining that the text entry field is a second type text entry field, the computer system cancels dictation input into the text entry field in response to detecting the user's attention being directed away from the text entry field. For example, the computer system deletes text entered in response to the speech input and ceases entering additional text in response to the continuation of the speech input.
[0311] Selectively maintaining or ceasing display of the text representation of the first speech input depending on the type of text entry field in response to detecting that the user's attention is not directed to the text entry field improves user interaction with the computer system by reducing the number of inputs required to perform an action.
[0312] In some embodiments, as in FIG. 9C , in accordance with a determination that the text entry field (e.g., 906) is a second type of text entry field, the computer system ...
Claims
1. 1. A method comprising:
1. A computer system in communication with a display generation component and one or more input devices, comprising: displaying a text entry field via said display generation component; While displaying the text entry field via the display generation component, detecting a first speech input from a user of the computer system via the one or more input devices; In response to detecting the first speech input from the user, displaying, via the display generation component, a text representation of the first speech input in the text entry field in accordance with a determination that the user's attention is directed to the text entry field when the first speech input from the user is received; and ceasing to display the text representation of the first speech input in the text entry field in accordance with a determination that the user's attention was not directed to the text entry field when the first speech input from the user was received, wherein detecting that the user's attention is directed to the text entry field includes detecting that the user's gaze is directed to the text entry field for longer than a time threshold, the method comprising: The method further includes, in response to detecting, while displaying the text entry field, the user's gaze directed toward the text entry field for longer than the time threshold before receiving the first speech input, presenting an indication that the user's gaze was directed toward the text entry field for longer than the time threshold.
2. 2. The method of claim 1 , wherein the determination that the user's attention is directed to the text entry field when the first speech input from the user is received comprises determining that the user's gaze is directed to the text entry field for at least a time threshold.
3. The method of claim 1, wherein the indication indicates that a first speech input is directed to the text entry field.
4. in response to detecting the first speech input from the user while displaying the text entry field; displaying, via the display generation component, a text cursor within the text entry field in accordance with the determination that the user's attention is directed to the text entry field, wherein the text representation of the first speech input is inserted into the text entry field at a location of the text cursor within the text entry field; 2. The method of claim 1, further comprising: ceasing to display the text cursor within the text entry field in accordance with the determination that the user's attention is not directed to the text entry field.
5. Detecting that the user's attention is directed to the text entry field includes detecting that the user's gaze is directed to the text entry field for longer than a time threshold, the method comprising: displaying, via the display generation component, the text entry field with a visual characteristic having a first value while the user's attention is directed away from the text entry field; detecting the user's gaze directed toward the text entry field via the one or more input devices while displaying the text entry field having the visual characteristic having the first value via the display generation component; In response to detecting the user's gaze directed toward the text entry field, 2. The method of claim 1, further comprising: gradually modifying, via the display generating component, a display of the text entry field having the visual characteristic having the first value to a display of the text entry field having the visual characteristic having a second value different from the first value, according to a duration that the user's gaze is directed at the text entry field.
6. 2. The method of claim 1, further comprising: in response to detecting the first speech input from the user while displaying the text entry field via the display generation component, displaying the text entry field via the display generation component with visual characteristics having discrete values that change over time in accordance with changes in characteristics of the first speech input over time in accordance with the determination that the user's attention was directed to the text entry field when the first speech input from the user was received.
7. while displaying the text entry field via the display generation component after detecting the first speech input from the user via the one or more input devices; detecting, via the one or more input devices, a second speech input from the user that is a continuation of the first speech input while the user's attention is not directed to the text entry field; in response to detecting the second speech input from the user while the attention of the user is not directed to the text entry field; displaying, via the display generation component, a text representation of the second speech input in the text entry field in accordance with the determination that the user's attention was directed to the text entry field when the first speech input from the user was received; 2. The method of claim 1 , further comprising: ceasing to display, via the display generation component, the text representation of the second speech input in the text entry field in accordance with the determination that the user's attention was not directed to the text entry field when the first speech input was received.
8. while displaying the text representation of the first speech input in the text entry field in response to detecting the first speech input from the user while the attention of the user is directed to the text entry field; receiving, via the one or more input devices, a second speech input from the user that is a continuation of the first speech input while the user's attention is directed away from the text entry field; The method of claim 1 , further comprising: in response to receiving the second speech input, displaying, via the display generation component, a textual representation of the second speech input.
9. while displaying the text representation of the first speech input in the text entry field via the display generation component, detecting, via the one or more input devices, a second speech input from the user that is a continuation of the first speech input while the user's attention is not directed to the text entry field; 10. The method of claim 1, further comprising: ceasing, via the display generation component, displaying the text representation of the first speech input in the text entry field in response to detecting the second speech input from the user while the user's attention is not directed to the text entry field.
10. in response to detecting the first speech input from the user, displaying, via the display generation component, the text entry field having a visual characteristic having a first value in accordance with the determination that the user's attention was directed to the text entry field when the first speech input was received; detecting, via the one or more input devices, that the user's attention is not directed toward the text entry field while displaying, via the display generation component, the text entry field having the visual characteristic with the first value; 2. The method of claim 1, further comprising: in response to detecting that the user's attention is not directed at the text entry field, displaying, via the display generation component, the text entry field having the visual characteristic with a distinct value that changes over time until it reaches a second value different from the first value.
11. while displaying the text representation of the first speech input in the text entry field in response to detecting the first speech input from the user; detecting a second speech input via the one or more input devices; In response to detecting the second speech input, performing the action on the text representation of the first speech input in the text entry field in accordance with a determination that the second speech input corresponds to a request to perform the action on the text representation of the first speech input in the text entry field and one or more criteria are satisfied; and 2. The method of claim 1 , further comprising: withdrawing from performing the action on the text representation of the first speech input in the text entry field in accordance with a determination that the second speech input does not correspond to the request to perform the action on the text representation of the first speech input in the text entry field or the one or more criteria are not met.
12. pursuant to a determination that the text entry field is a first type of text entry field, the determination that the second speech input corresponds to the request to perform the action on the text representation of the first speech input in the text entry field is based on one or more first criteria; 12. The method of claim 11 , wherein, pursuant to a determination that the text entry field is a second type of text entry field that is different from the first type of text entry field, the determination that the second speech input corresponds to the request to perform the action on the text representation of the first speech input in the text entry field is based on one or more second criteria that are different from the one or more first criteria.
13. 12. The method of claim 11, wherein the one or more criteria include a criterion that is satisfied when the user's gaze is directed toward the text entry field while the computer system detects the second speech input.
14. displaying, via the display generation component, a distinct piece of text in the text entry field prior to detecting the first speech input from the user; 2. The method of claim 1 , further comprising: in response to detecting the first speech input from the user, ceasing to display the individual text in the text entry field in accordance with the determination that the user's attention is directed to the text entry field, via the display generation component, and displaying the text representation of the first speech input in the text entry field.
15. displaying, via the display generation component, a separate text and a cursor at a first location within the text entry field prior to detecting the first speech input from the user; in response to detecting the first speech input from the user and in accordance with the determination that the user's attention is directed to the text entry field; maintaining, via the display generation component, a display of the individual text within the text entry field; ceasing to display the cursor within the text entry field via the display generation component; and 10. The method of claim 1, further comprising: displaying, via the display generation component, a visual indication at a second location within the text entry field, wherein the text representation of the first speech input is appended to the individual text at the second location within the text entry field.
16. pursuant to a determination that the user's gaze is directed toward a first portion of the text in the text entry field while the first speech input from the user is detected, the second location at which the text representation of the first speech input is added to the individual text is proximate to the first portion of the text; 16. The method of claim 15, wherein, in accordance with a determination that the user's gaze is directed toward a second portion of the text in the text entry field that is different from the first location in the text entry field while the first speech input from the user is detected, the second location at which the text representation of the first speech input is added to the individual text is proximate to the second portion of the text.
17. while, in response to detecting the first speech input from the user according to the determination that the attention of the user was directed to the text entry field when the first speech input from the user was received, displaying the text representation of the first speech input in a text entry field via the display generation component; detecting, via the one or more input devices, that the user's attention is not directed to the text entry field while detecting, via the one or more input devices, a second speech input from the user that is a continuation of the first speech input; in response to detecting that the user's attention is not directed to the text entry field; displaying, via the display generation component, a text representation of the continuation of the first speech input in the text entry field in accordance with determining that the text entry field is a first type of text entry field; 10. The method of claim 1, further comprising: withdrawing, via the display generation component, display of the text representation of the second speech input in the text entry field in accordance with a determination that the text entry field is a second type of text entry field that is different from the first type of text entry field.
18. in response to detecting that the user's attention is not directed to the text entry field; 18. The method of claim 17, further comprising: maintaining display of the text representation of the first speech input in the text entry field in accordance with the determination that the text entry field is the first type of text entry field; and ceasing display of the text representation of the first speech input in the text entry field in accordance with the determination that the text entry field is the second type of text entry field.
19. 18. The method of claim 17, wherein, in accordance with the determination that the text entry field is the second type of text entry field, the computer system, in response to detecting the first speech input from the user, displays the text representation of the first speech input in the text entry field via the display generation component, in accordance with the determination that the user's attention is directed to the text entry field when the first speech input from the user is received, regardless of whether the computer system detects a separate text entry input, distinct from the first speech input, via the one or more input devices prior to detecting the first speech input.
20. In accordance with the determination that the text entry field is the first type of text entry field, displaying the text representation of the first speech input in the text entry field via the display generation component is responsive to detecting, via the one or more input devices, a separate text entry input distinct from the first speech input prior to detecting the first speech input, the method comprising:
18. The method of claim 17, further comprising: in response to detecting the first speech input from the user, in accordance with the determination that the text entry field is the first type of text entry field, ceasing to display, via the display generation component, the text representation of the first speech input in the text entry field in accordance with a determination that the distinct text entry input is not detected prior to detecting the first speech input from the user.
21. Displaying the text representation of the first speech input includes displaying the text representation of the first speech input in a first appearance, the method comprising: receiving typed text entry input directed to the text entry field via the one or more input devices; 10. The method of claim 1, further comprising: in response to receiving the typed text entry input, displaying, via the display generation component, a text representation of the typed text entry input in the text entry field, wherein the text representation of the typed text entry input is displayed in a second appearance different from the first appearance.
22. 22. The method of claim 21 , wherein displaying the text representation of the first speech input in the first appearance comprises displaying the text representation of the first speech input with a glowing effect, and wherein displaying the text representation of the typed text entry input in the text entry field in the second appearance comprises displaying the text representation of the typed text entry input in the text entry field without the glowing effect.
23. Displaying the text representation of the first speech input in the first appearance includes: displaying, via the display generation component, individual portions of the text representation of the first speech input in one or more colors that vary over time for a period of time after displaying the individual portions of the text representation of the first speech input in the text entry field; and after the period of time has elapsed, displaying, via the display generation component, the distinct portions of the textual representation of the first speech input in distinct colors that do not change over time.
24. 24. The method of claim 23, wherein displaying the individual portions of the text representation of the first speech input in the colors that change over time comprises displaying the individual portions of the text representation of the first speech input in colors that change over time in response to changes in an audio level of the first speech input over time.
25. and displaying, in response to receiving text entry input via the display generation component, a text insertion marker within the text entry field indicating a location within the text entry field where additional text will be added; While detecting the first speech input, the text insertion marker is displayed with a distinct visual effect; The method of claim 1 , wherein the text insertion marker is displayed without the separate visual effect while not detecting the first speech input.
26. 26. The method of claim 25, wherein the distinct visual effect comprises a visual characteristic that changes over time in response to changes in audio level of the first speech input over time.
27. displaying the text entry field, displaying the text entry field with a distinct visual effect while detecting the first speech input; and displaying the text entry field without the distinct visual effect while not detecting the first speech input.
28. 28. The method of claim 27, wherein the distinct visual effect is a glowing visual effect.
29. 30. The method of claim 28, wherein the glowing visual effect comprises a visual characteristic having a value that changes over time in response to changes in audio level of the first speech input over time.
30. 28. The method of claim 27, wherein displaying the text entry field with the distinct visual effect comprises displaying the text entry field in a first color, and wherein displaying the text entry field without the distinct visual effect comprises displaying the text entry field in a second color different from the first color.
31. 31. The method of claim 30, wherein displaying the text entry field in the first color comprises changing the color of the text entry field over time in response to changes in an audio level of the first speech input over time.
32. 32. A program configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the program comprising instructions for performing the method of any one of claims 1 to 31.
33. 1. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 1 to 31.
Citation Information
Patent Citations
How to improve the distinction between dictation and commands
JP2004510239A
Modification of Visual Content to Facilitate Improved Speech Recognition
JP2017525002A
Apparatus and method for filling electronic document based on eye tracking and speech recognition
KR1020210123530A
Information processing device and information processing method
WO2018116556A1
Information processing device and information processing method
WO2020105349A1