User interface and techniques for interactions

By detecting the interaction type and adjusting the position of moving components, the interactive control and display of the computer system are optimized, solving the problems of complexity and inefficiency in the prior art, and improving the efficiency of the user interface and the energy utilization of battery-powered devices.

CN121925622APending Publication Date: 2026-04-24APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
APPLE INC
Filing Date
2024-09-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies for interactive control using computer systems are complex and inefficient, especially in battery-powered devices, wasting user time and energy, and existing interfaces may interfere with multiple media objects.

Method used

By detecting the type of interaction in a computer system and adjusting the position of moving components accordingly, the user interface display is optimized, reducing cognitive load and improving efficiency, making it suitable for battery-powered devices.

Benefits of technology

It enables faster and more efficient interactive control and display, reduces power consumption of battery-powered devices, extends battery life, and optimizes the user interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925622A_ABST
    Figure CN121925622A_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to user interfaces.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 541,843, filed September 30, 2023; U.S. Provisional Patent Application Serial No. 63 / 541,827, filed September 30, 2023; and U.S. Provisional Patent Application Serial No. 63 / 541,829, filed September 30, 2023, the entire contents of which are incorporated herein by reference for all purposes. Background Technology

[0002] Computer systems are frequently used during interactions. Such interactions include lectures, conversations, and meetings. Users often use computer systems to control the user interface. These controls in the user interface include interactive content. Computer systems often display multiple media objects simultaneously. Each displayed media object occupies a portion of the user interface and may therefore interfere with other displayed media objects. Summary of the Invention

[0003] Existing technologies for controlling computer systems based on interactions with electronic devices are often cumbersome and inefficient. For example, some existing technologies use complex and time-consuming user interfaces that may include multiple buttons or keystrokes. Some existing technologies require more time than necessary, wasting user time and device energy. This latter consideration is particularly important in battery-powered devices.

[0004] Therefore, this technology provides electronic devices with faster and more efficient methods and interfaces for interactively controlling computer systems and displaying overlays. Such methods and interfaces can optionally complement or replace other methods for interactively controlling computer systems and displaying overlays. These methods and interfaces reduce the cognitive burden on users and result in more efficient human-computer interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging. These methods and interfaces can complement or replace other methods for interactively controlling computer systems.

[0005] In some embodiments, a method is described that is performed at a computer system communicating with a mobile component. In some embodiments, the method includes: detecting the occurrence of a first interaction when the computer system is in a first location in an environment via the mobile component; and in response to detecting the first interaction: moving via the mobile component to a second location in an environment different from the first location in the environment based on determining that the first interaction is a first type of interaction; and abandoning the movement to the second location via the mobile component based on determining that the first interaction is a second type of interaction different from the first type of interaction.

[0006] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with a mobile component. In some embodiments, the one or more programs include instructions for: detecting the occurrence of a first interaction when the computer system is in a first location in an environment via the mobile component; and in response to detecting the first interaction: moving via the mobile component to a second location in an environment different from the first location in the environment based on determining that the first interaction is a first type of interaction; and abandoning the movement to the second location via the mobile component based on determining that the first interaction is a second type of interaction different from the first type of interaction.

[0007] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a mobile component. In some embodiments, the one or more programs include instructions for: detecting the occurrence of a first interaction when the computer system is in a first location in an environment via the mobile component; and in response to detecting the first interaction: moving via the mobile component to a second location in an environment different from the first location based on determining that the first interaction is a first type of interaction; and abandoning the movement to the second location based on determining that the first interaction is a second type of interaction different from the first type of interaction.

[0008] In some embodiments, a computer system communicating with a mobile component is described. In some embodiments, the computer system communicating with the mobile component includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting the occurrence of a first interaction when the computer system is in a first location in an environment via the mobile component; and in response to detecting the first interaction: moving via the mobile component to a second location in an environment different from the first location based on determining that the first interaction is a first type of interaction; and abandoning the movement to the second location based on determining that the first interaction is a second type of interaction different from the first type of interaction.

[0009] In some embodiments, a computer system communicating with a mobile component is described. In some embodiments, the computer system communicating with the mobile component includes components for performing each of the following steps: detecting the occurrence of a first interaction when the computer system is in a first location in an environment via the mobile component; and in response to detecting the first interaction: moving via the mobile component to a second location in an environment different from the first location in the environment based on determining that the first interaction is a first type of interaction; and abandoning the movement to the second location via the mobile component based on determining that the first interaction is a second type of interaction different from the first type of interaction.

[0010] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a mobile component. In some embodiments, the one or more programs include instructions for: detecting the occurrence of a first interaction when the computer system is in a first location in an environment via the mobile component; and in response to detecting the first interaction: moving via the mobile component to a second location in an environment different from the first location based on determining that the first interaction is a first type of interaction; and abandoning the movement to the second location via the mobile component based on determining that the first interaction is a second type of interaction different from the first type of interaction.

[0011] In some embodiments, a method is described that is executed at a computer system communicating with a display component and a microphone. In some embodiments, the method includes: detecting a first voice input via a microphone while a user interface is displayed via the display component; in response to detecting the first voice input, displaying, in a first manner, a first group of one or more words corresponding to the first voice input via the display component; detecting a second voice input via the microphone while the first group of one or more words corresponding to the first voice input is displayed; and in response to detecting the second voice input: based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, displaying, in a first manner, the new word corresponding to the second voice input along with the display of the first group of one or more words via the display component; and based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, displaying, in a first manner, a second group of one or more words including the new word corresponding to the second voice input via the display component, while stopping the first manner of displaying the first group of one or more words, wherein the second group of one or more words differs from the first group of one or more words.

[0012] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and a microphone. In some embodiments, the one or more programs include instructions for: detecting a first voice input via a microphone while a user interface is displayed via the display component; in response to detecting the first voice input, displaying, in a first manner, a first group of one or more words corresponding to the first voice input via the display component; detecting a second voice input via the microphone while the first group of one or more words corresponding to the first voice input is displayed; and in response to detecting the second voice input: based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, displaying, in a first manner, the new word corresponding to the second voice input along with the display of the first group of one or more words via the display component; and based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, displaying, in a first manner, a second group of one or more words including the new word corresponding to the second voice input via the display component, while stopping the first manner of displaying the first group of one or more words, wherein the second group of one or more words is different from the first group of one or more words.

[0013] In some embodiments, a transient computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and a microphone. In some embodiments, the one or more programs include instructions for: detecting a first voice input via a microphone while a user interface is displayed via the display component; in response to detecting the first voice input, displaying, in a first manner, a first group of one or more words corresponding to the first voice input via the display component; detecting a second voice input via the microphone while the first group of one or more words corresponding to the first voice input is displayed; and in response to detecting the second voice input: based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, displaying, in a first manner, the new word corresponding to the second voice input along with the display of the first group of one or more words via the display component; and based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, displaying, in a first manner, a second group of one or more words including the new word corresponding to the second voice input via the display component, while stopping the first manner of displaying the first group of one or more words, wherein the second group of one or more words is different from the first group of one or more words.

[0014] In some embodiments, a computer system communicating with a display component and a microphone is described. In some embodiments, the computer system communicating with the display component and microphone includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a first voice input via a microphone while a user interface is displayed via the display component; in response to detecting the first voice input, displaying, in a first manner, a first group of one or more words corresponding to the first voice input via the display component; detecting a second voice input via the microphone while the first group of one or more words corresponding to the first voice input is displayed; and in response to detecting the second voice input: based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, displaying, in a first manner, the new word corresponding to the second voice input along with the display of the first group of one or more words via the display component; and based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, displaying, in a first manner, a second group of one or more words including the new word corresponding to the second voice input via the display component, while stopping the first manner of displaying the first group of one or more words, wherein the second group of one or more words differs from the first group of one or more words.

[0015] In some embodiments, a computer system communicating with a display component and a microphone is described. In some embodiments, the computer system communicating with the display component and microphone includes components for performing each of the following steps: detecting a first voice input via the microphone while a user interface is displayed via the display component; in response to detecting the first voice input, displaying, in a first manner, a first group of one or more words corresponding to the first voice input via the display component; detecting a second voice input via the microphone while the first group of one or more words corresponding to the first voice input is displayed; and in response to detecting the second voice input: based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, displaying, in a first manner, the new word corresponding to the second voice input along with the display of the first group of one or more words via the display component; and based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, displaying, in a first manner, a second group of one or more words including the new word corresponding to the second voice input via the display component, while stopping the first manner of displaying the first group of one or more words, wherein the second group of one or more words differs from the first group of one or more words.

[0016] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with a display component and a microphone. In some embodiments, the one or more programs include instructions for: detecting a first voice input via a microphone while a user interface is displayed via the display component; in response to detecting the first voice input, displaying, in a first manner, a first group of one or more words corresponding to the first voice input via the display component; detecting a second voice input via the microphone while the first group of one or more words corresponding to the first voice input is displayed; and in response to detecting the second voice input: based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, displaying, in a first manner, the new word corresponding to the second voice input along with the display of the first group of one or more words via the display component; and based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, displaying, in a first manner, a second group of one or more words including the new word corresponding to the second voice input via the display component, while stopping the first manner of displaying the first group of one or more words, wherein the second group of one or more words differs from the first group of one or more words.

[0017] In some embodiments, a method is described that is performed at a computer system communicating with a display component and one or more input devices. In some embodiments, the method includes: detecting input corresponding to a user via one or more input devices; and, in conjunction with the detected input corresponding to the user, displaying via the display component a representation of a first portion of content related to the input and a representation of a second portion of content related to the input, including: visually grouping the representations of the first portion of content and the second portion of content based on determining that the first portion of content belongs to a first category of content and the second portion of content belongs to a second category of content different from the first category of content; and abandoning the visual grouping of the representations of the first portion of content and the second portion of content based on determining that the first portion of content belongs to the first category of content and the second portion of content belongs to a second category of content different from the first category of content.

[0018] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a user via one or more input devices; and, in conjunction with detecting the input corresponding to the user, displaying via the display component a representation of a first portion of content related to the input and a representation of a second portion of content related to the input, including: visually grouping the representations of the first portion of content and the second portion of content based on determining that the first portion of content belongs to a first category of content and the second portion of content belongs to a second category of content different from the first category of content; and abandoning the visual grouping of the representations of the first portion of content and the second portion of content based on determining that the first portion of content belongs to the first category of content and the second portion of content belongs to a second category of content different from the first category of content.

[0019] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a user via one or more input devices; and, in conjunction with detecting the input corresponding to the user, displaying via the display component a representation of a first portion of content related to the input and a representation of a second portion of content related to the input, including: visually grouping the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to a first category of content and the second portion of content belongs to a second category of content different from the first category of content; and abandoning the visual grouping of the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to the first category of content and the second portion of content belongs to a second category of content different from the first category of content.

[0020] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a user via one or more input devices; and, in conjunction with detecting the input corresponding to the user, displaying via the display component a representation of a first portion of content related to the input and a representation of a second portion of content related to the input, including: visually grouping the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to a first category of content and the second portion of content belongs to a second category of content different from the first category of content; and abandoning the visual grouping of the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to the first category of content and the second portion of content belongs to a second category of content different from the first category of content.

[0021] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes components for performing each of the following steps: detecting input corresponding to a user via one or more input devices; and, in conjunction with the detected input corresponding to the user, displaying via the display component a representation of a first portion of content related to the input and a representation of a second portion of content related to the input, including: visually grouping the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to a first category of content and the second portion of content belongs to a second category of content different from the first category of content; and abandoning the visual grouping of the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to the first category of content and the second portion of content belongs to a second category of content different from the first category of content.

[0022] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a user via one or more input devices; and, in conjunction with detecting the input corresponding to the user, displaying via the display component a representation of a first portion of content related to the input and a representation of a second portion of content related to the input, including: visually grouping the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to a first category of content and the second portion of content belongs to a second category of content different from the first category of content; and abandoning the visual grouping of the representation of the first portion of content and the representation of the second portion of content based on determining that the first portion of content belongs to the first category of content and the second portion of content belongs to a second category of content different from the first category of content.

[0023] In some embodiments, a method is described that is performed at a computer system communicating with a display component and one or more input devices. In some embodiments, the method includes: detecting a request corresponding to a previous interaction via one or more input devices; and in response to detecting the request corresponding to the previous interaction, displaying a user interface via the display component, the user interface including: a first representation of a first application corresponding to the previous interaction; a first representation of a first response to the request, wherein the first response is derived from the previous interaction; and a second representation of a second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response differs from the second representation of the second response.

[0024] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request corresponding to a previous interaction via one or more input devices; and, in response to detecting a request corresponding to a previous interaction, displaying a user interface via the display component, the user interface including: a first representation of a first application corresponding to the previous interaction; a first representation of a first response to the request, wherein the first response originates from the previous interaction; and a second representation of a second response to the request, wherein the second response originates from the previous interaction, and wherein the first representation of the first response differs from the second representation of the second response.

[0025] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request corresponding to a previous interaction via one or more input devices; and, in response to detecting the request corresponding to the previous interaction, displaying a user interface via the display component, the user interface including: a first representation of a first application corresponding to the previous interaction; a first representation of a first response to the request, wherein the first response originates from the previous interaction; and a second representation of a second response to the request, wherein the second response originates from the previous interaction, and wherein the first representation of the first response differs from the second representation of the second response.

[0026] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a request corresponding to a previous interaction via one or more input devices; and, in response to detecting a request corresponding to a previous interaction, displaying a user interface via the display component, the user interface including: a first representation of a first application corresponding to the previous interaction; a first representation of a first response to the request, wherein the first response originates from the previous interaction; and a second representation of a second response to the request, wherein the second response originates from the previous interaction, and wherein the first representation of the first response differs from the second representation of the second response.

[0027] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes components for performing each of the following steps: detecting a request corresponding to a previous interaction via one or more input devices; and, in response to detecting a request corresponding to a previous interaction, displaying a user interface via the display component, the user interface including: a first representation of a first application corresponding to the previous interaction; a first representation of a first response to the request, wherein the first response originates from the previous interaction; and a second representation of a second response to the request, wherein the second response originates from the previous interaction, and wherein the first representation of the first response differs from the second representation of the second response.

[0028] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request corresponding to a previous interaction via one or more input devices; and, in response to detecting a request corresponding to a previous interaction, displaying a user interface via the display component, the user interface including: a first representation of a first application corresponding to the previous interaction; a first representation of a first response to the request, wherein the first response is derived from the previous interaction; and a second representation of a second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response is different from the second representation of the second response.

[0029] In some embodiments, a method is described that is executed at a computer system communicating with a display component. In some embodiments, the method includes: detecting a first request corresponding to a previous interaction; and in response to detecting the request corresponding to the previous interaction: based on determining that the request does not correspond to new content, displaying a first summary of the previous interaction via the display component, the first summary including one or more representations of the previous interaction in a first orientation relative to a second set of one or more representations corresponding to the previous interaction; and based on determining that the request includes new content, displaying a second summary of the previous interaction via the display component, the second summary including one or more representations of the first set of one or more representations of the previous interaction in a second orientation relative to the second set of one or more representations, wherein the second orientation differs from the first orientation.

[0030] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display component. In some embodiments, the one or more programs include instructions for: detecting a first request corresponding to a previous interaction; and in response to detecting the request corresponding to the previous interaction: based on determining that the request does not correspond to new content, displaying a first summary of the previous interaction via the display component, the first summary including one or more representations of the previous interaction in a first orientation relative to a second set of one or more representations corresponding to the previous interaction; and based on determining that the request includes new content, displaying a second summary of the previous interaction via the display component, the second summary including one or more representations of the first set of one or more representations of the previous interaction in a second orientation relative to the second set of one or more representations, wherein the second orientation differs from the first orientation.

[0031] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display component. In some embodiments, the one or more programs include instructions for: detecting a first request corresponding to a previous interaction; and in response to detecting the request corresponding to the previous interaction: based on determining that the request does not correspond to new content, displaying a first summary of the previous interaction via the display component, the first summary including one or more representations of the previous interaction in a first orientation relative to a second set of one or more representations corresponding to the previous interaction; and based on determining that the request includes new content, displaying a second summary of the previous interaction via the display component, the second summary including one or more representations of the first set of one or more representations corresponding to the previous interaction in a second orientation relative to the second set of one or more representations, wherein the second orientation differs from the first orientation.

[0032] In some embodiments, a computer system communicating with a display component is described. In some embodiments, the computer system communicating with the display component includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a first request corresponding to a previous interaction; and in response to detecting a request corresponding to a previous interaction: based on determining that the request does not correspond to new content, displaying a first summary of the previous interaction via the display component, the first summary including one or more representations of the previous interaction in a first orientation relative to a second set of one or more representations corresponding to the previous interaction; and based on determining that the request includes new content, displaying a second summary of the previous interaction via the display component, the second summary including one or more representations of the previous interaction in a second orientation relative to the second set of one or more representations, wherein the second orientation differs from the first orientation.

[0033] In some embodiments, a computer system communicating with a display component is described. In some embodiments, the computer system communicating with the display component includes components for performing each of the following steps: detecting a first request corresponding to a previous interaction; and in response to detecting a request corresponding to a previous interaction: based on determining that the request does not correspond to new content, displaying a first summary of the previous interaction via the display component, the first summary including one or more representations of the previous interaction in a first orientation relative to a second group of one or more representations corresponding to the previous interaction; and based on determining that the request includes new content, displaying a second summary of the previous interaction via the display component, the second summary including one or more representations of the first group of one or more representations of the previous interaction in a second orientation relative to the second group of one or more representations, wherein the second orientation is different from the first orientation.

[0034] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display component. In some embodiments, the one or more programs include instructions for: detecting a first request corresponding to a previous interaction; and in response to detecting a request corresponding to a previous interaction: based on determining that the request does not correspond to new content, displaying a first summary of the previous interaction via the display component, the first summary including one or more representations of the previous interaction in a first orientation relative to a second set of one or more representations corresponding to the previous interaction; and based on determining that the request includes new content, displaying a second summary of the previous interaction via the display component, the second summary including one or more representations of the first set of one or more representations corresponding to the previous interaction in a second orientation relative to the second set of one or more representations, wherein the second orientation is different from the first orientation.

[0035] In some embodiments, a method is described that is performed at a computer system communicating with one or more output devices and one or more input devices including a display component. In some embodiments, the method includes: displaying visual content via the display component comprising one or more items of a first group, one or more items of a second group different from the first group, and an avatar closer to the first group than the second group; while displaying the visual content comprising the first group, the second group, and the avatar closer to the first group than the second group, outputting content corresponding to the first group items via one or more output devices; while outputting the content corresponding to the first group items and displaying the avatar closer to the first group than the second group, detecting that the content corresponding to the second group items will be output; and in response to detecting that the content corresponding to the second group items will be output, displaying the avatar positioned closer to the second group than the first group items via the display component.

[0036] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system including one or more output devices and one or more input devices, comprising a display component. In some embodiments, the one or more programs include instructions for: displaying visual content via the display component comprising one or more items of a first group, one or more items of a second group different from the first group, and an avatar closer to the first group than the second group; outputting content corresponding to the first group items via one or more output devices while displaying the visual content comprising the first group items, the second group items, and the avatar closer to the first group than the second group; detecting that content corresponding to the second group items will be output while outputting content corresponding to the first group items and displaying the avatar closer to the first group than the second group; and displaying the avatar positioned closer to the second group than the first group items via the display component in response to detecting that content corresponding to the second group items will be output.

[0037] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system including one or more output devices and one or more input devices, comprising instructions for: displaying visual content via the display component comprising one or more items of a first group, one or more items of a second group different from the first group, and an avatar closer to the first group than the second group; outputting content corresponding to the first group items via one or more output devices while displaying the visual content comprising the first group items, the second group items, and the avatar closer to the first group than the second group; detecting that content corresponding to the second group items will be output while outputting content corresponding to the first group items and displaying the avatar closer to the first group than the second group; and displaying the avatar positioned closer to the second group than the first group items via the display component in response to detecting that content corresponding to the second group items will be output.

[0038] In some embodiments, a computer system is described that communicates with one or more output devices and one or more input devices including a display component. In some embodiments, the computer system communicating with one or more output devices and one or more input devices including a display component includes one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: displaying visual content via the display component comprising one or more items of a first group, one or more items of a second group different from the first group, and an avatar closer to the first group than the second group; while displaying the visual content comprising the first group, the second group, and the avatar closer to the first group than the second group, outputting content corresponding to the first group items via one or more output devices; detecting that content corresponding to the second group items will be output while outputting content corresponding to the first group and displaying an avatar closer to the first group than the second group; and in response to detecting that content corresponding to the second group items will be output, displaying the avatar positioned closer to the second group than the first group via the display component.

[0039] In some embodiments, a computer system is described that communicates with one or more output devices and one or more input devices including a display component. In some embodiments, the computer system communicating with one or more output devices and one or more input devices including a display component includes components for performing each of the following steps: displaying visual content via the display component comprising one or more items of a first group, one or more items of a second group different from the first group, and an avatar closer to the first group than the second group; while displaying the visual content comprising the first group, the second group, and the avatar closer to the first group than the second group, outputting content corresponding to the first group items via one or more output devices; while outputting the content corresponding to the first group items and displaying the avatar closer to the first group than the second group, detecting that the content corresponding to the second group items will be output; and in response to detecting that the content corresponding to the second group items will be output, displaying the avatar positioned closer to the second group than the first group items via the display component.

[0040] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system including one or more output devices and one or more input devices, comprising instructions for: displaying visual content via the display component comprising one or more items of a first group, one or more items of a second group different from the first group, and an avatar closer to the first group than the second group; while displaying the visual content comprising the first group, the second group, and the avatar closer to the first group than the second group, outputting content corresponding to the first group items via one or more output devices; while outputting the content corresponding to the first group and displaying the avatar closer to the first group than the second group, detecting that the content corresponding to the second group will be output; and in response to detecting that the content corresponding to the second group will be output, displaying the avatar positioned closer to the second group than the first group via the display component.

[0041] In some embodiments, a method is described that is performed at a computer system communicating with a display component and one or more input devices. In some embodiments, the method includes: detecting input corresponding to a topic via one or more input devices while a first user interface object is displayed via the display component; and in response to detecting the input corresponding to the topic: abandoning an attempt to increase the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input; and increasing the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level above a threshold corresponding to the input.

[0042] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a topic via one or more input devices when a first user interface object is displayed via the display component; and in response to detecting the input corresponding to the topic: abandoning any attempt to increase the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input; and increasing the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level above a threshold corresponding to the input.

[0043] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a topic via one or more input devices when a first user interface object is displayed via the display component; and in response to detecting the input corresponding to the topic: abandoning any attempt to increase the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input; and increasing the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level above a threshold corresponding to the input.

[0044] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a topic via one or more input devices when a first user interface object is displayed via the display component; and in response to detecting the input corresponding to the topic: abandoning the increase in the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input; and increasing the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level above a threshold corresponding to the input.

[0045] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes components for performing each of the following steps: when a first user interface object is displayed via the display component, detecting input corresponding to a topic via one or more input devices; and in response to detecting input corresponding to a topic: abandoning an attempt to increase the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input; and increasing the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level above a threshold corresponding to the input.

[0046] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a topic via one or more input devices when a first user interface object is displayed via the display component; and in response to detecting input corresponding to the topic: abandoning the increase of the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input; and increasing the size of the first user interface object based on determining that a corresponding portion of the input is associated with a confidence level above a threshold corresponding to the input.

[0047] In some embodiments, a method is described that is performed at a computer system communicating with a display component and one or more input devices. In some embodiments, the method includes: detecting a request to display an animation via one or more input devices; initiating playback of the animation via the display component in response to detecting the request to display the animation; detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame; and displaying the overlay at a second position different from the first position via the display component in response to detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, wherein the second position is selected after initiating the playback of the animation.

[0048] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request to display an animation via one or more input devices; initiating playback of the animation via the display component in response to detecting the request to display the animation; detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, while the animation is being played back and an overlay is displayed via the display component; and displaying the overlay at a second position different from the first position via the display component in response to detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, wherein the second position is selected after the playback of the animation is initiated.

[0049] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request to display an animation via one or more input devices; initiating playback of the animation via the display component in response to detecting the request to display the animation; detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, while the animation is being played back and an overlay is displayed via the display component; and displaying the overlay at a second position different from the first position via the display component in response to detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, wherein the second position is selected after the playback of the animation is initiated.

[0050] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a request to display an animation via one or more input devices; initiating playback of the animation via the display component in response to detecting the request to display the animation; detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, while the animation is being played back and an overlay is displayed via the display component; and displaying the overlay at a second position different from the first position via the display component in response to detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, wherein the second position is selected after initiating the playback of the animation.

[0051] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system communicating with the display component and one or more input devices includes components for performing each of the following steps: detecting a request to display an animation via one or more input devices; initiating playback of the animation via the display component in response to detecting the request to display an animation; detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame; and displaying the overlay at a second position different from the first position via the display component in response to detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, wherein the second position is selected after initiating the playback of the animation.

[0052] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request to display an animation via one or more input devices; initiating playback of the animation via the display component in response to detecting the request to display the animation; detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame; and displaying the overlay at a second position different from the first position via the display component in response to detecting that an object in the animation will be within a distance of the first position when the animation is displayed in the first frame, wherein the second position is selected after initiating the playback of the animation.

[0053] In some embodiments, a method is described that is performed at a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the method includes: detecting input via one or more input devices corresponding to a request to view one or more previous interactions with an agent; and in response to detecting input corresponding to a request to view one or more previous interactions with an agent: outputting a first representation of a first previous interaction with the agent via one or more output devices based on determining that a first set of one or more criteria is met; and abandoning the output of the first representation of the first previous interaction with the agent via one or more output devices based on determining that a second set of one or more criteria is met.

[0054] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input via one or more input devices corresponding to a request to view one or more previous interactions with an agent; and in response to detecting input corresponding to a request to view one or more previous interactions with an agent: outputting a first representation of a first previous interaction with an agent via one or more output devices based on determining that a first set of one or more criteria is met; and abandoning the output of the first representation of the first previous interaction with an agent via one or more output devices based on determining that a second set of one or more criteria is met.

[0055] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input via one or more input devices corresponding to a request to view one or more previous interactions with an agent; and in response to detecting the input corresponding to the request to view one or more previous interactions with the agent: outputting a first representation of a first previous interaction with the agent via one or more output devices based on determining that a first set of one or more criteria is met; and abandoning the output of the first representation of the first previous interaction with the agent via one or more output devices based on determining that a second set of one or more criteria is met.

[0056] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system communicating with a display component and one or more input devices includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting input via one or more input devices corresponding to a request to view one or more previous interactions with an agent; and in response to detecting input corresponding to a request to view one or more previous interactions with an agent: outputting a first representation of a first previous interaction with an agent via one or more output devices based on determining that a first set of one or more criteria is met; and abandoning the output of the first representation of the first previous interaction with an agent via one or more output devices based on determining that a second set of one or more criteria is met.

[0057] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system communicating with a display component and one or more input devices includes components for performing each of the following steps: detecting input via one or more input devices corresponding to a request to view one or more previous interactions with an agent; and in response to detecting input corresponding to a request to view one or more previous interactions with an agent: outputting a first representation of a first previous interaction with an agent via one or more output devices based on determining that a first set of one or more criteria is met; and abandoning the output of the first representation of the first previous interaction with an agent via one or more output devices based on determining that a second set of one or more criteria is met.

[0058] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input via one or more input devices corresponding to a request to view one or more previous interactions with an agent; and in response to detecting input corresponding to a request to view one or more previous interactions with an agent: outputting a first representation of a first previous interaction with an agent via one or more output devices based on determining that a first set of one or more criteria is met; and abandoning the output of the first representation of the first previous interaction with an agent via one or more output devices based on determining that a second set of one or more criteria is met.

[0059] Executable instructions for performing these functions may optionally be included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Attached Figure Description

[0060] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, in which similar reference numerals indicate corresponding parts throughout the drawings.

[0061] Figure 1 This is a block diagram illustrating a computer system according to some implementation schemes.

[0062] Figures 2A to 2C These are illustrations of exemplary components and user interfaces of an electronic device 200 according to some implementation schemes.

[0063] Figure 3 This is a block diagram illustrating exemplary components of a device according to some implementation schemes.

[0064] Figure 4 This is a functional diagram of an exemplary actuator device according to some implementation schemes.

[0065] Figure 5 This is a functional diagram of an exemplary agent system based on some implementation schemes.

[0066] Figures 6A to 6D Exemplary user interfaces for engaging in interaction are illustrated according to some implementation schemes.

[0067] Figure 7 This is a flowchart illustrating a method for mobile positioning according to some implementation schemes.

[0068] Figure 8 This is a flowchart illustrating a method for displaying content according to some implementation schemes.

[0069] Figures 9A to 9J An exemplary user interface for controlling the user interface is illustrated according to some implementation schemes.

[0070] Figure 10 This is a flowchart illustrating a method for grouping content according to some implementation schemes.

[0071] Figure 11 This is a flowchart illustrating a method for displaying a response in response to a request corresponding to a previous interaction, according to some implementation schemes.

[0072] Figure 12 This is a flowchart illustrating a method for displaying a summary of previous interactions, according to some implementation schemes.

[0073] Figure 13 This is a flowchart illustrating methods for increasing the size of an object according to some implementation schemes.

[0074] Figure 14 This is a flowchart illustrating a method for displaying an incarnation closer to a set of items, according to some implementation schemes.

[0075] Figures 15A to 15D An exemplary user interface for displaying overlays is illustrated according to some implementation schemes.

[0076] Figure 16 This is a flowchart illustrating a method for displaying overlays according to some implementation schemes. Detailed Implementation

[0077] The following description illustrates exemplary methods, components, parameters, etc. While specific examples are described below, it should be understood that such examples should not be construed as limiting the scope of this disclosure to the explicit descriptions of the examples set forth herein, but rather as providing illustrative examples.

[0078] Each of the modules and applications identified herein corresponds to a set of executable instructions for performing one or more functions described above and methods described in this application (e.g., computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) may optionally not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and therefore various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. For example, a video player module may optionally be combined with a music player module into a single module. In some embodiments, memory may optionally store a subset of the modules and data structures identified above. Furthermore, memory may optionally store additional modules and data structures not described above.

[0079] One or more steps of the method described herein may depend on satisfying one or more conditions. In some embodiments, the method is performed through multiple iterative processes. In some embodiments, the conditional steps may be satisfied in different iterations of the same process and still remain within the scope of the method described herein. For example, for a given method comprising two steps depending on different conditions, those skilled in the art will understand that the given method should be considered performed even if the process is repeated multiple times until the conditional step is satisfied. In some embodiments, multiple iterations of the process are not required to practice the claims as set forth herein. For example, the claims of an electronic device, system, or computer-readable medium may be performed without iteratively repeating the process. In some embodiments, the claims of an electronic device, system, or computer-readable medium include instructions for performing one or more steps depending on satisfying one or more conditions. Because such instructions are stored in one or more processors and / or one or more memory locations, the claims of an electronic device, system, or computer-readable medium may include logic for determining whether one or more conditions have been satisfied without requiring the steps of the process to be repeated.

[0080] Although numerical descriptors such as "first" and / or "second" are used below to describe elements, these elements do not correspond to sequential or different representations and should not be limited to the stated numerical terms. In some embodiments, these terms are used only as prefixes to distinguish references to one element from references to another. For example, "first" device and "second" device can be two separate references to the same device. Conversely, for example, "first" device and "second" device can be references to two different devices (e.g., not the same device and / or not the same type of device). For example, a first computer system and a second computer system do not correspond to first and second in time and are merely used to distinguish the two computer systems. Therefore, without departing from the scope of the various described embodiments, a first computer system may be referred to as a second computer system, and a second computer system may be referred to as a first computer system.

[0081] In the description of various elements and examples, the use of certain terms is intended to provide a productive description of the following topics and should not be construed as restrictive. As used in describing the various examples herein, the singular forms “a,” “an,” and “the” should not be construed as excluding or precluding the plural forms unless the context clearly indicates otherwise. Similarly, “and / or” is used to cover any and all possible combinations of one or more associated listed items. For example, “x and / or y” should be interpreted as including “x” or “y” as well as “x and y” as a possible permutation. Furthermore, the terms “includes,” “including,” “comprises,” and / or “comprising” used in this specification specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0082] When describing choices and / or logical possibilities, the term "if" may optionally be interpreted, depending on the context, as meaning "when," "in response to determination," "in response to detection," or "according to determination." Similarly, depending on the context, the phrases "if determination..." or "if [the stated condition or event] is optionally interpreted as meaning "in response to determination," "in response to detection," "in response to detection," or "according to determination."

[0083] The processes described below enhance device operability and make user-device interfaces more efficient through a variety of technologies (e.g., by helping users provide correct input and reducing user errors when operating / interacting with the device). These technologies include: providing users with improved feedback (e.g., visual, tactile, audible, and / or haptic feedback); reducing the amount of input required to perform an operation; providing additional control options without cluttering the user interface with additional displayed controls; performing an operation without requiring additional input (e.g., user input) when a set of conditions has been met; and / or additional technologies (such as improving the security and / or privacy of the computer system and reducing the aging of one or more parts of the user interface on the display). These technologies also reduce power consumption and extend device battery life by enabling users to use the device more quickly and efficiently.

[0084] under, Figure 1 , Figures 2A to 2C and Figures 3 to 5 A description of an exemplary device for performing the techniques described herein is provided. Figures 6A to 6D Exemplary user interfaces for engaging in interaction are illustrated according to some implementation schemes. Figure 7 This is a flowchart illustrating a method for mobile positioning according to some implementation schemes. Figure 8 This is a flowchart illustrating a method for displaying content according to some implementation schemes. Figures 6A to 6D The user interface in the document is used to illustrate the processes described below, including Figure 7 and Figure 8 The process in. Figures 9A to 9J An exemplary user interface for controlling the user interface is illustrated according to some implementation schemes. Figure 10 This is a flowchart illustrating a method for grouping content according to some implementation schemes. Figure 11 This is a flowchart illustrating a method for displaying a response in response to a request corresponding to a previous interaction, according to some implementation schemes. Figure 12 This is a flowchart illustrating a method for displaying a summary of previous interactions, according to some implementation schemes. Figure 13 This is a flowchart illustrating methods for increasing the size of an object according to some implementation schemes. Figure 14 This is a flowchart illustrating a method for displaying an incarnation closer to a set of items, according to some implementation schemes. Figures 9A to 9J The user interface in the document is used to illustrate the processes described below, including Figure 10 , Figure 11 , Figure 12 , Figure 13 and Figure 14 The process in. Figures 15A to 15D An exemplary user interface for displaying overlays is illustrated according to some implementation schemes. Figure 16This is a flowchart illustrating a method for displaying overlays according to some implementation schemes. Figures 15A to 15D The user interface in the document is used to illustrate the processes described below, including Figure 16 The process in.

[0085] Figure 1 A block diagram depicts a computer system 100 (e.g., an electronic device and / or electronic system) comprising a set of electronic components communicating (e.g., connected) with each other (e.g., wired or wireless). It should be understood that computer system 100 is merely one example of a computer system that can be used to perform the functions described below, and one or more other computer systems can be used to perform the functions described below. Furthermore, although... Figure 1 The computer architecture of computer system 100 is described, but other computer architectures of computer systems (e.g., including more components, similar components and / or fewer components) may be used to perform the functionality described herein.

[0086] In some implementations, computer system 100 may correspond to (e.g., is and / or includes) a system-on-a-chip, a server system, a personal computer system, a smartphone, a smartwatch, a wearable device, a tablet computer, a laptop computer, a fitness tracker, a head-mounted display (HMD) device, a desktop computer, public equipment (e.g., smart speakers, connected thermostats and / or additional home-based computer systems), accessories (e.g., switches, lights, speakers, air conditioners, heaters, window covers, fans, locks, media playback devices, televisions, etc.), controllers, hubs and / or sensors.

[0087] In some embodiments, the sensor includes one or more hardware components capable of detecting (e.g., sensing, generating, and / or processing) information about the physical environment near the sensor. For example, the sensor may be configured to detect information around the sensor, detect information in one or more directions extending outward from the sensor, and / or detect information based on contact between the sensor and elements of the physical environment. In some embodiments, the hardware components of the sensor include sensing components (e.g., temperature and / or image sensors), transmitting components (e.g., radio and / or laser transmitters), and / or receiving components (e.g., laser and / or radio receivers). In some implementations, the sensors include angle sensors, breakage sensors, flow sensors, force sensors, gas sensors, humidity or moisture sensors, glass breakage sensors, chemical sensors, contact sensors, non-contact sensors, image sensors (e.g., RGB cameras and / or infrared sensors), particle sensors, photoelectric sensors (e.g., ambient light and / or sunlight), positioning sensors (e.g., GPS), precipitation sensors, pressure sensors, proximity sensors, radiation sensors, inertial measurement units, leak sensors, liquid level sensors, metal sensors, microphones, motion sensors, distance or depth sensors (e.g., RADAR, LiDAR), speed sensors, temperature sensors, time-of-flight sensors, torque sensors, ultrasonic sensors, vacancy sensors, presence sensors, voltage and / or current sensors, conductivity sensors, resistivity sensors, capacitance sensors, and / or water sensors. Although in Figure 1 Only a single computer system is depicted, but the functionality described below can be implemented using two or more computer systems operating together. Additionally, in some embodiments, computer system 100 includes one or more sensors as described above, and captures information about the physical environment by combining data from one sensor with data from one or more additional sensors (e.g., which are part of the computer and / or one or more additional computer systems).

[0088] like Figure 1As illustrated, computer system 100 comprises processor subsystem 110, memory 120, and I / O interface 130. Memory 120 corresponds to system memory that communicates with processor subsystem 110. Electronic components constituting computer system 100 are electrically connected via interconnect 150, which allows communication between components of computer system 100. For example, interconnect 150 may be a system bus, one or more memory locations, and / or additional electrical channels for connecting multiple components of computer system 100. Additionally, I / O interface 130 is connected to I / O device 140 via wired and / or wireless connections. In some embodiments, computer system 100 includes a component comprising I / O interface 130 and I / O device 140, such that the functionality of each component is included within that component. Furthermore, it should be understood that computer system 100 may include one or more I / O interfaces that communicate with one or more I / O devices. In some embodiments, computer system 100 comprises multiple processor subsystems 100s, each processor subsystem being electrically connected via interconnect 150.

[0089] In some embodiments, processor subsystem 110 includes one or more processors or separate processing units capable of executing instructions (e.g., programs, systems, and / or interrupts) to perform the functionality described herein. For example, operating system-level and / or application-level instructions executed by processor subsystem 110. In some embodiments, processor subsystem 110 includes one or more components (e.g., implemented as hardware, software, and / or combinations thereof) capable of supporting, interpreting, and / or executing machine learning instructions and / or operations. For example, computer system 100 may perform operations locally based on a machine learning model. Alternatively or additionally, computer system 100 may communicate with (e.g., perform computations thereto and / or execute corresponding instructions) a remote interactive knowledge base (e.g., processing resources implementing machine learning models, artificial intelligence models, and / or large language models) to perform operations that may otherwise be outside the set of capabilities of computer system 100. For example, computer system 100 may determine a set of inputs (e.g., instructions, data, and / or parameters) to an interactive knowledge base for performing desired machine learning operations.

[0090] The memory 120, which communicates with the processor subsystem 110, can be implemented using a variety of different physical, non-transitory memory media. In some embodiments, the computer system 100 includes multiple memory components and / or various types of memory components, each of which is directly and / or connected to the processor subsystem 110 via interconnect 150. For example, the memory 120 can be implemented using removable flash drives, storage arrays, storage area networks (e.g., SANs), flash memory, hard disk storage devices, optical drive storage devices, floppy disk storage devices, removable disk storage devices, random access memory (e.g., SDRAM, DDR SDRAM, RAM-SRAM, EDO RAM, and / or RAMBUS RAM) and / or read-only memory (e.g., PROM and / or EEPROM). Additionally, in some embodiments, the processor subsystem 110 and / or interconnect 150 are connected to a memory controller, which is electrically connected to the memory 120.

[0091] In some embodiments, the instructions may be executed by processor subsystem 110. In this example, memory 120 may include a computer-readable medium (e.g., a non-transitory or transient computer-readable medium) that can be used to store (e.g., configured to store, assigned to store, and / or store) instructions executable by processor subsystem 110. In some embodiments, each instruction stored by memory 120 and executed by processor subsystem 110 corresponds to an operation for performing the functionality described herein. For example, memory 120 may store program instructions to implement the methods described below (including methods 700 and 800). Figure 7 and Figure 8 Related functionality.

[0092] As mentioned above, I / O interface 130 may be one or more types of interfaces that enable computer system 100 to communicate with other devices. In some embodiments, I / O interface 130 includes a bridge chip (e.g., a southbridge) connecting a front-side bus to one or more back-side buses. In some embodiments, I / O interface 130 enables communication with one or more I / O devices (exemplified as I / O device 140) via one or more corresponding buses or other interfaces. For example, I / O devices may include one or more of the following: physical user interface devices (e.g., physical keyboard, mouse, and / or joystick), storage devices (e.g., as described above with respect to memory 120), network interface devices (e.g., to a local area network or wide area network), sensor devices (e.g., as described above with respect to sensors), and / or auditory and / or visual output devices (e.g., screens, speakers, lamps, and / or projectors). In some embodiments, the visual output device is referred to as a display component. For example, a display component may be configured to provide visual output, such as displaying images on a physical visual medium via an LED display or image projection. As used herein, “display” content includes content that is displayed by sending data (e.g., image data and / or video data) to an integrated or external display component via a wired or wireless connection to visually generate content (e.g., video data rendered and / or decoded by a display controller).

[0093] In some embodiments, computer system 100 includes a component that integrates I / O device 140 with other components (e.g., a component including I / O interface 130 and I / O device 140). In some embodiments, I / O device 140 is separate from other components of computer system 100 (e.g., it is a discrete component). In some embodiments, I / O device 140 includes a network interface device that allows computer system 100 to connect to a network or other computer system (e.g., communicate with it) via wired or wireless means. In some embodiments, the network interface device may include Wi-Fi, Bluetooth, NFC, USB, Thunderbolt, Ethernet, etc. For example, computer system 100 may utilize NFC connectivity to facilitate banking, credit, financial, token (e.g., fungible or non-fungible tokens) and / or cryptocurrency transactions between computer system 100 and another nearby computer system.

[0094] In some embodiments, I / O device 140 includes components for detecting users (e.g., users, people, animals, another computer system different from the computer system, and / or objects) and / or input from the detected users (e.g., tap input and / or non-tap input (e.g., verbal input, audible requests, audible commands, audible statements, swipe input, press and drag input, gaze input, air gestures, and / or mouse clicks)). In some embodiments, I / O device 140 enables computer system 100 to identify users associated with and / or not having accounts within the environment. In some embodiments, computer system 100 may detect known users (e.g., users corresponding to accounts) and access information about the users using the known users' accounts. In some embodiments, as part of computer system 100's user detection, computer system 100 detects that a user's account is associated with a group of users (e.g., included in and / or identified relative to that group of users). For example, computer system 100 may access information associated with accounts in a family defined as a group of accounts in response to detecting a member of that family. In some implementations, the user's account may be connected to additional accounts and / or additional computer systems. For example, computer system 100 may detect such additional computer systems and / or detect such computer systems used to detect users. In some implementations, computer system 100 detects unknown users and enables guest accounts of unknown users to utilize computer system 100.

[0095] In some embodiments, I / O device 140 includes one or more cameras. In some embodiments, the camera includes an image sensor (e.g., one or more optical sensors and / or one or more depth camera sensors) that provides computer system 100 with the ability to detect user and / or user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, air gestures are gestures detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)). In some implementations, one or more cameras enable computer system 100 to send image and / or video information to an application. For example, image data captured by a camera can enable computer system 100 to complete a video call by sending video data to an application used to perform the video call.

[0096] In some embodiments, I / O device 140 includes one or more microphones. For example, the microphone may be used by 100 to obtain data and / or information from a user without contact input. In some embodiments, the microphone enables computer system 100 to detect verbal and / or speech input from a user. In some embodiments, computer system 100 utilizes speech input to enable personal assistant functionality. For example, a user makes a request to computer system 100 to perform an action and / or obtain information from the user. In some embodiments, computer system 100 utilizes speech input (e.g., in conjunction with one or more other input and / or output technologies) to request and / or detect information from a user without requiring physical contact between the user and computer system 100.

[0097] In some embodiments, I / O device 140 includes physical input media for a user to interact directly with computer system 100. In some embodiments, the physical input media includes one or more physical buttons (e.g., tactilely pressable buttons and / or touch-sensitive non-pressable components) on and / or connected to computer system 100, mouse and keyboard input methods (e.g., connected to computer system 100 together with and / or separately from one or more I / O interfaces), and / or touch-sensitive display components.

[0098] In some embodiments, I / O device 140 includes one or more components for outputting information (e.g., display components, audio generation components, speakers, haptic output devices, displays, projectors, and / or touch-sensitive displays). In some embodiments, computer system 100 uses I / O device 140 to transmit information and / or the state of computer system 100. In some embodiments, I / O device 140 includes haptic output components. For example, the haptic output component may be a haptic generation component that enables computer system 100 to transmit information to a user who is in contact with computer system 100 (e.g., holding, touching, and / or near it). In some embodiments, I / O device 140 includes one or more components for outputting visual output (e.g., video, images, animations, 3D rendering, augmented reality overlay, motion graphics, data visualization, digital art, etc.). For example, displaying content from one or more applications and / or system applications, and / or displaying widgets corresponding to one or more applications (e.g., controls displaying real-time information and / or data).

[0099] In some implementations, I / O device 140 includes one or more components for outputting audio (e.g., smart speaker, home theater system, soundbar, headphones, earphones, earbuds, speaker, TV speaker, augmented reality headphone speaker, audio jack, optical audio output, Bluetooth audio output, HDMI audio output, audio sensor, etc.). In some implementations, computer system 100 is capable of outputting audio through one or more speakers. For example, computer system 100 outputs audio-based content and / or information to a user. In some implementations, one or more speakers enable spatial audio (e.g., audio output corresponding to the environment (e.g., computer system 100 detects materials and / or objects in the environment and / or computer system 100 changes audio modes, intensities, and / or waveforms to compensate for changing environmental characteristics)).

[0100] Figure 2 to Figure 5Exemplary components and user interfaces of an electronic device 200 according to some embodiments are illustrated. The electronic device 200 (sometimes referred to herein as device 200) may include one or more features of the computer system 100. (Referring to Figures 2 to...) Figure 5 In the described example, device 200 is a laptop computer. In some embodiments, device 200 is not limited to a laptop computer, and those skilled in the art will recognize that device 200 can be one or more other devices (e.g., one or more of the components and / or functions described herein with respect to device 200). For example, device 200 can be a public device (such as a smart display, smart speaker, and / or television) and / or a personal device (such as a smartphone, smartwatch, tablet, desktop computer, fitness tracker, and / or head-mounted display). In some embodiments, the public device is configured to provide functionality to multiple users (e.g., simultaneously and / or at different times). In such embodiments, the public device can be managed and / or set by a single user. In some embodiments, the personal device is configured to provide functionality to a single user (e.g., once, such as when a single user logs into the personal device).

[0101] Figures 2A to 2C An example is shown of a device 200 located in three different physical locations. For example... Figure 2A As illustrated, device 200 is a laptop computer (also referred to herein as a "laptop"), which includes a base portion 200-2 (e.g., as shown in the image). Figure 2A The device 200 is horizontally placed on a surface such as a table and connected to a base portion 200-2 at a connection 200-3 (e.g., one or more connection points, motor arms, hinges, and / or joints). This connection allows the display portion 200-1 to pivot and / or change orientation relative to the base portion 200-2. For example, the device 200 may pivot at the connection 200-3 to rotate the display portion 200-1 and / or the device 200 to one or more positions corresponding to the “closed” internal state (e.g., as described below regarding...). Figure 2C(Further description). In some embodiments, the positioning corresponding to the "off" internal state is the positioning of the device 200 in a predetermined pose. For example, the predetermined pose may include a display portion 200-1 positioned parallel to the base portion 200-2 or forming a predetermined angle (e.g., 60 degrees) with respect to the base portion 200-2. In some embodiments, in the "off" internal state, the area of ​​the device 200 in which content is displayed is positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., an area facing downwards, not visible, and / or obscuring the displayed content). In some embodiments, in the "off" internal state, the area of ​​the device 200 in which content is displayed is not positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., instead positioned in a manner corresponding to the "on" internal state). For example, when not in a "closed" internal state, device 200 can be positioned within a range of different open positions (e.g., where display portion 200-1 is not parallel to base portion 200-2, and where the area where the content displayed by device 200 is visible and / or unobstructed). It should be recognized that display portion 200-1 being parallel to base portion 200-2 is an example of positioning corresponding to a "closed" internal state of device 200 (e.g., closed positioning). In some embodiments, another configuration may set another orientation of display portion 200-1 relative to base portion 200-2 as a closed positioning of device 200, such as... Figure 2C exemplified.

[0102] Figure 2A The left side illustrates display screen 200-4 (representing the area where device 200 displays content), and the right side illustrates device 200 in the corresponding pose. For example... Figure 2A As illustrated, device 200 is in a first position (e.g., display portion 200-1 is perpendicular to base portion 200-2, forming a 90-degree angle). Figure 2A In this context, display screen 200-4 represents the content currently being displayed (e.g., via a display component) when device 200 is first activated. Figure 2A In this embodiment, display screen 200-4 illustrates the device 200 in an "on" internal state (e.g., operable, powered, awake, higher power and / or more resource-intensive than the "off" state, and / or activated). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces (e.g., user interface objects, windows, application user interfaces, system user interfaces, controls, and / or other visual content). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces while in an "on" internal state. For example, in Figure 2AIn this configuration, device 200 is in an "on" internal state, and display screen 200-4 shows a desktop user interface 200-5, including an application window. In some embodiments, the user interface includes (and / or) one or more user interface objects (e.g., windows, icons, and / or other graphical objects). For example, the user interface (e.g., 200-5) may include one or more graphical objects that are different from and / or the same as the application window.

[0103] Figure 2B Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2B As illustrated, device 200 is in a second position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming an angle of 120 degrees (e.g., more than). Figure 2A (at a larger angle). Figure 2B In the diagram, display screen 200-4 represents the content being displayed when device 200 is in the second position. Display screen 200-4 illustrates the internal state of device 200 being "on" (e.g., with...). Figure 2A (The top diagram shows the same internal state). Figure 2B In the process, device 200 displays (e.g., via display screen 200-4) a desktop user interface 200-5 (e.g., with...). Figure 2A (The same as shown in the image). In some implementations, device 200 displays a different user interface (e.g., different from desktop user interface 200-5). For example, although... Figure 2B Example of device 200 in a state of being with Figure 2A Different positioning displays and Figure 2A The same desktop user interface 200-5 exists, but device 200 may display different user interfaces. In some embodiments, device 200 displays a user interface corresponding to (e.g., based on, due to, caused by, involved in, and / or configured to accompany) a physical state (e.g., positioning, location, and / or orientation), including content specific to a particular angle or specific to the current context.

[0104] Figure 2C Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2C As illustrated, device 200 is in a third position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming a 60-degree angle (e.g., compared to...). Figure 2A and Figure 2B (smaller angles)). Figure 2C In the diagram, display screen 200-4 shows the content being displayed when device 200 is in the third position. Figure 2CIn the diagram, displays 200-4 illustrate an internal state in which device 200 is "off" (e.g., not operating, not powered, not woken up, not activated, powered off, asleep, hibernating, inactive, and / or disabled). In some embodiments, device 200 does not display (e.g., via displays 200-4) one or more user interfaces (e.g., no visual content is displayed) when it is in the "off" internal state. In some embodiments, device 200 displays (e.g., via displays 200-4) one or more user interfaces (e.g., the same as and / or different from one or more user interfaces displayed when it is in the "on" internal state) (e.g., a user interface specific to the "off" state and / or a way of displaying a user interface not specific to the "off" internal state). Figure 2C In this case, display screen 200-4 is blank because nothing is displayed on the monitor of device 200 (e.g., display screen 200-4 is off and / or does not display the user interface) (e.g., desktop user interface 200-5 is not displayed on display screen 200-4).

[0105] In some embodiments, device 200 includes one or more components (referred herein also as “moving components”) that enable device 200 to perform (e.g., cause and / or control) movement (and / or be moved). For example, performing movement may include a portion of mobile device 200 (e.g., less or all components of the device moving), all of mobile device 200 (e.g., the entire device (including all its components) moving, such as by changing position), and / or moving one or more other devices and / or components (e.g., communicating with device 200 and / or the moving components of device 200). For example, device 200 may move automatically (e.g., pivot), cause and / or control movement of display portion 200-1 relative to base portion 200-2, such as moving to... Figures 2A to 2CAny of the illustrated locations. In some embodiments, device 200 performs movement based on its internal state. Performing movement based on internal state enables device 200 to perform new (e.g., otherwise unavailable) interactions. For example, such new interactions of device 200 can be configured using special features, functions, patterns, and / or procedures that leverage device 200's ability to perform movement. Examples of such interactions include using movement to (e.g., to a user) convey the device's internal state (e.g., on, off, sleep, and / or hibernate) to assist user input (e.g., shorten the distance to the user) and / or enhance the device's interactive behavior (e.g., moving in a specific manner during interaction with the user, conveying information such as importance and / or direction of attention). In some embodiments, the performed movement corresponds to (e.g., caused by, responded to, and / or determined and / or performed based on) one or more of the following: detected input, detected context (e.g., environmental context and / or user context), and / or the device 200's internal state (e.g., internal state and / or a set of multiple internal states). For example, device 200 can move the display portion, causing device 200 to move from a position where... Figure 2A The illustrated first positioning moves to the position where Figure 2B The illustrated second positioning. In this example, device 200 can detect that the user has repositioned relative to device 200 (e.g., the user stands up), and in response, device 200 can perform a movement to the second positioning such that the display is at an optimized viewing angle based on the height and / or angle of the user's eye relative to the display of device 200. As another example, device 200 can perform a movement such that device 200 moves from a position... Figure 2A The illustrated first positioning moves to the position where Figure 2C The illustrated third location. In this example, device 200 may perform a movement to a third location in response to detecting an internal state with reduced activity (e.g., an "off" internal state as described above). In this way, movement of device 200 to one or more locations can indicate the internal state of device 200.

[0106] Figures 2A to 2C An example is illustrated of a device 200 having a display portion capable of moving with one degree of freedom via a connection 200-3 (e.g., a hinge) connecting the display portion 200-1 to a base portion 200-2. In some embodiments, the device 200 includes one or more components having one or more degrees of freedom. For example, a moving component of the device 200 (e.g., an output component that causes and / or allows movement) (e.g., Figure 5Device 200-26C may include multiple degrees of freedom (e.g., six degrees of freedom including three translational components and three rotational components). For example, device 200 may be implemented to move the display portion by telescopic forward or backward movement (e.g., display portion 200-1 moves forward in space relative to the base portion while the base portion 200-2 remains stationary (e.g., to shorten and / or lengthen the user's viewing distance)). As yet another example, device 200 may be implemented to move the display portion to rotate about an axis perpendicular to the hinge, such that the display portion can rotate to position the display to follow the user as the user walks around device 200. Although Figures 2A to 2C The example shown illustrates a hinge, but other moving components may be included in device 200, such as actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), movable bases, rotatable components, and / or rotatable bases. In some embodiments, one or more moving components may enable device 200 to move in different ways, such as rotation (e.g., 0 to 360 degrees), lateral movement (e.g., to the right, left, down, up, and / or any combination thereof), and / or tilting (e.g., 0 to 360 degrees).

[0107] Figure 3 An exemplary block diagram of device 200 is illustrated. In some embodiments, device 200 includes... Figure 1 A, Figure 1 B. Figure 3 and Figure 5 B describes some or all of the components. For example... Figure 3 As illustrated, device 200 has a bus 200-13 that operatively couples I / O segments 200-12 (also referred to as I / O sub-segments and / or I / O interfaces) to processor 200-11 and memory 200-10. For example... Figure 3 As illustrated, I / O section 200-12 is connected to output device 200-16 (also referred to herein as "output component"). In some embodiments, output device 200-16 includes one or more visual output devices (e.g., display components such as monitors, displays, projectors, and / or touch-sensitive displays), one or more tactile output devices (e.g., devices that cause vibration and / or other tactile outputs), one or more audio output devices (e.g., speakers), and / or one or more moving components (e.g., actuators, motors, mechanical linkages, devices that cause and / or allow movement, and / or one or more moving components as described above). Figure 3As illustrated, output device 200-16 includes two exemplary moving components (e.g., a movement controller 200-17 and an actuator 200-18). Actuator 200-18 can be any component that performs (e.g., partial and / or overall) physical movement of a device (e.g., device 200 and / or devices coupled to and / or in contact with that device). Movement controller 200-17 can be any component (e.g., a control device) that controls actuator 200-18 (e.g., provides control signals to it). For example, movement controller 200-17 can provide control signals that actuate actuator 200-18 (e.g., cause physical movement). In some embodiments, movement controller 200-17 includes one or more logic components (e.g., a processor), one or more feedback components (e.g., sensors), and / or one or more control components (e.g., for applying control signals, such as relays, switches, and / or control lines). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in the same device and / or component (e.g., a dedicated onboard motion controller 200-17 attached to the actuator 200-18). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in different devices and / or components (e.g., one or more processors 200-11 may serve as the motion controller 200-17 for the actuator 200-18). In some embodiments, the motion controller 200-17 and / or the actuator 200-18 are embodied in a device (or one or more devices) other than device 200 (e.g., device 200 is coupled to (e.g., temporarily and / or removably) another device and may instruct the motion controller 200-17 and / or the actuator 200-18 to control the other device). Actuator 200-18 can be used to induce one or more types of mechanical movement (e.g., linear and / or rotary movement) in one or more ways (e.g., using electric, magnetic, hydraulic and / or pneumatic power). Examples of actuator 200-18 may include electromechanical actuators, linear actuators and / or rotary actuators.

[0108] like Figure 3As illustrated, I / O section 200-12 is connected to input device 200-14. In some embodiments, input device 200-14 includes one or more visual input devices (e.g., cameras and / or light sensors), one or more physical input devices (e.g., buttons, sliders, switches, touch-sensitive surfaces, and / or rotatable input mechanisms), one or more audio input devices (e.g., microphones), and / or other input devices (e.g., accelerometers, pressure sensors (e.g., contact strength sensors), distance sensors, temperature sensors, GPS sensors, accelerometers, orientation sensors (e.g., compasses), gyroscopes, motion sensors, and / or biometric sensors). Furthermore, I / O section 200-12 may be connected to communication unit 200-15 for receiving application and operating system data using Wi-Fi, Bluetooth, Near Field Communication (NFC), cellular, and / or other wireless (and / or wired) communication technologies.

[0109] The memory 200-10 of the personal electronic device 200 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions, which, when executed by one or more computer processors 200-11, cause the computer processors to perform, for example, the techniques described below, including processes 700 and 800. Figure 7 and Figure 8 A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device. In some embodiments, the storage medium is a transient computer-readable storage medium. In some embodiments, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, and Blu-ray technologies, and persistent solid-state storage such as flash memory and solid-state drives. Electronic device 200 is not limited to... Figure 3 The components and configurations may include, but may include, other and / or additional components in a variety of possible configurations, all of which are intended to fall within the scope of this disclosure.

[0110] Figure 4 A functional diagram of actuator 200-18B according to some embodiments is illustrated. As described above, actuator 200-18B can be any component that performs physical movement. In some embodiments, actuator 200-18B is operated using inputs including control signal 200-18A and / or energy source 200-18B. For example, actuator 200-18B can be a rotary actuator that converts electrical energy into rotational movement. This rotational movement can cause the aforementioned... Figures 2A to 2CThe movement of the display portion of the described device 200 (e.g., the counterclockwise rotation of the actuator moves the device 200 to a position with a large angle) Figure 2B The illustrated second positioning), and the clockwise (e.g., counterclockwise) rotational movement of the actuator moves the device 200 to a positioning with a smaller angle (e.g., Figure 2C (The illustrated third positioning). Control signal 200-18A may indicate one or more start and / or stop commands, movement and / or actuation direction, movement and / or actuation speed, movement and / or actuation time, target positioning (e.g., pose and / or position) of movement and / or actuation, and / or one or more other characteristics of movement and / or actuation. In some embodiments, the control signal and the energy source are the same signal and / or input. In some embodiments, one or more additional components (e.g., mechanical and / or electrical) (e.g., removably or permanently) are coupled to actuator 200-18B to influence movement and / or actuation (e.g., mechanical linkages such as lead screws, gears, and / or other components for changing (e.g., switching) the characteristics of movement and / or actuation). In some embodiments, actuator 200-18B includes one or more feedback components (e.g., a positioning sensor, an encoder, an overcurrent sensor, and / or a force sensor) that form part of a feedback loop for modifying and / or stopping movement and / or actuation (e.g., slowing down actuation upon reaching a target position and / or stopping actuation if physical resistance to actuation is detected via a sensor). In some embodiments, one or more feedback components are included (e.g., partially and / or entirely) in a motion controller (e.g., motion controller 200-13) operatively coupled to the actuator.

[0111] Now turn attention to the functionality (e.g., features and / or capabilities) of one or more devices (e.g., computer system 100 and / or electronic device 200). One such functionality is the implementation of an “agent,” which may alternatively be referred to as a software agent, intelligent agent, interactive agent, virtual assistant, intelligent virtual assistant, interactive virtual assistant, personal assistant, intelligent personal assistant, interactive personal assistant, intelligent interactive personal assistant, and / or artificial intelligence (AI) assistant. In some implementations, an agent refers to one or more sets of functions implemented in hardware and / or software (e.g., local and / or remote) on an agent system (e.g., a single device and / or multiple devices). In some implementations, the agent performs operations to perceive the environment, acquire knowledge, retrieve knowledge, learn skills, interact with the user, and / or perform tasks. The agent may perform these (and / or other) operations, for example, in response to user input and / or automatically (e.g., at an appropriate time determined based on the perceived context). An incomplete list of exemplary operations that an agent may be used for and / or used with includes: tracking a user’s eyes, face, and / or body (e.g., to move with the user and / or identify the user’s intentions and / or activities); detecting, identifying, and / or classifying users in the environment; detecting and / or responding to input (e.g., verbal input, air gestures, and / or physical input, such as touch input and / or force input to physical hardware components (e.g., buttons, knobs, and / or sliders); detecting context (e.g., user context, operational context, and / or environmental context); moving (e.g., changing pose, orientation, orientation, and / or location); performing one or more operations in response to input, context, and / or stimuli (e.g., objects or events that elicit one or more responsive operations on the device (e.g., outside and / or inside the device)); providing intelligent interaction capabilities (e.g., in part due to one or more machine learning (“ML”) models, such as large language models (“LLM”)) to respond to and / or perform operations; and / or (e.g., automatically and / or intelligently) performing tasks (e.g., a set of operations for achieving a specific goal). In some implementations, the agent performs actions in response to contactless input (e.g., air gestures and / or natural language commands). The foregoing list is intended to exemplify the actions that can be performed by the agent, but is not intended to be an exhaustive list. Other actions fall within the intended scope of the agent's capabilities. Furthermore, for the purposes of this disclosure, the agent need not include all the functionalities mentioned herein, but may include fewer or more functionalities (e.g., the agent may be implemented on an agent system that does not have mobile functionality but otherwise includes an intelligent personal assistant capable of interacting with the user).

[0112] In some implementations, a user is one or more of a user, person, object, and / or animal in an environment (e.g., a device) that is perceived (e.g., detected by the device, one or more other devices, and / or one or more of its components). In some implementations, a user is an entity that is perceived (e.g., detected by the device, one or more other devices, and / or one or more of its components). In some implementations, an entity is something distinguishable from surrounding entities (e.g., components of the environment and / or other users) and / or something that is considered to be a discrete logical construct via one or more components (e.g., a sensing component and / or other components). In some implementations, a user is physical and / or virtual. For example, a physical user may represent a user standing in front of the device and perceived by the device. As another example, a virtual user may represent an avatar in a virtual scene perceived by the device (e.g., an avatar detected in a media stream received by the device and / or captured by the device's camera). Although presented above as an example of “user,” throughout this disclosure the terms and / or concepts referred to as “user,” “person,” “object,” and / or “animal” are interchangeable with “user” unless otherwise expressly indicated.

[0113] As an example, and to revisit Figures 2A to 2C An agent, at least partially implemented on device 200, can perform operations that cause the display portion 200-1 of device 200 to move relative to the base portion 200-2. For example, the agent detects (e.g., senses and determines that it has occurred) contexts including a user standing up (e.g., based on face detection and tracking); and in response, the agent causes device 200 to open and / or device 200 to open the display portion 200-1 to a greater angle. As another example, the agent can detect verbal input corresponding to (e.g., interpreted as and / or implying including) a request to move the display (e.g., “Please move my display” or “Please enter sleep mode”); and in response, the agent causes device 200 to move and / or device 200 to move the display portion 200-1.

[0114] Figure 5 A functional diagram of an exemplary agent system 200-20 is shown. Figure 5 As illustrated, agent system 200-20 has a dashed box boundary surrounding input component 200-22, agent component 200-24, and output component 200-26. In some embodiments, agent system 200-20 includes more than Figure 5 Fewer, more, and / or different components are illustrated. In some embodiments, agent system 200-20 is implemented on a single device (e.g., computer system 100 and / or electronic device 200). In some embodiments, agent system 200-20 is implemented on multiple devices. In some embodiments, in Figure 5One or more components of the agent system 200-20 illustrated and / or described with respect to this figure are external to but operatively coupled to the agent system (e.g., accessories, external devices, external sensors, external actuators, external display components, external speakers, and / or external databases). In some embodiments, one or more components of the agent system 200-20 are local to one or more other components of the agent system 200-20. In some embodiments, one or more components of the agent system 200-20 are remote from one or more other components of the agent system 200-20.

[0115] In some implementations, input components 200-22 include components for performing sensing and / or communication functions of the agent system 200-20. For example... Figure 5 As illustrated, input components 200-22 include one or more sensors 200-22A. The one or more sensors 200-22A may include any components for detecting data corresponding to the physical environment. Examples of the one or more sensors 200-22A may include: cameras, light sensors, microphones, accelerometers, positioning sensors, pressure sensors, temperature sensors, olfactory sensors, and / or contact sensors. This list is not intended to be exhaustive, and the one or more sensors 200-22A may include other sensors not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to detect data corresponding to the physical environment. Figure 5 As illustrated, input component 200-22 includes one or more communication components 200-22B. The one or more communication components 200-22B may include any component (e.g., antenna, modem, network interface component, encoder, decoder, and / or communication protocol stack) for transmitting and / or receiving communications within and / or outside the agent system 200-20. Communication components 200-22B may be between different devices and / or between components within the same device. Communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, input component 200-22 includes more than Figure 5 The components illustrated herein may be fewer, more, and / or different. In some implementations, input components 200-22 are implemented in hardware and / or software.

[0116] In some implementations, agent components 200-24 include components that manage and / or perform the agency functions of agent system 200-20. For example... Figure 5As illustrated, agent components 200-24 include the following functional components: task flow, coordination and / or orchestration component 200-24A, management component 200-24B, perception component 200-24C, evaluation component 200-24D, interaction component 200-24E, strategy and decision-making component 200-24F, knowledge component 200-24G, learning component 200-24H, model component 200-24I, and API component 200-24J. Each of these components is briefly described below. It is important to note that this list of agent components 200-24 is not intended to be exhaustive, and agent components 200-24 may include other functional components not explicitly identified herein, which may be used (e.g., to process, store, and / or transform) any functions of the agent, such as those described herein. In some embodiments, agent components 200-24 include more than Figure 5 The illustrated components may be fewer, more, and / or different. In some implementations, agent components 200-24 are implemented in hardware and / or software.

[0117] In some implementations, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to handle coordination between various components. For example, operations may include handling data processing task flows to move from perception components 200-24C (e.g., those detecting speech input) to model components 200-24I (e.g., for processing the detected speech input using a large language model to determine the content and / or intent of the speech input). In some implementations, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to handle coordination between one or more external components (e.g., resources). For example, Figure 5 Examples of external components (such as external databases 200-30) are illustrated. In some embodiments, management component 200-24B includes functionality performed by the operating system of the device implementing agent system 200-20. In some embodiments, management component 200-24B includes functionality performed by one or more applications of the device implementing agent system 200-20.

[0118] In some implementations, management component 200-24B performs operations that enable the agent system to handle administrative tasks, such as managing system and / or component updates, managing user accounts, and managing system settings and / or component settings. In some implementations, management component 200-24B includes functionality performed by the operating system of the device implementing agent system 200-20. In some implementations, management component 200-24B includes functionality performed by one or more applications of the device implementing agent system 200-20.

[0119] In some embodiments, the sensing components 200-24C perform operations that enable the agent to perceive environmental input. For example, operations may include detecting that context and / or environmental conditions have occurred, detecting the presence of a user (e.g., a person, object, and / or animal in the environment), detecting input including speech, detecting input including air gestures, detecting facial expressions, detecting user characteristics (e.g., visible and / or invisible), and / or detecting verbal and / or physical cues. In some embodiments, the sensing components 200-24C include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the sensing components 200-24C include functionality performed by one or more applications of the device implementing the agent system 200-20.

[0120] In some embodiments, evaluation component 200-24D performs operations that enable the agent to process evaluation data (e.g., to determine context, such as user context, environmental context, and / or operational context). For example, operations may include evaluating data collected from perception component 200-24C, knowledge component 200-24G, external database 200-30, and / or remote processing resource 200-32. In some embodiments, evaluation component 200-24D includes functionality performed by the operating system of the device implementing agent system 200-20. In some embodiments, evaluation component 200-24D includes functionality performed by one or more applications of the device implementing agent system 200-20.

[0121] This document refers to an environmental context (also referred to herein as "context of the environment" and / or "context corresponding to the environment"). In some embodiments, an environmental context is a context based on one or more characteristics of the environment (e.g., user, location, time, weather, and / or lighting). For example, an environmental context may include rain outside, daytime, and / or the device currently being in a park. In some embodiments, the device (e.g., using an agent) uses one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device) to determine the environmental context (e.g., currently true, happening, and / or applicable).

[0122] This document refers to user context (also referred to herein as "user context" and / or "context corresponding to the user") (and / or user context). In some embodiments, user context is a context based on one or more characteristics of the user. In some embodiments, user context may include the user's appearance and / or clothing, personality, actions, behaviors, movement, location, and / or pose. In some embodiments, the device (e.g., using an agent) determines user context (e.g., currently true, happening, and / or applicable) using one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, the device determines user context based on historical context and / or learned user characteristics, wherein one or more user characteristics are learned and / or stored by the device over a period of time.

[0123] This document refers to an operational context (also referred to herein as "the context of operation" and / or "operational context"). In some embodiments, an operational context is a context based on one or more characteristics of the device's operation (e.g., the device and / or one or more other devices that determine and / or access the operational context). For example, an operational context may include the internal state of the device (and / or one or more components of the device), the device's internal dialogue (e.g., the device's understanding of the context), the operations performed by the device, and applications and / or processes executed on the device (e.g., running and / or opening). In some embodiments, the device (e.g., using an agent) uses one or more of the following to determine the operational context (e.g., currently true, happening, and / or applicable): detected input (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, the device (e.g., using an agent) uses one or more internal states (e.g., accessed, retrieved, and / or queried by the device's processes) to determine the operational context (e.g., currently true, happening, and / or applicable).

[0124] In some embodiments, interaction components 200-24E perform operations that enable the agent to manage and / or perform interactions with the user. In some embodiments, operations may include determining an appropriate interaction model for a specific context and / or in response to specific input. In some embodiments, interaction components 200-24E include functionality performed by the operating system of the device implementing agent system 200-20. In some embodiments, interaction components 200-24E include functionality performed by one or more applications of the device implementing agent system 200-20.

[0125] In some implementations, the policy and decision components 200-24F perform operations that enable the agent to take actions based on available data. For example, operations may include determining which actions to perform and / or which functional components to utilize in response to detected context. In some implementations, the policy and decision components 200-24F include functionality performed by the operating system of the device implementing the agent system 200-20. In some implementations, the policy and decision components 200-24F include functionality performed by one or more applications of the device implementing the agent system 200-20.

[0126] In some implementations, the knowledge component 200-24G performs operations that enable the agent to access and use the stored knowledge. For example, operations may include indexing, storing, and / or retrieving data from a data repository, database, and / or other resource. In some implementations, the knowledge component 200-24G includes functionality performed by the operating system of the device implementing the agent system 200-20. In some implementations, the knowledge component 200-24G includes functionality performed by one or more applications of the device implementing the agent system 200-20.

[0127] In some implementations, the learning components 200-24H perform operations that enable the agent to learn through experience. For example, operations may include observing and / or tracking data, including preferences, routines, user characteristics, and / or environmental characteristics, in a way that allows the data to inform the agent and / or its components about future actions (e.g., when performing tasks and / or interacting with a user). In some implementations, the learning components 200-24H include functionality performed by the operating system of the device implementing the agent system 200-20. In some implementations, the learning components 200-24H include functionality performed by one or more applications of the device implementing the agent system 200-20.

[0128] In some implementations, model components 200-24I perform operations that enable the agent to apply ML models (e.g., large language models (LLMs)) to process data. For example, operations may include storing the ML model, executing the ML model, training and / or retraining the ML model, and / or otherwise managing aspects of implementing the ML model. In some implementations, model components 200-24I include functionality performed by the operating system of the device implementing agent system 200-20. In some implementations, model components 200-24I include functionality performed by one or more applications of the device implementing agent system 200-20.

[0129] In some implementations, agent system 200-20 responds to natural language input. For example, agent system 200-20 responds to natural language input in the form of statements, questions, commands, and / or requests. In some implementations, agent system 200-20 outputs text and / or speech output provided in natural language or mimicking a natural language style. For example, agent system 200-20 may use a speech response indicating the current outside temperature at the user's location (e.g., "18 degrees outside") to handle the natural language question "How hot is it outside?". In some implementations, agent system 200-20 responds to natural language input by providing information (e.g., weather, travel, and / or calendar information) and / or performing tasks (e.g., opening a document, searching a database, and / or opening an application).

[0130] In some implementations, agent system 200-20 includes and / or relies on one or more data models to process inputs (e.g., natural language input, gesture input, visual input, and / or other data input) and / or provide outputs (e.g., information output via natural language output, visual output, audio output, and / or text output). Such data models may include user data (e.g., data based on a specific interaction and / or from the user with whom the interaction takes place) and / or global data (e.g., general data based on the interaction and / or data from many users) and / or be trained using user data and / or global data. For example, user data (e.g., preferences, prior use of language and / or phrases, calendar entries, contact lists, and / or activity data) can be used to better infer user intent and / or provide responses more likely to resolve user requests. In some implementations, the data models used by agent system 200-20 include one or more machine learning components (e.g., hardware and / or software) (e.g., one or more neural networks), which are used by and / or implemented using them. Such machine learning components can be used to process spoken input to determine words and / or phrases therein, one or more contexts corresponding to the words, user intent corresponding to the words, one or more confidence scores, and / or a set of one or more actions to be taken in response to the spoken input. Similar operations can be performed to process other types of input, such as visual input, data input, and / or text input. Such data models may include machine learning and / or data processing models, including but not limited to natural language processing models, language models, speech recognition models, object recognition models, visual processing models, ontology, task flow models, and / or intent recognition models (e.g., for determining user intent).

[0131] In some implementations, the Application Programming Interface (API) components 200-24J perform operations that enable the agent to interface with services, devices, and / or components. For example, operations may include relaying data (e.g., requests, responses, and / or other messages) between data interfaces (e.g., between software programs, between system processes and application processes, between system processes, between application processes, between communication protocols, between clients and servers, between file systems, and / or between components on different sides of a trust boundary). In some implementations, the data interfaces served by the API components 200-24J are local (e.g., for a device, such as two application processes exchanging data) and / or remote (e.g., from a device, such as interfacing with a web service via a remote server). In some implementations, the API components 200-24J include functionality performed by the operating system of the device implementing the agent system 200-20. In some implementations, the API components 200-24J include functionality performed by one or more applications of the device implementing the agent system 200-20.

[0132] In some implementations, output components 200-26 include components for performing the output functions of agent system 200-20. A brief description follows. Figure 5 The exemplary output components are illustrated herein. In some embodiments, output components 200-26 include... Figure 5 The components illustrated may include fewer components, more components, and / or different components. In some implementations, the input components are implemented in hardware and / or software.

[0133] like Figure 5 As illustrated, output components 200-26 include one or more visual output components 200-26A. One or more visual output components 200-26A may include any component used for outputting (e.g., generating, creating, and / or displaying) and / or causing visual output (e.g., visually perceptible output, such as a graphical user interface, playback of visual media content, and / or lighting). Examples of one or more visual output components 200-26A may include: display components, projectors, head-mounted displays (HMDs), light-emitting diodes (“LEDs”), and / or components that create visually perceptible effects (e.g., movement). This list is not intended to be exhaustive, and one or more visual output components 200-26A may include other visual output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output visual output.

[0134] like Figure 5As illustrated, output components 200-26 include one or more audio output components 200-26B. One or more audio output components 200-26B may include any component for outputting (e.g., generating and / or creating) and / or causing audio output (e.g., audibly perceptible output, such as sound, music, speech, and / or audio media content). Examples of one or more audio output components 200-26B may include: speakers, audio amplifiers, tone generators, and / or components that produce audibly perceptible effects (e.g., movement, such as vibration). This list is not intended to be exhaustive, and one or more audio output components 200-26B may include other audio output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) the output audio output.

[0135] like Figure 5 As illustrated, output components 200-26 include one or more motion output components 200-26C (also referred to herein as "motion components"). One or more motion output components 200-26C may include any component for outputting (e.g., generating and / or creating) and / or causing motion output (e.g., output including physical movement of a device and / or another device / component). Examples of one or more motion output components 200-26C may include: motion controllers, actuators, mechanical linkages, electromechanical devices, and / or components that generate physical movement. This list is not intended to be exhaustive, and one or more motion output components 200-26C may include other motion output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output motion output. Figure 5 As illustrated, output components 200-26 include one or more haptic output components 200-26D. One or more haptic output components 200-26D may include any component for outputting (e.g., generating, creating, and / or displaying) and / or causing haptic output (e.g., output using haptically perceptible means, such as vibration, pressure, texture, and / or shape). Examples of one or more haptic output components 200-26D may include: speakers, components that generate vibrations, components that generate texture changes, components that generate pressure changes, and / or components that create perceptible haptic effects. This list is not intended to be exhaustive, and one or more haptic output components 200-26D may include other haptic output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output haptic output.

[0136] like Figure 5As illustrated, output components 200-26 include one or more communication components 200-26E. The one or more communication components 200-26E may include any components (e.g., antennas, modems, network interface components, encoders, decoders, and / or communication protocol stacks) for transmitting and / or receiving communications internal and / or external to the agent system 200-20. In some embodiments, communication may be between different devices and / or between components within the same device. In some embodiments, communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, the one or more communication components 200-26E include one or more features of one or more communication components 200-22B (e.g., as described above). In some embodiments, the one or more communication components 200-26E are identical to one or more communication components 200-22B (e.g., handling communication inputs and outputs and therefore considered as one or more of input and output components and / or both).

[0137] Throughout this disclosure, reference may be made to moving output (e.g., referred to in various forms such as: movement, device movement, moving output, device motion, motion output, and / or motion output). In some embodiments, output movement (e.g., an output that causes movement) refers to movement of an electronic device (e.g., a portion or component thereof relative to another portion and / or the entire electronic device). For example, refer again to Figure 2B The movable output can refer to the device 200 actuating the movable component 200-3 to move the display part 200-1 to... Figure 2B The illustrated location (e.g., from) Figure 2A (Positioning within). In some embodiments, the motion output is not (e.g., excluding and / or not only including) tactile output (e.g., tactile motion output). In some embodiments, the motion output is not (e.g., excluding and / or not only including) vibration output. In some embodiments, the motion output is not (e.g., excluding and / or not only including) oscillatory motion (e.g., movement of an actuator that causes vibration solely by repeatedly moving a component along a path within the device). In some embodiments, the motion output includes (e.g., requiring and / or causing) a change in the position and / or pose of at least a portion (and / or all) of a component or electronic device. In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device from a first position and / or a first pose to a second position and / or a second pose. For example, relative to Figures 2A to 2C ,exist Figure 2A , Figure 2B and Figure 2CIn each of these embodiments, display portion 200-1 is shown in a different position (e.g., in space) and pose (e.g., relative to base portion 200-2). In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device to a third position and / or a third pose (e.g., from a first position and / or a first pose and / or from a second position and / or a second pose). In some embodiments, the third position and / or the third pose is the same as the first position and / or the first pose and / or the second position and / or the second pose. For example, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement is made to return to... Figure 2A The illustrated first positioning. For example, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement continues until it stops at... Figure 2C The illustrated third position.

[0138] Throughout this disclosure, electronic devices can be exemplified (and / or described) as being in different positions and / or poses at different times. For example, Figure 2A Example of device 200 in the first position, Figure 2B An example is shown of device 200 in the second position, and Figure 2A A device 200 in a third position is illustrated. In some embodiments, the electronic device moves itself between such positions and / or poses (e.g., using a movement output). For example, device 200 moves from a first position to a second position under its own power (e.g., using a power supply and one or more actuators to induce movement). Specifically, any examples of electronic devices illustrated and / or described herein in different positions and / or poses (e.g., at different times) should be understood to cover scenarios where the device moves itself between such positions and / or poses (e.g., unless otherwise explicitly stated).

[0139] Throughout this disclosure, reference may be made to “performing output,” “causing output,” and / or “output” (e.g., via one or more output generating devices and / or via one or more output generating components) (and / or similar phrases). In some embodiments, the output (e.g., or variations thereof) includes (and / or) output movement (e.g., moving the output as described above).

[0140] Throughout this disclosure, references may be made to “display,” “cause display,” and / or “output visual content” (e.g., via one or more display components) (and / or similar phrases). In some embodiments, display (e.g., or variations thereof) includes displaying visual content in conjunction with output movement (e.g., moving output as described above).

[0141] Throughout this disclosure, reference may be made to "output audio," "output that causes audio," and / or "provide audio output" (e.g., via one or more audio generation components and / or via one or more audio output devices) (and / or similar phrases). In some embodiments, outputting audio (e.g., or variations thereof) includes outputting audio content in conjunction with output movement (e.g., movement output as described above).

[0142] Throughout this disclosure, reference may be made to the movement (and / or similar phrases) of an avatar (e.g., or other representations of a displayed user, agent, and / or role) (e.g., via one or more display components). In some embodiments, moving an avatar (e.g., or a variant thereof) includes movement in conjunction with output movement (e.g., movement output as described above) to display visual content. For example, displaying an avatar nodding in agreement may include an electronic device moving in a manner similar to avatar movement (e.g., simulating a nod). In some embodiments, moving an avatar (e.g., or a variant thereof) includes movement with output movement (e.g., movement output as described above) without displaying visual content. For example, a device may perform a simulated nod without moving the displayed avatar's movement output (e.g., the avatar does not move relative to the display). Figure 5As illustrated, agent system 200-20 may optionally interface with external components such as external database 200-30, remote processing component 200-32, and / or remote management component 200-34. In some embodiments, external database 200-30 represents one or more functions that provide data storage resources accessible to agent system 200-20. In some embodiments, access to data in external database 200-30 is provided directly to agent system 200-20 (e.g., the agent system manages the database) and / or indirectly to agent system 200-20 (e.g., the database is managed by a different system, but the data stored therein can be provided and / or stored for use by agent system 200-20). In some embodiments, external database 200-30 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., is a database of web services accessible to different agent systems), and / or a combination of dedicated and non-dedicated database resources. In some embodiments, remote processing component 200-32 represents one or more components that serve as data processing resources accessible to agent system 200-20. In some embodiments, access to remote processing component 200-32 is provided directly to agent system 200-20 (e.g., the agent system manages processing resources) and / or indirectly to agent system 200-20 (e.g., processing resources managed by a different system, but which can provide data processing for the benefit of agent system 200-20). In some embodiments, remote processing component 200-32 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., a processing resource for a web service accessible to a different agent system), and / or a combination of both dedicated and non-dedicated processing resources. Examples of data processing include processing image data (e.g., for feature extraction and / or object detection), processing audio data (e.g., for processing natural language speech input via a large language model), and / or training machine learning algorithms and / or models. In some implementations, remote management component 200-34 represents functions including management functions and / or functions related to management functions. For example, such management functions may include providing component updates (e.g., software and / or firmware updates) to agent system 200-20, managing accounts (e.g., associated licenses, access controls, and / or preferences), synchronizing between different agent systems and / or their components (e.g., enabling agents accessible via multiple devices of a user to provide a consistent user experience across such devices), managing collaboration with other services and / or agent systems, error reporting, managing backup resources to maintain agent system reliability and / or agent availability and / or other functions required for agent system 200-20 to perform operations, such as those described herein.

[0143] The above text is about Figure 5 The various components of the described agent system 200-20 represent functional blocks that represent functionality. This functionality can be implemented on the same and / or different hardware (e.g., physical components) and / or by the same and / or different software. For example, a functional block can be implemented using one or more physical components, devices (e.g., computer system 100 and / or electronic device 200), and / or software programs. In other words, each functional block does not necessarily represent a single, discrete physical component, device, and / or software program, but can be implemented using one or more of these. Furthermore, agent system 200-20 may include multiple implementations of the functionality represented by the respective functional blocks. For example, agent system 200-20 may include multiple different model components representing ML models used in different contexts, multiple different API components representing different APIs for different services, and / or multiple different visual output components for outputting different types of visual output.

[0144] Now let’s turn our attention to a discussion of the concepts that may arise regarding the operation of agents.

[0145] As discussed throughout, the agent may be able to interact with the user. In some implementations, this capability includes the ability to process explicit requests, commands, and / or statements. In some implementations, explicit requests, commands, and / or statements include and / or are interpreted as instructions relating to completing a task (e.g., displaying X, completing task Y, and / or performing operation Z). In some implementations, the agent includes the ability to process implicit requests, commands, and / or statements. In some implementations, implicit requests, commands, and / or statements do not include explicit requests, commands, and / or statements. For example, “I like to go to Europe” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays the itinerary in response to the statement. As another example, “This picture is for my grandmother” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays a suggestion to modify the picture. As another example, “I am tired” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 causes a sleep meditation application to start a meditation session. As another example, "I miss my grandfather" can be interpreted as an implicit request, command, and / or statement, which, upon detection, device 200 can initiate a real-time communication session with the grandfather (e.g., a phone call, video call, and / or text messaging session). In some implementations, implicit requests are more likely to be processed based on one or more current contexts, operational contexts, and / or user contexts, while explicit requests are less likely to be processed based on one or more current contexts, operational contexts, and / or user contexts. For example, the phrase "Call my grandfather" can be an explicit request, and in response to detecting such a request, device 200 will initiate a real-time communication session with the grandfather regardless of one or more current contexts, operational contexts, and / or user contexts. However, the phrase "I miss my grandfather" can be an implicit request, and in response to detecting such a request, device 200 can display a list of gifts to buy for the grandfather if the user has recently been discussing buying gifts, or can call the grandfather in a different context that does not include the user's recent discussions about buying gifts. In some implementations, a request can include one or more explicit requests and one or more implicit requests. In some implementations, implicit requests are responded to independently of explicit requests; in other implementations, responses to implicit requests depend on explicit requests.

[0146] This document may refer to responses from an agent output by a device. In some embodiments, the response includes an audio component (e.g., audio output, audible output, sound and / or speech) (also referred to herein as a “verbal response,” “audio response,” and / or “audible response”) and / or a visual component (e.g., a display and / or movement of a representation and / or avatar). In some embodiments, the response includes a motion component (e.g., movement of the device). In some embodiments, the response includes a tactile component (e.g., touch and / or vibration).

[0147] This document may refer to internal dialogues, internal contexts, and / or operational contexts, which may refer to the dynamic context or dynamic decision-making process of a device, the internal state of device 200, and / or internal data of the device based in part on its decisions. In some embodiments, internal dialogues include a set of one or more rules, features, detections, and / or observations used by a computer system to generate responses to one or more commands, questions, and / or statements. In some embodiments, the set of one or more rules, features, detections, and / or observations is learned and / or generated via deep learning and / or one or more machine learning algorithms and / or using one or more machine learning and / or system agents. In some embodiments, internal dialogues are generated in real time. In some embodiments, internal dialogues are stored locally and / or via cloud storage. In some embodiments, internal dialogues can be modified, updated, and / or deleted. In some embodiments, internal dialogues are generated based on other internal dialogues.

[0148] This document may refer to (e.g., agent, user, and / or role) personality and / or behavior (or a representation of personality / behavior). In some embodiments, personality and / or behavior refers to one or more characteristics that a device detects, understands, conforms to, applies, and / or tracks. In some embodiments, personality or behavior is used as the basis for performing operations. For example, an agent may detect a user's personality and respond in a personality-based manner (e.g., outputting different responses in response to different user personalities). As another example, an agent may output responses having characteristics corresponding to one or more characteristics corresponding to personality and / or behavior (e.g., outputting responses in different ways depending on the agent's personality). In some embodiments, such characteristics represent and / or simulate a user's personality, such as how a user acts and / or speaks. In some embodiments, such characteristics approximate a user's personality.

[0149] In some implementations, the agent is a system agent. In some implementations, the system agent is an agent corresponding to the operating system originating from the device (e.g., the device implementing the agent) and / or a process controlled by the device's operating system. In some implementations, the agent is an application agent. In some implementations, the application agent is an agent corresponding to an application originating from the device (e.g., the device implementing the agent) (e.g., installed on the device and / or executed by the device) and / or a process controlled by the device's application.

[0150] This document may refer to the representation (e.g., avatar and / or avatar representation) of an agent (e.g., and / or user (e.g., person, object and / or animal) and / or user interface object (e.g., animated character)). In some embodiments, the representation of an agent refers to a set of output characteristics (e.g., visual and / or audio) of the agent (and / or user and / or user interface object). For example, the representation of an agent may include (and / or correspond to) a set of one or more visual characteristics (e.g., facial features of an animated face) and / or one or more audio characteristics (e.g., language and speech characteristics of audio output). In some embodiments, the representation (e.g., of the agent) is used to represent the agent's output. For example, a device implementing an interactive agent outputs audio in the agent's voice and displays an animated face of the agent moving in a manner that simulates the agent speaking the audio output. In this way, the user can feel that they are having a normal conversation with the agent. In some embodiments, the representation of an agent includes (or does not include) personality and / or behavioral characteristics (e.g., as described above). For example, the representation of an agent may include (and / or correspond to) a set of visual characteristics (e.g., facial features of an animated face) and a set of personality characteristics. In some implementations, the representation of the agent includes a set of user characteristics corresponding to the user's visual representation (e.g., representations of the user's appearance, voice, and / or personality used as avatars that appear to move and / or speak). In some implementations, the representation is a facial representation (e.g., a user interface object outputting features that simulate a human face and / or facial expressions (e.g., for conveying information to a viewer)).

[0151] In some implementations, a role (e.g., a role representing an agent and / or an avatar) refers to a specific set of characteristics. For example, an avatar may embody the characteristics (e.g., use, application, interaction with, and / or output) of a fictional and / or non-fictional character (e.g., from a movie, show, book, TV series, and / or popular culture).

[0152] In some implementations, (e.g., agent and / or avatar) voice refers to one or more characteristics corresponding to a set of sound outputs that are similar to (e.g., represent, imitate, and / or reproduce) spoken utterances (e.g., attributable to and / or simulated as output by an agent and / or avatar). For example, device 200 may output sentences that sound different depending on the voice used. In some implementations, a particular character and / or avatar may be configured to use a particular voice (e.g., have a corresponding voice). In some implementations, the particular voice may mimic the user's voice.

[0153] In some implementations, (e.g., the appearance of an agent and / or avatar) refers to a set of one or more characteristics corresponding to the visual output representing the avatar (and / or agent). For example, device 200 may output an avatar having a set of facial features that form an appearance similar to a specific character from a movie.

[0154] In some implementations, an avatar's expression refers to one or more characteristics corresponding to a specific visual appearance of a user, avatar, and / or agent. For example, device 200 may output an avatar having a set of facial features arranged in a specific manner to give the appearance of a facial expression (e.g., which can be used as a form of nonverbal communication to a user) (e.g., a frown is an expression of sadness, a smile is an expression of happiness, and / or wide eyes are an expression of surprise). As another example, device 200 may output an avatar having a set of body features (e.g., arms and / or legs) arranged in a specific manner to give the appearance of a body expression (e.g., which can be used as a form of nonverbal communication to a user) (e.g., a gesture is an expression of approval, covering the eyes is an expression of fear, and / or shrugging is an expression of lack of awareness). In some implementations, expressions include avatar movement (e.g., a nod is an expression of agreement and / or disagreement). In some implementations, device 200 may be movable via a movement component to indicate expressions with or without avatar movement. In some implementations, the agent performs one or more actions that depend on the user's facial expressions (e.g., detecting whether a person is sad and responding with a kind statement or question). In some implementations, facial expressions (e.g., whether and / or how they are used and / or how they are output) depend on personality. For example, a first-person sex may use more specific facial expressions than a second-person sex. As another example, a first-person sex's expressions (e.g., frowning, smiling, and / or widening their eyes) may look different from a second-person sex's expressions (and / or similar and / or equivalent expressions) (e.g., a first-person sex smiles with their teeth showing, but a second-person sex smiles without showing their teeth).

[0155] In some implementations, an agent (e.g., an avatar of the agent and / or an agent system implementing the agent (e.g., hardware and / or software)) mimics the characteristics of another user, agent, and / or role (e.g., in terms of personality, behavior, facial expressions, and / or voice). In some implementations, mimicry includes mirroring the user (e.g., replicating phrases and / or movements detected from a user interacting with the agent). In some implementations, simulating user characteristics includes attempting to reproduce the user's characteristics (e.g., in exactly the same way and / or in a way that is similar to the characteristics but not an exact reproduction of the characteristics). For example, an agent mimicking voice and / or facial expressions does not require the agent to have exactly the same voice and / or facial expressions as the user being mimicked (e.g., simply to resemble the user's voice and / or facial expressions).

[0156] In some implementations, components and / or devices use (e.g., performing actions, making decisions, and / or determining context based on them) learned characteristics (e.g., characteristics of the context, user, and / or environment learned by the device over time (e.g., via detection, prior experience, and / or feedback (e.g., from one or more users))). For example, characteristics learned over time may include user routines. In such an example, if a particular user requests a summary of any new messages for that user from an agent at the same time every day, the agent may learn to automate actions based on the characteristics of the learned routines (e.g., what data is needed, when data is needed, and / or for which user). In some implementations, the learned characteristics enable the agent (and / or device) to improve its understanding (and / or response to) of the context, user, and / or environment, and / or its understanding of context, user, and / or environment that is otherwise not (and / or will not) understood (e.g., not responded to or responded to incorrectly). In some implementations, reinforcement learning (e.g., by and / or for the agent) is used to form the learned characteristics. In some implementations, the learned features correspond to one or more confidence levels, determinisms, and / or rewards (e.g., shaped by one or more reward functions). In some implementations, the learned features (and / or how they are used to influence the output of the agent and / or device) can change over time (e.g., confidence levels, determinisms, and / or rewards change over time). For example, the output of a device before learning a set of learned features may differ from the output of a device after learning a set of learned features. In some implementations, components and / or devices use the learned knowledge. For example, similar to what is described above regarding learned features, learned knowledge may refer to information used to update (e.g., enhance, add to, and / or expand) the device's knowledge base (e.g., for use by agents implemented thereon). In some implementations, multiple sets of learned features for a user may be stored and / or used. In some implementations, different sets of learned features for different users may be stored and / or used.

[0157] This document may refer to interactions with an agent (and / or device). In some embodiments, an interaction refers to a set of one or more inputs and / or outputs from a device implementing the agent and one or more users. In some embodiments, an interaction may be a user input (e.g., “Please turn on the light”) and a corresponding output (e.g., turning on the light and / or the device’s response “OK”). In some embodiments, an interaction may include multiple inputs / outputs performed by one or more parties to the interaction (e.g., the device and / or the user). In some embodiments, an interaction may include a first user input (e.g., “Please turn on the light”) and a corresponding first output (e.g., “Which lights?”), and also includes a second user input (e.g., “Kitchen light”) and a second output from the device (e.g., “OK”). In some embodiments, which inputs and / or outputs are considered together as an interaction is based on logical and / or contextual grouping (e.g., interactions within the previous thirty (30) seconds and / or interactions related to turning on the light). As those skilled in the art will understand, interactions may be considered in an implementation-dependent manner (e.g., determining when an interaction is complete may involve determining whether the user is still present (e.g., is still talking) and / or whether the user is still talking about the light or has moved on to a different topic). In some implementations, the interaction is the current interaction (e.g., ongoing, currently occurring, and / or active). In some implementations, the interaction is a previous interaction. The examples above describe a device that engages in dialogue with a user. In some implementations, the dialogue is between two or more users (e.g., users in an environment). For example, the device may detect dialogue between users (e.g., users directing speech and responses to each other, rather than to the device).

[0158] In some implementations, the agent (and / or device) determines and / or performs an action based on an intent corresponding to the user. In some implementations, the device detects user input and outputs a response that depends on the intent of the user input. For example, the device detects user input including a pointing gesture detected along with a verbal command to “turn on the light,” and in response, the device turns on the light determined to correspond to the intent of the input (e.g., the light the pointing gesture is pointing to). In some implementations, intent is determined using one or more of the following (e.g., determined by the device detecting the input and / or by one or more other devices): one or more inputs, knowledge (e.g., knowledge about the user learned based on observed behavior, personality, and history of interactions), learned characteristics, and / or context. In some implementations, intent is determined based on one or more types of input (e.g., verbal input, visual input via a camera, and / or contextual input).

[0159] Now turn our attention to the implementation of user interfaces (“UIs”) and associated processes on electronic devices (such as computer system 100 and / or electronic device 200).

[0160] Figures 6A to 6D Exemplary user interfaces for engaging in interaction are illustrated according to some implementation schemes. The user interfaces in these figures are used to illustrate the processes described below, including... Figure 7 and Figure 8 The process in.

[0161] Specifically, at least two distinct features of the exemplary computer system will be described below. The first feature relates to the physical movement of the computer system in connection with the detection of different types of interactions occurring in the environment. The second feature relates to displaying one or more word clouds in response to the detection of one or more inputs. For ease of discussion, [the following will be discussed regarding...] Figures 6A to 6D Discuss the first feature, and then... Figures 6A to 6D The second feature is discussed. However, in implementation, the computer system may execute one or more techniques related to these two features concurrently and / or independently.

[0162] Regarding the first characteristic, Figures 6A to 6DExamples of one or more scenarios of computer system movement based on the type of interaction are illustrated. In some embodiments, one type of interaction may be a specific user talking to the computer system. In some embodiments, one type of interaction may be whether one or more users are having a conversation with the computer system, and another type of interaction may be whether one or more users are interacting with each other (and in some embodiments, not interacting with the computer system). In some embodiments, another type of interaction is an interaction between a user and themselves, such as a person talking aloud to themselves. In the scenario where a user is talking to themselves, the computer system may temporarily face the user who is speaking and, in response to determining that the user is not addressing the computer system, return to its original location. In some embodiments, one type of interaction is an interaction between a user and another device and / or computer system (e.g., a person talking to another device). In some embodiments, one type of interaction may be a private interaction, while another type of interaction may be a non-private interaction, such as a conversation between two or more individuals in a group and a conversation between all and / or most people in a group. In some embodiments, the interaction is a conversation involving one or more users talking. In some embodiments, the interaction may differ from a conversation, such as one or more users looking at each other, gesturing towards each other, touching each other, and / or talking to each other. In some implementations, the computer system detects interaction when one or more users are silent and / or when no users are talking in the environment. In some implementations, the computer system may move differently in response to detecting different interactions. For example, the computer system may move to face the user, not face the user, face a specific area of ​​the environment, or face another type of user in the environment.

[0163] Go to Figure 6A , Figure 6A The right side illustrates environment 602. Within environment 602 are computer system 600, user 608 located at the leftmost position, and user 604 located at the rightmost position. The dashed lines at an angle to computer system 600 represent the visibility area of ​​computer system 600. In some embodiments, the display of computer system 600 is visible to elements within the dashed lines of environment 602. In some embodiments, the dashed lines at an angle to computer system 600 represent the detection field of the computer system, such as the field of view of one or more cameras of computer system 600. Figure 6A At this point, computer system 600 faces user 604 because computer system 600 has detected that user 604 is interacting with the computer system. It is worth noting that... Figure 6A At this point, computer system 600 is not oriented towards user 608, which is represented by user 608 being outside the dashed lines representing the visibility area and / or detection field of computer system 600.

[0164] In some implementations, the computer system 600 moves in response to detecting different types of interactions. For example, in Figure 6B At this point, computer system 600 detects input 618 from user 608, which is related to... Figure 6A This is a different interaction detected from the interaction detected by user 604 (e.g., input 606). Here, computer system 600 detects that a different interaction has occurred because a new user (e.g., user 608) has begun interacting with computer system 600. Figures 6A to 6B In response to the detection of different types of interaction (e.g., interaction between user 608 and computer system 600), computer system 600 rotates counterclockwise, so that user 608 is within the detection field and / or visibility area of ​​computer system 600. This is determined by user 608 in... Figure 6B The dotted line in the text and user 604 are in Figure 6B Examples are shown outside the dotted lines.

[0165] In some implementations, the computer system 600 moves in different ways in response to the detection of an interaction. For example, instead of rotating counterclockwise or in addition to counterclockwise rotation, the computer system 600 may tilt (e.g., 0 to 270 degrees) to turn from user-facing 604 to user-facing 608 and / or move left, up, down, and / or in any combination thereof to move from user-facing 604 to user-facing 608. As used herein, it should be understood that "facing one or more users" is used to indicate that one or more users are within the visibility area and / or detection field of the computer system 600.

[0166] In some implementations, the computer system 600 interacts through different modal detection. For example, in Figure 6BIn some embodiments, computer system 600 detects different interactions based on received voice input (e.g., input 618). However, in some implementations, computer system 600 may also detect different interactions based on detecting one or more air gestures (such as a user pointing at another user and / or the computer system and / or waving to another user and / or the computer system). In some implementations, the use of different air gestures allows computer system 600 to detect different types of interactions. In some implementations, computer system 600 detecting that a user is waving to another user may cause computer system 600 to move differently than computer system 600 detecting that a user is high-fiving another user. In some implementations, computer system 600 may also detect different interactions based on detecting one or more other types of input (such as one or more gaze inputs (e.g., whether one or more users are looking at each other and / or the computer system), sound inputs (e.g., whether one or more users are making noise relative to another user, such as bumping a physical object and / or opening a physical object), and / or inputs pointing to one or more hardware components (such as buttons and / or rotatable input mechanisms)).

[0167] In some implementations, computer system 600 moves to face the user based on detecting an interaction that does not involve direct communication between the user and computer system 600. In some implementations, computer system 600 responds to detecting a mention by user 604. Figure 6B The user 604 is moved to face user 608. In some implementations, if user 604 is... Figure 6A or Figure 6B If a user says, "This is my friend John," then computer system 600 can redirect to user 608 (e.g., assuming user 608 is "John"), without user 608 needing to interact directly with computer system 600 first. In some implementations, detecting that user 604 is mentioning user 608 includes detecting that user 604 is performing an air gesture, such as pointing at user 608, waving to user 608, and / or gesturing for user 608 to come over.

[0168] Although Figure 6B In some embodiments, the computer system 600 does not move when facing the user 608 (e.g., shaking and / or bowing), but in some implementations, the computer system 600 may move when facing the user and / or one or more users (e.g., if no distinct interaction is detected). In some implementations, the computer system 600... Figure 6B Shaking and / or bowing when facing user 608. In some implementations, the movements performed (such as bowing and / or shaking) instruct computer system 600 to reverse and / or interact with user 608 (e.g., while user 608 is interacting with computer system 600).

[0169] In some implementations, computer system 600 detects a new interaction when it detects that one or more users (and / or all users) are relatively silent and / or not talking in environment 602. For example, as Figure 6C As illustrated, computer system 600 expands its visibility area to face both user 604 and user 608 simultaneously because computer system 600 has determined that a new interaction has occurred since no input from user 604 or user 608 has been detected for a predetermined period of time (e.g., 1 second to 600 seconds). In some implementations, computer system 600 detects a new interaction when it detects that one or more users have stopped talking, gesturing, interacting, and / or gazing at each other and / or the computer system for a period of time (e.g., 1 second to 600 seconds).

[0170] In some implementations, the computer system 600 detects a new interaction occurring when users in the environment are interacting with each other without interacting with the computer system. Figure 6D At this point, computer system 600 detects input 622, which includes the term "we" to indicate that users 604 and 608 are talking to themselves (e.g., with...). Figure 6A and Figure 6B (The use of "I" differs in the text). Therefore, it can be seen that in Figure 6D At this point, computer system 600 detects that users 604 and 608 are talking to each other, therefore computer system 600 faces users 604 and 608. If Figure 6C If skipped (e.g., the user never remains silent), then in response to detecting that the users are interacting with each other without interacting with computer system 600, computer system 600 will remove the user from the computer system 600. Figure 6B The position facing user 608 is rotated clockwise to computer system 600. Figure 6D Positioning is targeted at multiple users. In some implementations, if... Figure 6C If not skipped, computer system 600 can access it from... Figure 6C Move downwards or upwards at the location Figure 6D A flatter positioning (e.g., 0-degree tilt). In some embodiments, computer system 600 tilts up or down (e.g., in a less flat positioning) to represent that computer system 600 monitors interactions less and / or has less interest in interactions compared to when computer system 600 is tilted to a flatter positioning. In some embodiments, in response to detecting that users 604 and 608 are talking to each other without talking to computer system 600, computer system 600 moves to a position where computer system 600 is not facing users 608 and / or user 604.

[0171] exist Figure 6D At this point, computer system 600 detects input 622 from user 604. It should be noted that although user 604 is speaking, computer system 600 does not move to face user 604, but remains simultaneously facing both user 604 and user 608. In some implementations, computer system 600 reacts to different types of interactions. In some implementations, computer system 600 moves to face user 604 in response to the initial detection of input 622, but turns away in response to detecting that the user is talking to themselves (e.g., not addressing computer system 600).

[0172] Moving on to the second feature, the discussion below includes... Figures 6A to 6D The description illustrates one or more scenarios in which a computer system 600 displays a set of words (referred to herein as a word cloud) in response to the detection of one or more inputs. In some embodiments, a word cloud is defined as a set of words grouped together under a common category, topic, and / or theme. In some embodiments, in response to the detection of specific words and context based on one or more inputs, the computer system 600 displays words under headings representing the common category, topic, and / or theme of the word cloud.

[0173] like Figure 6A As illustrated, computer system 600 detects input 606 from user 604 (e.g., “Let’s go to the beach”). In response to the context of detecting input 606, computer system 600 determines that the word “beach” should be added to word cloud 614 (e.g., a visual representation of a group titled “Azores trip”). Therefore, based on the context of input 606, computer system 600 determines that user 604 is referring to a trip to the Azores and displays word representation 616, which is the word “beach” detected by computer system 600 in the phrase of input 606 and determined to be relevant to the Azores trip word cloud context. In some embodiments, computer system 600 uses one or more machine learning algorithms to determine which words are relevant to the word cloud and / or should be added and / or used to generate the word cloud or another group. In some embodiments, words may be relevant based on one or more contexts involving the current interaction between computer system 600 and the user, previous interactions between computer system 600 and the user, and / or an explicit request from the user to add words to the word cloud. Note that in Figure 6A At this point, user 604 did not explicitly request the word "beach" to be added to word cloud 614. However, computer system 600 intuitively determined that the user made an implicit request to add the word "beach" from the phrase included in input 606 to word cloud 614.

[0174] In some implementations, the computer system 600 adds additional words to one or more word clouds based on context. For example, as Figure 6B As illustrated, computer system 600 detects input 618 from user 608 (e.g., "I like hiking"). Based on input 618, computer system 600 determines that "hiking" is a keyword (and / or a word that should be added to the word cloud based on context). In some implementations, keywords are words from input 618 that are relevant to the context determined by computer system 600. For example, in Figures 6A to 6B In the example, users 604 and 608 are discussing a trip to the Azores. Because the word "hiking" from input 618 is determined to be relevant to the Azores trip, computer system 600 adds "hiking" to word cloud 614, which is displayed as a word representation 620 on user interface 612. Alternatively, if user 604 mentions an item from a grocery list, computer system 600 does not add the item from the grocery list to the list discussing the Azores trip (e.g., as described below regarding...). Figure 6C (As discussed).

[0175] In some implementations, the computer system 600 does not include words that have not been identified as keywords as part of the word cloud. For example, such as Figure 6B As illustrated, upon detecting input 618, computer system 600 includes the keyword "hiking" in the list of words, but excludes the word "likes" because the word "likes" is irrelevant to the topic indicated by word cloud 614. In some implementations, if computer system 600 detects input including the phrase "I like hiking in the mountains," computer system 600 may add the keywords "hiking" and "mountain" to word cloud 614, but will not include the word "on," because it is irrelevant to the topic indicated by word cloud 614 and / or is not a keyword relative to that topic.

[0176] In some implementations, the computer system 600 dynamically displays word representations. In some implementations, if the computer system 600 detects that the user 608 slowly speaks input 618, the computer system 600 slowly displays the word "hiking" and / or adds the word "hiking" to the word cloud at a slower rate compared to if the user 608 speaks the input quickly. In some implementations, the computer system 600 displays words as part of the word cloud at different sizes. In some implementations, if the computer system 600 detects that the user 608 speaks low-pitched, low-voice, and / or low-volume input 618, the computer system 600 displays the word "hiking" at a smaller size compared to if the computer system 600 detects that the user 608 speaks high-pitched, high-voice, and / or high-volume input 618. In some implementations, if the computer system 600 detects that the word "hiking" is more relevant to the theme of the word cloud 614 than the word "beach," the computer system 600 displays "hiking" higher than the word "beach" on the user interface 612. In some implementations, if the computer system 600 detects that the word "hiking" is less relevant to the topic of the word cloud 614 than the word "beach," then the computer system 600 displays "hiking" lower than the word "beach" on the user interface 612. Figure 6C At this point, computer system 600 did not detect any voice input and therefore did not add any word representation to word cloud 614.

[0177] In some implementations, computer system 600 generates new word clouds and / or adds words to an existing word cloud that is different from the currently displayed word cloud (e.g., word cloud 614) and / or the word cloud to which computer system 600 recently added words. For example, as Figure 6D As illustrated, in response to detecting input 622 from user 604, computer system 600 creates a new word cloud (e.g., word cloud 626) and adds a relative word (e.g., "egg") to word cloud 626. Computer system 600 adds the word "egg" to word cloud 626 when it detects that the keyword "egg" should be related to different categories and / or groups of words (e.g., not just the word "egg" is related to the topic indicated by word cloud 614). Figure 6DAs illustrated, in response to the determination of input 622 mentioning a grocery list, computer system 600 creates word cloud 626 (e.g., titled "groceries"), which is displayed on user interface 624. In some embodiments, computer system 600 detects a second input related to word cloud 626 and adds words from the other input (e.g., "milk," "bread," and / or "cheese"). In some embodiments, while displaying word cloud 626 (e.g., a list and / or group of words displayed via a display component), computer system 600 detects input related to a travel word cloud for the Azores, such as "we should book a hotel." In some embodiments, computer system 600 redisplays word cloud 614 and / or increases the size of word cloud 614 to add relative words (e.g., "hotel") to the group of words. In some implementations, if computer system 600 detects input such as “we should book a hotel and buy milk today”, computer system 600 concurrently adds “hotel” to word cloud 614 and “milk” to word cloud 626, because the input includes “hotel” (which is a word related to word cloud 614) and “milk” (which is a word related to word cloud 626).

[0178] In some implementations, computer system 600 does not emphasize one or more word clouds (e.g., one or more word clouds that are not currently being focused on). For example, as Figure 6D As illustrated, in response to displaying word cloud 626, computer system 600 collapses word cloud 614 to the lower right corner of user interface 624. In some embodiments, in response to input that adds a word to a different word cloud (e.g., "egg" is added to word cloud 626), computer system 600 completely stops displaying word cloud 614. In some embodiments, if computer system 600 detects that user 604 begins discussing shopping at a shopping mall, computer system 600 stops displaying word cloud 626 and begins displaying word clouds related to shopping at a shopping mall. In some embodiments, if computer system 600 detects that user 604 begins discussing shopping at a shopping mall, computer system 600 minimizes word cloud 626 while displaying word clouds related to shopping at a shopping mall, but still displays the word cloud.

[0179] In some embodiments, computer system 600 displays word clouds differently. In some embodiments, computer system 600 displays word representations and / or groups of word representations in word cloud 614 as square formations, while displaying word representations and / or groups of word representations in word cloud 626 as circular formations. In some embodiments, computer system 600 displays word representations and / or groups of word representations in word cloud 614 as circular formations, while displaying word representations and / or groups of word representations in word cloud 626 as square formations. In some embodiments, the position of a word cloud within the word cloud indicates its relevance to the word cloud. In some embodiments, more relevant words are displayed at the top of the word cloud, and less relevant words are displayed at the bottom, and / or vice versa. In some embodiments, computer system 600 moves differently when displaying different word clouds. In some embodiments, computer system 600 moves relative to the shape of the word cloud and / or based on the characteristics of the words in the word cloud (e.g., moving to indicate mountains and / or hills) and / or the number of words.

[0180] In some embodiments, computer system 600 includes different images in the word cloud. In some embodiments, computer system 600 includes images of beaches and / or hiking trails with word cloud 614, but these images are not included with word cloud 626. In some embodiments, computer system 600 includes images of grocery stores and / or eggs with word cloud 626, but these images are not included with word cloud 614.

[0181] Figure 7 This is a flowchart illustrating a method for mobile positioning using a computer system according to some embodiments. Process 700 is executed at a computer system (e.g., 100, 200, and 600). Some operations in process 700 may be combined, some operations may be changed in order, and some operations may be omitted.

[0182] As described below, process 700 provides an intuitive approach to mobile positioning. This method reduces the cognitive burden on users, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to position themselves faster and more efficiently, saving power and increasing the time interval between battery charging.

[0183] In some embodiments, process 700 is performed at a computer system (e.g., 600) that communicates with a movable component (e.g., an actuator, a movable base, a rotatable component, and / or a rotatable base). In some embodiments, the computer system is a watch, telephone, tablet, fitness tracker, processor, head-mounted display (HMD) device, public utility, media device, speaker, television, and / or personal computing device. In some embodiments, the computer system communicates with one or more output devices (e.g., a display component, an audio generation component, a speaker, a haptic output device, a display screen, a projector, and / or a touch-sensitive display). In some embodiments, the computer system communicates with a movable component (e.g., an actuator (e.g., a pneumatic actuator, a hydraulic actuator, and / or an electric actuator), a movable base, a rotatable component, and / or a rotatable base).

[0184] When a computer system (e.g., 600) is in a first location within an environment (e.g., 602) (e.g., a physical environment and / or a virtual environment) via a moving component, the computer system detects (702) the occurrence of a first interaction (e.g., 606, 618, and / or 622) (e.g., a single instance and / or a dissimilar event) (e.g., such as...). Figures 6A to 6D Location description (e.g., one or more users (e.g., animals, users, people and / or objects) looking at, speaking, gesturing and / or moving in one or more directions and / or about each other) (and in some embodiments, when the computer system is facing a second location in the environment).

[0185] In response to (704) detecting a first interaction (e.g., 606, 618, and / or 622), based on the determination that the first interaction (e.g., 606, 618, and / or 622) is a first type of interaction (e.g., a dialogue type, such as back-and-forth dialogue between multiple people or a dialogue with a single person and / or a dialogue with a person physically moving relative to the computer system, and / or an activity type, such as two people playing a board game and / or watching a movie), the computer system moves (706) (e.g., changes and / or repositions) via a moving component to a second location in an environment different from the first location in the environment (e.g., 602) (e.g., as...). Figures 6A to 6D (Location description) (and in some embodiments, the computer system is moved to face (e.g., the orientation of the moving components and / or the orientation of the display components that communicate with the computer system) in a corresponding orientation).

[0186] In response to (704) detecting a first interaction, and based on determining that the first interaction (e.g., 606, 618, and / or 622) is a second type of interaction different from the first type of interaction (e.g., a person talking to themselves, and / or interaction with a person's digital representation), the computer system abandons (708) moving via the moving component to a second location (e.g., as...). Figures 6A to 6D (e.g., continuing to face a certain direction and / or moving to face a direction different from the corresponding direction). In response to detecting a first interaction, moving to a second location in the environment based on determining that the first interaction is a first type of interaction, and not moving to the second location based on determining that the first interaction is a second type of interaction, enables the computer system to change its location for certain types of interactions but not for others, thereby performing an operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0187] In some embodiments, prior to detecting the first interaction (e.g., 606, 618, and / or 622), a first portion of the computer system (e.g., 600) (e.g., display components, screen, the center of the screen, a corner of the screen, and / or hardware components (e.g., those fixedly positioned on the casing of the computer system and / or currently in a specific position on the casing of the computer system) (e.g., buttons, microphone indicators, status indicators, and / or lights)) faces a first direction. In some embodiments, moving to a second position causes the first portion to face a second direction different from the first direction (e.g., as shown in the image). Figures 6A to 6D (Location description) (and / or make the first part not face and / or stop facing the first direction). Moving from facing the first direction to facing the second direction or not moving from facing the first direction to facing the second direction when the specified conditions are met enables the computer system to change direction during a type of interaction, thereby performing an operation without additional user input when a set of conditions have been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0188] In some implementations, a first interaction (e.g., 604 and / or 608) (e.g., people, animals, users, and / or objects) is classified as a first type of interaction when the number of users (e.g., 604 and / or 608) involved in a first interaction (e.g., contributing, talking, listening, and / or participating) is determined to be above a threshold (e.g., two or more individuals interacting, two or more people participating in an interaction and / or conversation, and / or two or more people participating in an activity together) (e.g., 2 to 100). In some implementations, a first interaction is classified as a second type of interaction when the number of people involved in a first interaction is determined to be below a threshold. Moving or not moving to a second location in the environment based on whether the number of people involved in the first interaction is above / below a threshold enables the computer system to change the location according to a specific number of users and control the number of users facing that part of the computer system, thereby performing operations without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform operations, improving security, and providing improved visual feedback to the user.

[0189] In some implementations, at least a portion (e.g., a segment, fragment, section, and / or time period) of the first interaction (e.g., 606, 618, and / or 622) is directed toward the computer system (e.g., 600) (e.g., in the direction of and / or associated with the computer system) (e.g., one or more users are looking at the computer system, talking to the computer system, gesturing toward the computer system, and / or moving toward the computer system). In some implementations, the first interaction is a dialogue and / or interaction between the user and the computer system (and in some implementations, a back-and-forth dialogue and / or interaction), wherein the user gives commands, questions, and / or statements to the computer system, and the computer system responds. Moving to a second location in the environment based on the interaction directed toward the computer system enables the computer system to change its location, allowing one or more users to further interact with the computer system and / or to better interact with the computer system, thereby performing actions without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform actions, and providing the user with improved visual feedback.

[0190] In some implementations, at least a portion of the first interaction (e.g., 606, 618, and / or 622) does not refer to the computer system (e.g., 600). In some implementations, no portion of the first interaction refers to the computer system. In some implementations, the first interaction is between two or more users, where the two or more users are conversing and / or interacting with each other without interacting with the computer system. In some of these implementations, the computer system records the interaction between the two or more users, and in some implementations, displays content based on the interaction between the two or more users. However, in some implementations, the computer system does not actively respond with audio output to the context of the context and / or interaction between the computer system and the two or more users.

[0191] In some implementations, in response to detecting a first interaction (e.g., 606, 618, and / or 622), based on determining that the first interaction (e.g., 606, 618, and / or 622) is a second type of interaction, the computer system moves via a moving component to a third location different from the first and second locations (e.g., such as...). Figures 6A to 6D (Location description). By determining that the first interaction is a second type of interaction, moving to a third location different from the first and second locations enables the computer system to automatically move to a certain location for different types of interactions, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0192] In some implementations, prior to detecting the first interaction (e.g., 606, 618, and / or 622), a second portion of the computer system (e.g., 600) (e.g., display components, screen, the center of the screen, a corner of the screen, and / or hardware components (e.g., those fixedly positioned on the computer system's casing and / or currently in a specific position on the computer system's casing) (e.g., buttons, microphone indicators, status indicators, and / or lights)) faces a third-party orientation. In some implementations, in response to detecting the first interaction (e.g., 606, 618, and / or 622), based on determining that the first interaction (e.g., 606, 618, and / or 622) is a second-type interaction, the computer system continues to orient the second portion of the computer system (e.g., 600) towards a third-party orientation (e.g., such as...). Figures 6A to 6D (Location description). Based on determining that the first interaction is a second type of interaction, the second part of the computer system continues to be oriented towards a third party, enabling the computer system to maintain its orientation for certain types of interactions, thereby performing operations without additional user input when a set of conditions have been met, reducing the amount of input required to perform operations, and providing improved visual feedback to the user.

[0193] In some implementations, prior to detecting the first interaction (e.g., 606, 618, and / or 622), a third portion of the computer system (e.g., 600) (e.g., a display component, screen, the center of the screen, a corner of the screen, and / or a hardware component (e.g., which is fixedly positioned on the casing of the computer system and / or currently in a specific position on the casing of the computer system) (e.g., a button, microphone indicator, status indicator, and / or light)) faces a fourth direction. In some implementations, in response to detecting the first interaction (e.g., 606, 618, and / or 622), based on determining that the first interaction (e.g., 606, 618, and / or 622) is a second type of interaction, the computer system moves via a moving component to a fourth position different from the first position, while continuing to have the third portion of the computer system (e.g., 600) facing a fourth direction (e.g., as shown in the image). Figures 6A to 6D (Location description). In some embodiments, moving to a fourth location while maintaining the computer system facing a fourth direction includes changing the computer system's location in the environment without causing the third part of the computer system to face a different direction. In some embodiments, the direction the third part of the computer system is facing includes a focal point (e.g., an object and / or point in the environment that the eighth direction points to), and moving to a fourth location while continuing to cause the third part of the computer system to face a fourth direction includes changing the location of the third part of the computer system in the environment while maintaining the focal point. Moving to a fourth location different from the first location based on determining that the first interaction is a second type of interaction, while continuing to cause the third part of the computer system to face a fourth direction, allows the computer system to face the same direction when moving for certain types of interactions (e.g., to view the interaction), thereby performing an operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0194] In some embodiments, prior to detecting the first interaction (e.g., 606, 618, and / or 622), the fourth part of the computer system (e.g., 600) (e.g., display components, screen, the center of the screen, a corner of the screen, and / or hardware components (e.g., those in a fixed position on the casing of the computer system and / or currently in a specific position on the casing of the computer system) (e.g., buttons, microphone indicators, status indicators, and / or lights)) faces the fifth direction. In some embodiments, in response to detecting the first interaction (e.g., 606, 618, and / or 622), based on determining that the first interaction (e.g., 606, 618, and / or 622) is a second type of interaction, the computer system abandons moving the computer system (e.g., 600) via the moving components (e.g., does not move, and / or abandons moving to the second position and / or additional position), while continuing to have the fourth part of the computer system facing the fifth direction. Based on the determination that the first interaction is a second type of interaction, the computer system is not moved, while the fourth part of the computer system continues to face the fifth direction, so that the computer system can face a certain direction for certain types of interactions without moving, thereby performing operations without additional user input when a set of conditions have been met, reducing the amount of input required to perform operations, and providing improved visual feedback to the user.

[0195] In some implementations, in response to the detection of a first interaction (e.g., 606, 618, and / or 622), based on the determination that the first interaction (e.g., 606, 618, and / or 622) is a third type of interaction (e.g., a dialogue type, such as back-and-forth dialogue between multiple people or a dialogue with a single person and / or a dialogue with a person physically moving relative to the computer system, and / or an activity type, such as two people playing a board game and / or watching a movie), different from the first and second types of interactions, the computer system moves (e.g., changes and / or repositions) to a second location in the environment (e.g., 602) via a mobile component. Moving to a second location in the environment based on the determination that the first interaction is a third type of interaction enables the computer system to move to a specific location for multiple types of interactions, thereby performing an operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0196] In some implementations, when the first interaction is determined to include a first type of dialogue (e.g., back-and-forth dialogue and / or dialogue directed to another person), the first interaction (e.g., 606, 618, and / or 622) is a first-type interaction. In some implementations, when the first interaction is determined to include a second type of dialogue different from the first type of dialogue (e.g., single-speaker dialogue and / or dialogue directed to a computer system), the first interaction is a second-type interaction (e.g., as...). Figures 6A to 6D (Location description). In some implementations, when it is determined that the first interaction includes a second type of dialogue, the first interaction is not a first type of interaction. In some implementations, when it is determined that the first interaction includes a first type of dialogue, the first interaction is not a second type of interaction.

[0197] In some implementations, before detecting a first interaction (e.g., 606, 618, and / or 622) and while the computer system (e.g., 600) is in a first position, a fifth part of the computer system (e.g., 600) (e.g., display components, screen, center of the screen, corner of the screen, and / or hardware components (e.g., those in a fixed position on the casing of the computer system and / or currently in a specific position on the casing of the computer system) (e.g., buttons, microphone indicators, status indicators, and / or lights)) faces a first user (e.g., 604 and / or 608) who is currently communicating (e.g., speaking, conversing, gesturing, nodding, and / or gesturing) and / or indicating. In some implementations, after moving to a second position in response to detecting a first interaction (e.g., 606, 618, and / or 622) and based on determining that the first interaction is a first type of interaction, the fifth part of the computer system (e.g., 600) faces a second user (e.g., 604 and / or 608) different from the first user when the computer system is in the second position. In some implementations, the second user is communicating while the computer system is in the second location. In some implementations, the second user is not communicating while the computer system is in the second location. Responding to the detection of a first interaction, moving to a second location to face the second user after determining that the first interaction is a first type of interaction, and not moving to the second location after determining that the first interaction is a second type of interaction, enables the computer system to face different users for certain types of interactions. This allows it to perform operations without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform operations and providing improved visual feedback to the user.

[0198] In some implementations, communication between the computer system (e.g., 600) and one or more input devices (e.g., speakers, touch-sensitive displays, and / or cameras) that detect the occurrence of a first interaction (e.g., 606, 618, and / or 622) includes receiving input via one or more input devices that positively mentions a second user (e.g., 604 and / or 608) (e.g., guiding communication with the second user, pointing to the second user, or speaking a phrase including the second user (e.g., “This is my friend, second user,” “Hello, second user,” and / or “Thank you, second user”)). In response to detecting a first interaction including receiving input mentioning the second user, moving to a second location to face the second user after determining that the first interaction is a first type of interaction, and not moving to the second location after determining that the first interaction is a second type of interaction, enables the computer system to face different users for certain types of interactions mentioning a particular user. This allows operations to be performed without additional user input when a set of conditions has already been met, reducing the amount of input required to perform operations and providing improved visual feedback to the user.

[0199] In some implementations, detecting the occurrence of a first interaction (e.g., 606, 618, and / or 622) includes receiving an indication that a third user (e.g., 604 and / or 608) (e.g., a person, animal, user, and / or object) is not communicating (e.g., no longer singing, no longer speaking, no longer gesturing, no longer conversing, no longer nodding, and / or no longer gesturing). In some implementations, receiving an indication that a third user is not communicating includes detecting that the third user is silent and / or has been silent for a predetermined period of time (e.g., 1 second to 1000 seconds). Responding to the detection of a first interaction including receiving an indication that a third user is not communicating, moving to a second location in the environment based on determining that the first interaction is a first type of interaction and not moving to the second location based on determining that the first interaction is a second type of interaction enables the computer system to change location during one type of interaction when the user stops communicating, thereby performing an operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0200] In some implementations, detecting the occurrence of a first interaction (e.g., 606, 618, and / or 622) includes detecting that a fourth user (e.g., 604 and / or 608), different from a third user (e.g., 604 and / or 608), is communicating (and / or has been communicating for more than a predetermined time period (e.g., 1 second to 1000 seconds)). In response to detecting a first interaction including receiving an indication that a third user is not communicating and a fourth user is communicating, moving to a second location in the environment based on determining that the first interaction is a first type of interaction and not moving to the second location based on determining that the first interaction is a second type of interaction enables the computer system to change location during one type of interaction when a user stops communicating and another user is communicating, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing improved visual feedback to the user.

[0201] In some embodiments, when the computer system is in a first position, the computer system (e.g., 600) is in a first tilt position (and / or angle (e.g., 0 to 360 degrees)). In some embodiments, moving to a second position in the environment (e.g., 602) via a moving component includes tilting from the first tilt position to a second tilt position (and / or angle (e.g., 0 to 360 degrees) different from the first tilt position) via the moving component. Tiltping to a second position in the environment in response to detecting a first interaction enables the computer system to change its tilt position during a type of interaction, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing improved visual feedback to the user.

[0202] In some embodiments, when the computer system is in a first position, the computer system (e.g., 600) is in a first rotational position (and / or angle (e.g., 0 to 360 degrees)). In some embodiments, moving to a second position in the environment (e.g., 602) via a moving component includes rotating from the first rotational position to a second rotational position (and / or angle (e.g., 0 to 360 degrees) different from the first rotational position). Rotating to a second position in the environment in response to detecting a first interaction enables the computer system to change its rotational position during a type of interaction, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0203] In some embodiments, the first positioning includes a first lateral positioning. In some embodiments, the second positioning, moved via a moving component into the environment (e.g., 602), includes moving via a moving component from the first lateral positioning to a second lateral positioning different from the first lateral positioning (e.g., as shown in the image). Figures 6A to 6D (Location description). In response to detecting a first interaction, moving laterally to a second location in the environment based on determining that the first interaction is a first type of interaction, and not moving to the second location (where the second location is a lateral location) based on determining that the first interaction is a second type of interaction, enables the computer system to change the lateral location during one type of interaction, thereby performing an operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0204] It should be noted that the above text regarding process 700 (for example, Figure 7 The details of the process described herein also apply in a similar manner to the methods described below / above. For example, process 800 may optionally include one or more features of the various methods described above with reference to process 700. For example, a new word may be added to a word cloud using one or more techniques of process 800, and the new word may be moved to a second location to be displayed using one or more techniques of process 700. For the sake of brevity, these details will not be repeated below.

[0205] Figure 8 This is a flowchart illustrating a method for displaying content using a computer system according to some embodiments. Process 800 is executed at a computer system (e.g., 100, 200, and / or 600). Some operations in process 800 may be combined, some operations may be changed in order, and some operations may be omitted.

[0206] As described below, process 800 provides an intuitive way to display content. This method reduces the cognitive burden on the user when displaying content, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to display content faster and more efficiently, saving power and increasing the time interval between battery charging.

[0207] In some embodiments, process 800 is performed at a computer system (e.g., 600) that communicates with display components (e.g., projectors, displays, and / or touch-sensitive displays) and a microphone. In some embodiments, the computer system is a watch, phone, tablet, fitness tracker, processor, head-mounted display (HMD) device, public utility, media device, speaker, television, and / or personal computing device. In some embodiments, the computer system communicates with one or more output devices (e.g., display components, audio generation components, speakers, haptic output devices, displays, projectors, and / or touch-sensitive displays). In some embodiments, the computer system communicates with movable components (e.g., actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), movable bases, rotatable components, and / or rotatable bases).

[0208] When a user interface (e.g., 612 and / or 624) (e.g., home screen, application and / or user interface object) is displayed via a display component, the computer system detects (802) first voice input (e.g., 606, 618 and / or 622) (e.g., phrase, statement, question and / or answer) via a microphone.

[0209] In response to the detection of a first voice input (e.g., 606, 618 and / or 622), the computer system displays (804) a first set of one or more words (e.g., text and / or symbols) corresponding to the first voice input via a display component in a first manner (e.g., at a first size, as a highlighted set of one or more words, and / or as an emphasized set of one or more words).

[0210] While displaying a first group of one or more words corresponding to the first voice input (e.g., 606, 618 and / or 622), the computer system detects (806) a second voice input (e.g., 606, 618 and / or 622) (e.g., a phrase, a question and / or an answer) via a microphone.

[0211] In response to (808) detecting a second voice input (e.g., 606, 618, and / or 622), based on the determination that the second voice input (e.g., 606, 618, and / or 622) includes new words (e.g., 616, 620, and / or 628) (and in some embodiments, the new words are not included in the first voice input) and the new words (e.g., 616, 620, and / or 628) corresponding to the second voice input (e.g., 606, 618, and / or 622) should be added to a first group of one or more words, the computer system displays (810) the new words (e.g., 616, 620, and / or 628) corresponding to the second voice input (e.g., 606, 618, and / or 622) together with the display of the first group of one or more words (e.g., as part of the first group of one or more words and / or when the first group of one or more words is displayed and / or presented concurrently) in a first manner via a display component.

[0212] In response to (808) detecting a second voice input, and based on the determination that the second voice input (e.g., 606, 618, and / or 622) includes a new word corresponding to the second voice input (e.g., 616, 620, and / or 628) and that the new word corresponding to the second voice input should not be added to the first group of one or more words, the computer system displays (812) the second group of one or more words including the new word corresponding to the second voice input in a first manner via a display component, while stopping the first manner of displaying the first group of one or more words, wherein the second group of one or more words differs from the first group of one or more words (e.g., as described above in...). Figures 6A to 6D(Location description). In some embodiments, the second group of one or more words differs from the first group of one or more words (and in some embodiments, includes two or more words that differ from the first group of one or more words). In some embodiments, the second voice input differs from the first voice input. In some embodiments, the second voice input is separate from and identical to the first voice input. Based on determining that the second voice input includes new words and that the new words corresponding to the second voice input should be added to the first group of one or more words, the new words corresponding to the second voice input are displayed in a first manner along with the display of the first group of one or more words. Based on determining that the second voice input includes new words corresponding to the second voice input and that the new words corresponding to the second voice input should not be added to the first group of one or more words, the second group of one or more words including the new words corresponding to the second voice input are displayed in a first manner, while stopping the first manner of displaying the first group of one or more words. This allows the computer system to automatically add new words from the input to a group of words and display a new group of words when new words should not be added to that group, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing improved visual feedback to the user.

[0213] In some implementations, the new word (e.g., 616, 620, and / or 628) is the first new word. In some implementations, when displaying a second group of one or more words including the first new word (e.g., 616, 620, and / or 628) corresponding to the second voice input (e.g., 606, 618, and / or 622), the computer system detects a third voice input (e.g., 606, 618, and / or 622) via a microphone (e.g., different from the first voice input, different from the second voice input, separate from and the same as the first voice input, and / or separate from and the same as the second voice input). In some implementations, in response to the detection of a third voice input (e.g., 606, 618, and / or 622), and based on the determination that the third voice input (e.g., 606, 618, and / or 622) includes a second new word (e.g., 616, 620, and / or 628) different from the first new word, and that the second new word corresponding to the third voice input (e.g., 616, 620, and / or 628) should be added to a first group of one or more words, the computer system displays the second new word corresponding to the third voice input (e.g., as described above) via a display component together with the first group of one or more words (e.g., in a first manner). Figures 6A to 6D(Location description) (and in some embodiments, the display of the second group of one or more words in the first manner is simultaneously stopped). In some embodiments, based on the determination that the third voice input includes a second new word and that the second new word corresponding to the third voice input should not be added to the first group of one or more words, the computer system does not display the second new word corresponding to the third voice input along with the first group of one or more words via a display component (and in some embodiments, the computer system continues to display the second group of one or more words (e.g., in the first manner)). In some embodiments, based on the determination that the third voice input includes a second new word and that the second new word corresponding to the third voice input should not be added to the first group of one or more words, the computer does not display the second new word corresponding to the third voice input along with the first group of one or more words (e.g., in the first manner) via a display component (and in some embodiments, the computer system displays a new group of one or more words including the second new word). Based on the determination that the third voice input includes a second new word different from the first new word, and that the second new word corresponding to the third voice input should be added to the first group of one or more words to display the second new word corresponding to the third voice input together with the first group of one or more words, the computer system is able to detect additional input, and when it is determined that a new word corresponding to the new spoken input should be added to the first group of words, a group of words is redisplayed together with that new word, thereby performing the operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0214] In some implementations, the new word (e.g., 616, 620, and / or 628) is the third new word. In some implementations, when displaying a second group of one or more words including the third new word (e.g., 616, 620, and / or 628) corresponding to the second voice input (e.g., 606, 618, and / or 622), the computer system detects a fourth voice input (e.g., 606, 618, and / or 622) via a microphone (e.g., different from the first voice input, different from the second voice input, separate from and the same as the first voice input, and / or separate from and the same as the second voice input). In some implementations, in response to the detection of a fourth speech input (e.g., 606, 618, and / or 622) and based on the determination that the fourth speech input (e.g., 606, 618, and / or 622) includes a fourth new word (e.g., 616, 620, and / or 628) that is different from the third new word (e.g., 616, 620, and / or 628), and that the fourth new word corresponding to the fourth speech input should be added to a second group of one or more words, the computer system displays the fourth new word corresponding to the fourth speech input via a display component (e.g., as described above) together with the display of the second group of one or more words (e.g., in a first manner). Figures 6A to 6D (Location description) (and in some embodiments, the first group of one or more words is not displayed simultaneously). In some embodiments, based on the determination that the fourth voice input includes a fourth new word and that the fourth new word corresponding to the fourth voice input should not be added to the second group of one or more words, the computer system does not display the fourth new word corresponding to the fourth voice input via a display component together with the display of the second group of one or more words. In some embodiments, based on the determination that the fourth voice input includes a fourth new word and that the fourth new word corresponding to the fourth voice input should not be added to the second group of one or more words, the computer system (e.g., in a first manner) displays a group of one or more words including the fourth new word (e.g., different from the first group of one or more words and the second group of one or more words). Displaying the fourth new word corresponding to the fourth voice input together with the display of the second group of one or more words, based on the determination that the fourth voice input includes a fourth new word and that the fourth new word corresponding to the fourth voice input should be added to the second group of one or more words, enables the computer system to add new words by automatically adding new words for additional input together with the previous group of words, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0215] In some implementations, the new word (e.g., 616, 620, and / or 628) is the fifth new word. In some implementations, when displaying a second group of one or more words including the fifth new word (e.g., 616, 620, and / or 628) corresponding to the second voice input (e.g., 606, 618, and / or 622), the computer system detects the fifth voice input (e.g., 606, 618, and / or 622) via a microphone (e.g., different from the first voice input, different from the second voice input, separate from and the same as the first voice input, and / or separate from and the same as the second voice input). In some implementations, in response to the detection of a fifth voice input (e.g., 606, 618, and / or 622) and based on the determination that the fifth voice input (e.g., 606, 618, and / or 622) includes a sixth new word (e.g., 616, 620, and / or 628) that is different from the fifth new word (e.g., 616, 620, and / or 628), and that the sixth new word corresponding to the fifth voice input should not be added to a second group of one or more words, the computer system displays a third group of one or more words including the sixth new word corresponding to the fifth voice input via a display component (e.g., in a first manner), while stopping the display of the second group of one or more words in the first manner, wherein the third group of one or more words is different from the second group of one or more words (and in some implementations, different from the first group of one or more words). In some implementations, based on the determination that the fifth voice input includes a sixth new word and that the sixth new word corresponding to the fifth voice input should be added to a second group of one or more words, the computer system does not display a third group of one or more words including the sixth new word corresponding to the fifth voice input via a display component (e.g., in a first manner) (and in some implementations, does not stop the display of the second group of one or more words in the first manner). Based on the determination that the fifth voice input includes the sixth new word and that the sixth new word corresponding to the fifth voice input should not be added to the second group of one or more words, the system displays the third group of one or more words including the sixth new word corresponding to the fifth voice input, while stopping the display of the second group of one or more words in the first manner. This allows the computer system to continuously add new words to multiple groups of words used for additional input, thereby performing operations without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform operations, and providing improved visual feedback to the user.

[0216] In some implementations, the new word (e.g., 616, 620, and / or 628) is the seventh new word. In some implementations, when displaying a second group of one or more words including the seventh new word (e.g., 616, 620, and / or 628) corresponding to the second voice input (e.g., 606, 618, and / or 622), the computer system detects a sixth voice input (e.g., 606, 618, and / or 622) via a microphone (e.g., different from the first voice input, different from the second voice input, separate from and the same as the first voice input, and / or separate from and the same as the second voice input). In some implementations, in response to the detection of a sixth voice input (e.g., 606, 618, and / or 622) and based on the determination that the sixth voice input (e.g., 606, 618, and / or 622) includes an eighth new word (e.g., 616, 620, and / or 628) that is different from the seventh new word (e.g., 616, 620, and / or 628), and that the eighth new word corresponding to the sixth voice input should not be added to the corresponding set of one or more words (e.g., any set of words, a second set of one or more words, and / or a first set of one or more words), the computer system abandons displaying the seventh new word corresponding to the sixth voice input via a display component. In some implementations, when it is determined that the eighth new word corresponding to the sixth voice input is not a significant word, keyword, main word, and / or relevant word in relation to the context, interaction, and / or dialogue, it is determined that the eighth new word should not be added to the corresponding set of one or more words. In some implementations, the eighth word is a preposition, conjunction, and / or another part of speech that is considered unimportant. By determining that the sixth voice input includes an eighth new word that is different from the seventh new word, and that the eighth new word corresponding to the sixth voice input should not be added to the corresponding group of one or more words, the computer system can detect additional input and not add additional words that should not be added to the group of one or more words. This allows the system to perform operations without additional user input when a set of conditions has been met, reducing the amount of input required to perform operations and providing improved visual feedback to the user.

[0217] In some implementations, in response to the detection of a sixth voice input (e.g., 606, 618, and / or 622) and based on the determination that the sixth voice input (e.g., 606, 618, and / or 622) includes an eighth new word and that the eighth new word (e.g., 616, 620, and / or 628) should not be added to the word list, the computer system continues to display a second group of one or more words in a first manner via a display component. Displaying a second group of one or more words in a first manner based on the determination that the sixth voice input includes an eighth new word and that the eighth new word should not be added to the word list enables the computer system to maintain the display of the group of one or more words even when new input includes a new word but does not include words that should be added to that group of one or more words. This allows the system to perform an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0218] In some implementations, the second voice input (e.g., 606, 618 and / or 622) includes phrases containing new words (e.g., a fourth group of one or more words and / or short spoken expressions).

[0219] In some implementations, the new word (e.g., 616, 620, and / or 628) is the ninth new word. In some implementations, in response to detecting a second voice input (e.g., 606, 618, and / or 622) and based on determining that the second voice input (e.g., 606, 618, and / or 622) includes a tenth new word different from the ninth new word, the tenth new word corresponding to the second voice input (e.g., 616, 620, and / or 628) should be added to a first group of one or more words, the second voice input including the ninth new word, and the ninth new word corresponding to the second voice input (e.g., 616, 620, and / or 628) should be added to the first group of one or more words, the computer system concurrently displays the ninth new word corresponding to the second voice input and the tenth new word corresponding to the second voice input (e.g., as described above) via a display component, together with the display of the first group of one or more words, in a first manner. Figures 6A to 6D(Location description). In some embodiments, the phrase includes a tenth neologism. In some embodiments, based on the determination that the second voice input includes a tenth neologism, the tenth neologism corresponding to the second voice input should not be added to the first group of one or more words; the second voice input includes a ninth neologism, and the ninth neologism corresponding to the second voice input should be added to the first group of one or more words; the computer system displays the ninth neologism corresponding to the second voice input via a display component (e.g., in a first manner) when displaying the first group of one or more words, and does not display the tenth neologism corresponding to the second voice input. In some embodiments, based on the determination that the second voice input includes a tenth neologism, the tenth neologism corresponding to the second voice input should be added to the first group of one or more words; the second voice input includes a ninth neologism, and the ninth neologism corresponding to the second voice input should not be added to the first group of one or more words; the computer system displays the tenth neologism corresponding to the second voice input via a display component (e.g., in a first manner) when displaying the first group of one or more words, and does not display the ninth neologism corresponding to the second voice input. In some implementations, based on the determination that the second voice input includes a tenth new word, the tenth new word corresponding to the second voice input should not be added to the first group of one or more words; the second voice input includes a ninth new word, and the ninth new word corresponding to the second voice input should not be added to the first group of one or more words; the computer system, when displaying the first group of one or more words, does not display the ninth new word corresponding to the second voice input via a display component (e.g., in a first manner), and does not display the tenth new word corresponding to the second voice input. Based on the determination that the second voice input includes a tenth new word, the tenth new word corresponding to the second voice input should be added to the first group of one or more words, the second voice input includes a ninth new word, and the ninth new word corresponding to the second voice input should be added to the first group of one or more words, displaying the ninth new word corresponding to the second voice input and the tenth new word corresponding to the second voice input concurrently in a first manner along with the display of the first group of one or more words, the computer system is able to concurrently add multiple new words to a group of words for input with multiple new words, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing improved visual feedback to the user.

[0220] In some embodiments, the second voice input (e.g., 606, 618, and / or 622) includes an eleventh new word (e.g., 616, 620, and / or 628) between the ninth new word (e.g., 616, 620, and / or 628) and the tenth new word (e.g., 616, 620, and / or 628) in the second voice input. In some embodiments, in response to detecting the second voice input (e.g., 606, 618, and / or 622) (e.g., and based on determining that the eleventh new word should not be added to the first group of one or more words and / or any corresponding group of one or more words), the computer system refrains from displaying the eleventh new word (e.g., 616, 620, and / or 628) corresponding to the second voice input via a display component (and in some embodiments, in the first group of one or more words, the second group of one or more words, or additional groups of one or more words) (e.g., simultaneously displaying the ninth new word corresponding to the second voice input and the tenth new word corresponding to the second voice input via the display component together with the display of the first group of one or more words (e.g., in a first manner)). In some implementations, in response to the detection of a second voice input, the computer system does not add an eleventh new word corresponding to the second voice input to a group of one or more words (a first group of one or more words, a second group of one or more words, or additional groups of one or more words). Not displaying the eleventh new word corresponding to the second voice input in response to its detection allows the computer system to automatically ignore some words between other new words that should be added, thereby performing the operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0221] In some implementations, the new word (e.g., 616, 620, and / or 628) is the twelfth new word. In some implementations, in response to detecting a second voice input (e.g., 606, 618, and / or 622), based on the determination that the second voice input (e.g., 606, 618, and / or 622) includes a thirteenth new word (e.g., 616, 620, and / or 628) that is different from the twelfth new word (e.g., 616, 620, and / or 628), and the thirteenth new word corresponding to the second voice input should not be added to the first group of one or more words, and the twelfth new word corresponding to the second voice input should not be added to the first group of one or more words, the computer system concurrently displays the thirteenth new word corresponding to the second voice input and the twelfth new word corresponding to the second voice input as part of the second group of one or more words via a display component (e.g., in a first manner) (e.g., as described above). Figures 6A to 6D(Location description) (e.g., simultaneously stopping the display of one or more words in the first group in a first manner). In some embodiments, based on determining that the second voice input includes a thirteenth new word, the thirteenth new word corresponding to the second voice input should be added to one or more words in the first group, and the twelfth new word corresponding to the second voice input should not be added to one or more words in the first group. The computer system displays one or more words in the second group including the twelfth new word corresponding to the second voice input via a display component (e.g., in a first manner), while stopping the display of one or more words in the first group and not displaying the thirteenth new word corresponding to the second voice input. In some embodiments, based on determining that the second voice input includes a thirteenth new word, the thirteenth new word corresponding to the second voice input should not be added to one or more words in the first group, and the twelfth new word corresponding to the second voice input should be added to one or more words in the first group. The computer system displays one or more words in the second group including the thirteenth new word corresponding to the second voice input via a display component (e.g., in a first manner), while stopping the display of one or more words in the first group and not displaying the twelfth new word corresponding to the second voice input. In some implementations, based on the determination that the second voice input includes a thirteenth new word, the thirteenth new word corresponding to the second voice input should not be added to the first group of one or more words, and the twelfth new word corresponding to the second voice input should not be added to the first group of one or more words. The computer system displays a second group of one or more words excluding the twelfth new word corresponding to the second voice input via a display component, while simultaneously stopping the display of the first group of one or more words and not displaying the thirteenth new word corresponding to the second voice input. By concurrently displaying the thirteenth new word corresponding to the second voice input and the twelfth new word corresponding to the second voice input as part of the second group of one or more words, based on the determination that the second voice input includes a thirteenth new word, and that the thirteenth new word corresponding to the second voice input should not be added to the first group of one or more words, and that the twelfth new word corresponding to the second voice input should not be added to the first group of one or more words, the computer system can simultaneously display multiple new words including additional new words to that group of one or more words. This allows the system to perform operations without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0222] In some embodiments, the second voice input (e.g., 606, 618, and / or 622) includes a fourteenth new word (e.g., 616, 620, and / or 628) that is different from the thirteenth new word (e.g., 616, 620, and / or 628) and the twelfth new word in the second voice input. In some embodiments, in response to the detection of the second voice input (e.g., 606, 618, and / or 622) (e.g., and based on the determination that the fourteenth new word should not be added to the first group of one or more words and / or any corresponding group of one or more words), the computer system refrains from displaying the fourteenth new word (e.g., 616, 620, and / or 628) corresponding to the second voice input via a display component (and in some embodiments, in the first group of one or more words, the second group of one or more words, or additional groups of one or more words) (e.g., simultaneously displaying the thirteenth and twelfth new words corresponding to the second voice input via a display component together with the display of the first group of one or more words (e.g., in a first manner)). In some implementations, in response to the detection of a second voice input, the computer system does not add a fourteenth new word corresponding to the second voice input to a group of one or more words (a first group of one or more words, a second group of one or more words, or additional groups of one or more words). Not displaying the fourteenth new word corresponding to the second voice input in response to its detection allows the computer system to ignore words in the input that are between other words being displayed (e.g., adding relevant context for important words but not adding other words that depend on the context), thereby performing an operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0223] In some implementations, the second voice input (e.g., 606, 618, and / or 622) does not include adding new words (e.g., 616, 620, and / or 628) to a specific set of one or more words (e.g., as described above). Figures 6A to 6DThe system displays explicit instructions (e.g., a group of one or more words explicitly mentioning and / or commanding) for the location description (e.g., a first group of one or more words and / or a second group of one or more words). Based on the determination that the second voice input includes a new word and that the new word corresponding to the second voice input, which does not explicitly indicate that a new word should be added, it displays the new word corresponding to the second voice input in a first manner along with the display of the first group of one or more words. Based on the determination that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, it displays the second group of one or more words including the new word corresponding to the second voice input in a first manner via a display component, while stopping the first manner of displaying the first group of one or more words. This allows the computer system to automatically add a new word to the group of one or more words even when the input does not explicitly indicate that a word should be added, and to display a new group of one or more words when it is determined that a new word should be added to the group of one or more words. This allows the system to perform an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0224] In some implementations, in response to the detection of a second voice input (e.g., 606, 618, and / or 622), based on the determination that the second voice input (e.g., 606, 618, and / or 622) includes a fifteenth new word (e.g., 616, 620, and / or 628) and that the new word corresponding to the second voice input should be added to a first group of one or more words, the computer system displays one or more indications of the first group of one or more words (e.g., list headers, representations of the first group of one or more words, and / or text), and simultaneously displays the fifteenth new word (e.g., 616, 620, and / or 628) corresponding to the second voice input in a first manner together with the display of the first group of one or more words (e.g., as part of the first group of one or more words and / or when the first group of one or more words is displayed concurrently and / or presented). In some implementations, in response to the detection of a second voice input, based on the determination that the second voice input (e.g., 606, 618, and / or 622) includes a fifteenth new word corresponding to the second voice input (e.g., 616, 620, and / or 628) and that the fifteenth new word corresponding to the second voice input should not be added to the first group of one or more words, the computer system displays a second group of instructions corresponding to the second group of one or more words, different from the first group of instructions, while simultaneously displaying the new word corresponding to the second voice input (e.g., 616, 620, and / or 628) in a first manner (e.g., as described above). Figures 6A to 6D(Location description) (e.g., and simultaneously not displaying the first group of one or more words in the first manner). Based on determining that the second voice input includes a fifteenth new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, a first group of one or more instructions corresponding to the first group of one or more words is displayed, while the fifteenth new word corresponding to the second voice input is displayed in the first manner along with the display of the first group of one or more words. Based on determining that the second voice input includes a fifteenth new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, a second group of instructions corresponding to the second group of one or more words is displayed, while the new word corresponding to the second voice input is displayed in the first manner. This allows the computer system to display different instructions for the new word displayed with that group of one or more words and for a new group of one or more words to be displayed with the new word, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing improved visual feedback to the user.

[0225] In some embodiments, a first group of one or more words is displayed in a first arrangement (e.g., the organization, order, sequence, spacing, and / or shape of the displayed words). In some embodiments, a second group of one or more words is displayed in a second arrangement different from the first arrangement (e.g., as described above). Figures 6A to 6D (Location description). Based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to a first group of one or more words, the system displays the new word corresponding to the second voice input in a first manner and a first arrangement, together with the display of the first group of one or more words. Based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, the system displays a second group of one or more words including the new word corresponding to the second voice input via a display component in a first manner and a second arrangement, while simultaneously stopping the display of the first group of one or more words in the first manner. This allows the computer system to automatically display the new word in the first arrangement when it is determined that a new word should be added to the group of one or more words, and to display a new group of one or more words in a different arrangement. This allows the system to perform an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0226] In some implementations, displaying a first group of one or more words includes displaying a first group of one or more media representations corresponding to the first group of one or more words (e.g., video, image, animation, 3D rendering, augmented reality overlay, motion graphics, data visualization, and / or digital art). In some implementations, displaying a second group of one or more words includes displaying a second group of one or more media representations corresponding to the second group of one or more words, wherein the second group of one or more media representations differs from the first group of one or more media representations (e.g., as described above). Figures 6A to 6D (Location description). Based on determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to a first group of one or more words, the system displays the new word corresponding to the second voice input in a first manner, including displaying a representation of the first group of one or more media, together with the display of the first group of one or more words. Simultaneously, based on determining that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, the system displays a second group of one or more words including the new word corresponding to the second voice input via a display component in a first manner, including a representation of a second group of one or more media, while stopping the display of the first group of one or more words in the first manner. This allows the computer system to automatically add the new word to the group of one or more words including the first media when it is determined that the new word should be added to the group of one or more words including the second media, and to display a new group of one or more words. This allows the system to perform an operation without requiring additional user input when a set of conditions has already been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0227] In some implementations, determining that new words (e.g., 616, 620, and / or 628) corresponding to the second speech input (e.g., 606, 618, and / or 622) should be added to a first group of one or more words (e.g., in a first manner) includes determining that the new words are key (e.g., relevant, important) words (e.g., core terms, central words, and / or basic communication elements) in the second speech input. The system determines that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to a first group of one or more words, including determining that the new word is a keyword in the second voice input, and displays the new word corresponding to the second voice input in a first manner together with the display of the first group of one or more words. It also determines that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, and displays a second group of one or more words including the new word corresponding to the second voice input via a display component in a first manner, while stopping the first manner of displaying the first group of one or more words. This allows the computer system to automatically add the new keyword to the group of one or more words when it is determined that the new word should be added, and to display a new group of one or more words. This allows the system to perform an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0228] In some implementations, determining whether a new word (e.g., 616, 620, and / or 628) is a key (e.g., relevant, important) word (e.g., core term, central word, and / or basic communication element) in the second voice input (e.g., 606, 618, and / or 622) includes: determining that the new word (e.g., 616, 620, and / or 628) is a keyword (e.g., in the second voice input) based on determining that the current context is a first context (e.g., based on previous voice input to the second voice input and / or the first voice input, and / or the presence of a user (e.g., user, person, animal, another computer system and / or object different from the computer system)) (e.g., the context in which the computer system is operating, the context of internal dialogue within the computer system, and / or the environmental context); and determining that the new word is not a keyword (e.g., in the second voice input) based on determining that the current context is a second context different from the first context. The system determines that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to a first group of one or more words, including determining that the new word is a keyword based on the current context being a first context. The new word corresponding to the second voice input is displayed in a first manner along with the display of the first group of one or more words. It also determines that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, including determining that the new word is not a keyword based on the current context being a second context. Simultaneously, the system stops displaying the first group of one or more words in the first manner. This allows the computer system to automatically add the new word to the group of one or more words and display a new group of one or more words when the context determines that the new word should be added to the group of one or more words. This enables the system to perform the operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0229] In some implementations, determining whether a new word (e.g., 616, 620, and / or 628) corresponding to a second voice input (e.g., 606, 618, and / or 622) should be added to a first group of one or more words includes: determining that the new word (e.g., 616, 620, and / or 628) should be added to a first group of one or more words based on the determination that the new word (e.g., 616, 620, and / or 628) is relevant to the context of the first group of one or more words (e.g., previous groups of one or more words and / or the user); and determining that the new word (e.g., 616, 620, and / or 628) should not be added to a first group of one or more words based on the determination that the new word (e.g., 616, 620, and / or 628) is not relevant to the context of the first group of one or more words. The system determines that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to a first group of one or more words, including determining that the new word is context-dependent with the first group of one or more words, and displays the new word corresponding to the second voice input in a first manner along with the display of the first group of one or more words. It also determines that the second voice input includes a new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, including determining that the new word is not context-dependent with the first group of one or more words, and displays a second group of one or more words including the new word corresponding to the second voice input in a first manner via a display component, while stopping the first manner of displaying the first group of one or more words. This allows the computer system to automatically add the new word to the group of one or more words that is context-dependent with the first group of one or more words when it is determined that the new word should be added, and to display a new group of one or more words. This allows the system to perform an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation and providing improved visual feedback to the user.

[0230] In some embodiments, upon detecting a second voice input (e.g., 606, 618, and / or 622) (and in some embodiments, and based on determining that the second voice input includes one or more new words (and in some embodiments, one or more new words are not included in the first voice input) and the one or more new words corresponding to the second voice input should be added to a first group of one or more words): Firstly, a first portion of the second voice input (e.g., 606, 618, and / or 622) is detected; in response to detecting the first portion of the second voice input (e.g., 606, 618, and / or 622), words corresponding to the first portion of the second voice input (e.g., 616, 620, and / or 628) are displayed via a display component along with the first group of one or more words (and in some embodiments, based on determining that the second voice input includes one or more new words (and in some embodiments, one or more new words are not included in the first voice input) and the one or more new words corresponding to the second voice input should be added to a first group of one or more words): The words corresponding to the first part of the second voice input should be added to the first group of one or more words; at a second time (e.g., after the first time), a second part of the second voice input (e.g., 606, 618, and / or 622) different from the first part is detected; and in response to the detection of the second part of the second voice input (e.g., 606, 618, and / or 622), (and in some embodiments, based on determining that the words corresponding to the second part of the second voice input should be added to the first group of one or more words), the words corresponding to the second part of the second voice input (e.g., 616, 620, and / or 628) are displayed via a display component together with the first group of one or more words (e.g., in a first manner) (e.g., concurrently with the display of the words corresponding to the first part of the second voice input and the first group of one or more words). In some embodiments, the computer system changes the first group of one or more words in response to the detection of the second voice input and / or upon the detection of the second voice input. In some embodiments, changing the first group of one or more words includes moving at least one word in the first group of one or more words, changing the first group of one or more words to the second group of one or more words, and / or adding a new word corresponding to the second voice input to the first group of one or more words. In response to detecting a second part of the second voice input, the system displays words corresponding to the second part of the second voice input along with one or more words in a first group, and in response to detecting a first part of the second voice input, the system displays words corresponding to the first part of the second voice input along with one or more words in a first group, enabling the computer system to dynamically and automatically display words when input is received, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0231] In some embodiments, based on the determination that the second voice input (e.g., 606, 618, and / or 622) has a first speed, the first time and the second time are separated by a first time interval. In some embodiments, based on the determination that the second voice input (e.g., 606, 618, and / or 622) has a second speed different from the first speed, the first time and the second time are separated by a second time interval different from the first time interval. In some embodiments, the faster the voice input, the faster the words are displayed. In some embodiments, the slower the voice input, the slower the words are displayed. In response to detecting a second part of the second voice input, words corresponding to the second part of the second voice input are displayed together with one or more words in a first group, and in response to detecting a first part of the second voice input, words corresponding to the first part of the second voice input are displayed together with one or more words in a first group, wherein the second voice input is determined to have a first speed, a first time and a second time are separated by a first time interval, and the second voice input is determined to have a second speed different from the first speed, the first time and the second time are separated by a second time interval, enabling the computer system to automatically display words at a dynamic speed based on the speed of the input, thereby performing an operation without requiring additional user input when a set of conditions has been met, reducing the amount of input required to perform the operation, and providing the user with improved visual feedback.

[0232] In some embodiments, based on the determination that a first portion of the second speech input (e.g., 606, 618, and / or 622) has a first set of one or more characteristics (e.g., pitch, tone, and / or volume), the word corresponding to the first portion of the second speech input (e.g., 616, 620, and / or 628) is a first size (e.g., text width and / or height). In some embodiments, based on the determination that a first portion of the second speech input (e.g., 606, 618, and / or 622) has a second set of one or more characteristics different from the first set of one or more characteristics (e.g., pitch, tone, and / or volume), the word corresponding to the first portion of the second speech input (e.g., 616, 620, and / or 628) is a second size different from the first size. In some embodiments, the louder the corresponding portion of the speech input, the louder the word corresponding to the corresponding portion of the speech input.

[0233] In some embodiments, based on determining that a word (e.g., 616, 620, and / or 628) corresponding to a first portion of the second speech input (e.g., 606, 618, and / or 622) has a first relevance score relative to a first group of one or more words, the word (e.g., 616, 620, and / or 628) corresponding to a first portion of the second speech input (e.g., within and / or relative to the first group of one or more words) is displayed at a first location relative to the first group of one or more words (e.g., within and / or relative to the first group of one or more words). In some embodiments, based on determining that a word (e.g., 616, 620, and / or 628) corresponding to a first portion of the second speech input (e.g., 606, 618, and / or 622) has a second relevance score different from the first relevance score relative to the first group of one or more words, the word corresponding to the first portion of the second speech input is displayed at a second location different from the first location relative to the first group of one or more words (e.g., within and / or relative to the first group of one or more words). In some implementations, words displayed at the top (and / or right and / or left) may be more or less relevant to the corresponding set of words compared to words displayed at the bottom (and / or left and / or right).

[0234] In some implementations, stopping the display of one or more words in the first manner includes removing the display of one or more words in the first group.

[0235] In some implementations, stopping the display of a first group of one or more words in a first manner includes displaying the first group of words in a second manner different from the first manner via a display component (e.g., as described above). Figures 6A to 6D (Location description). In some embodiments, when displayed in a first manner, the first group of one or more words is visually more prominent and / or emphasized compared to when the first group of one or more words is displayed in a second manner.

[0236] It should be noted that the above text regarding process 800 (for example, Figure 8 The details of the process described herein also apply in a similar manner to the methods described below / above. For example, process 700 may optionally include one or more features of the various methods described above with reference to process 800. For example, a new word may be added to a word cloud using one or more techniques of process 800, and the new word may be moved to a second location to be displayed using one or more techniques of process 700. For the sake of brevity, these details will not be repeated below.

[0237] Figures 9A to 9J Exemplary user interfaces for controlling a user interface according to some embodiments are illustrated. The user interfaces in these figures are used to illustrate the processes described below, including... Figures 10 to 14 The process in.

[0238] Figures 9A to 9J The computer system 900 is illustrated as a tablet computer displaying different user interfaces. It should be understood that the computer system 900 can be other types of computer systems, such as smartphones, smartwatches, laptops, public devices, smart speakers, accessories, personal gaming systems, desktop computers, fitness trackers, and / or head-mounted display (HMD) devices. In some embodiments, the computer system 900 includes one or more sensors (e.g., one or more cameras, one or more LiDAR detectors, one or more motion sensors, one or more infrared sensors, and / or one or more microphones) and / or communicates with them. In some embodiments, the computer system 900 includes one or more output devices (e.g., displays, projectors, touch-sensitive displays, and / or speakers) and / or communicates with them. In some embodiments, the computer system 900 includes one or more moving components (e.g., actuators, movable bases, rotatable components, and / or rotatable bases) and / or communicates with them. In some examples, the computer system 900 includes one or more components and / or features described above with respect to computer system 100 and / or electronic device 200.

[0239] Figures 9A to 9J Examples illustrate one or more scenarios in which computer system 900 views and / or recalls previous interactions. In some embodiments, previous interactions include previous conversations that the user has had with computer system 900 and / or another user (e.g., where computer system 900 has recorded previous conversations), previous presentations of one or more topics that computer system 900 has given to the user and / or that the user has given to computer system 900, and / or previous sets of one or more inputs provided by the user and previous sets of one or more outputs generated by the computer system. In some cases, managing interactions involves the user utilizing a digital assistant of computer system 900. In some embodiments, the digital assistant is represented by an avatar (such as avatar 904). In some embodiments, computer system 900 updates avatar 904 to indicate to the user that computer system 900 is interacting with one or more users in the environment. For example, computer system 900 may update avatar 904 such that avatar 904 appears to be looking at one or more users in the environment, looking away, talking to them, nodding, and / or gesturing to them. Figure 9AAs illustrated, Avatar 904 is a face with one or more human characteristics. In some embodiments, Avatar 904 has different appearances (e.g., different colors (e.g., color groups, skin color, red, orange, yellow, green, blue and / or purple), textures (e.g., skin, hair, fur, scales, plastic, glass, feathers and / or wood), accessories (e.g., hats, glasses, monocles, wands, books, collars, bows, wings, halos and / or crowns) and / or facial types (e.g., human, animal, anthropomorphic object, alien, mundane face, fantasy creature and / or a collection of similar faces)).

[0240] Regarding interaction, in some embodiments, the user provides verbal input or some other type of input (such as touch input, air gestures, and / or input to one or more hardware buttons) to interact with computer system 900 and / or a digital assistant represented by avatar 904. In some embodiments, interaction with the digital assistant provided by computer system 900 is initiated in response to the detection of input. While the interaction is in progress, computer system 900 may display content and provide one or more audio and / or haptic outputs to interact with the user, such as guiding the user to view trip details and / or answering one or more questions from the user. In some embodiments, interactions that occur between computer system 900 and the user may be stored and later invoked by detecting one or more inputs from the user. In some embodiments, in response to the detection of one or more inputs from the user (such as verbal input), computer system 900 displays a summary of the previous interaction. In some embodiments, the summary is dynamically generated using audio output, wherein computer system 900 provides an interactive overview of the previous interaction. In some implementations, the summary may include one or more content items (e.g., application and / or media items) for completing tasks related to the previous interaction and / or one or more other highlights, such as relevant content, updated content, and / or content that was not initially included in the previous interaction.

[0241] In some implementations, a verbal request to discuss a previous interaction can be an explicit (e.g., clear, unambiguous) request. For example, a user might audibly state, "Show me the music recommendations again" or "Do you remember our conversation about music yesterday?" (e.g., a statement directly related to the previous interaction). In some implementations, a verbal request to discuss a previous interaction can be an implicit request. For example, a user might audibly state, "Your previous music recommendations were good" (e.g., not necessarily a statement loosely related to the previous interaction). Figure 9AAt this point, computer system 900 detects verbal input 905a (e.g., “What should I look at?”). In some embodiments, computer system 900 detects one or more other types of input, such as tapping input, air gestures, mouse clicks, and / or gaze input, and performs techniques similar to those described herein for verbal input (such as verbal input 905a).

[0242] like Figure 9B As illustrated, in response to the detection of verbal input 905a, computer system 900 reduces the size of avatar 904 and displays a first content item 906 (e.g., "First Movie") (e.g., movie recommendations). Figure 9B As illustrated, when computer system 900 presents (e.g., introduces, describes) a first content item 906, computer system 900 displays an avatar 904 facing (e.g., looking at) the first content item 906 in the upper left corner (e.g., avatar 904 may face the content item and / or category as an indication that computer system 900 is presenting the content item and / or category (e.g., focusing on the content item and / or category and outputting an audio description corresponding to the content item and / or category)). In some embodiments, if computer system 900 detects touch input, in response to the detection of touch input, computer system 900 may display the content item at the location where the touch input was detected. For example, if touch input is detected in the upper right corner, computer system 900 may display the content item and / or category in the upper right corner. Figure 9B At that point, computer system 900 outputs an audible description corresponding to the first content item 906.

[0243] In some implementations, if input is detected pointing to a content item and / or category, the computer system 900 may stop displaying any other content items and / or categories that are currently displayed in response to the detection of input. For example, in a scenario where the computer system 900 detects input pointing to a first content item, the computer system 900 may stop displaying a second content item that is displayed before or after the first content item in response to the detection of input.

[0244] like Figure 9CAs illustrated, computer system 900 automatically displays a second content item 908 (e.g., "Second Movie") (e.g., movie recommendation) after detecting verbal input 905a and no additional input is detected. Furthermore, computer system 900 outputs an audible description corresponding to the second content item and stops outputting an audible description corresponding to the first content item. To display the second content item 908 in the center of the user interface, computer system 900 moves the first content item 906 to the upper left corner and moves the avatar 904 to the lower left corner. In this example, computer system 900 attempts to display the currently discussed content item (e.g., for which audio is being output) in the center of the user interface and other items (e.g., previously discussed content items) in other locations on the user interface. In some embodiments, displaying the currently discussed content item in the center allows the user to focus on those items compared to those not displayed in the center.

[0245] In addition, Figure 9C In this case, computer system 900 dynamically moves avatar 904 when displaying different content items. For example... Figure 9C As illustrated, computer system 900 stops displaying avatar 904 facing the first content item 906, and instead displays avatar 904 facing the second content item 908. In some embodiments, when computer system 900 presents the second content item 908 (e.g., via audio output), avatar 904 can be updated to stop facing the old content item and / or category and instead face the new content item and / or category. Therefore, in Figure 9C At this point, avatar 904 appears to be facing the second content item 908 because computer system 900 is outputting audio corresponding to the second content item 908. In some embodiments, when a new content item and / or category is displayed, avatar 904 is updated to appear to be facing the new content item and / or category before moving. In some embodiments, when a new content item and / or category is displayed, avatar 904 is updated to appear to be facing the new content item and / or category while moving. In some embodiments, when a new content item and / or category is displayed, avatar 904 is updated to appear to be facing the new content item and / or category after moving.

[0246] In some implementations, computer system 900 may display avatar 904 on different sides of content items and / or categories. For example, if displaying a music content item, computer system 900 may display avatar 904 above, below, to the right, to the left, at the top, behind, and / or at an angle to the content item and / or category. In some implementations, computer system 900 may not display avatar 904 in the same placement for different content items and / or categories. For example, if avatar 904 is displayed to the right of the music category, computer system 900 may display avatar 904 below the photo category. In some implementations, avatar 904 may face or turn away from the user (e.g., avatar 904 may look back and forth between the user and the content item currently being presented by computer system 900). In some implementations, after a predetermined amount of time, computer system 900 updates avatar 904 such that avatar 904 appears to stop facing the content item and / or category. In this example, avatar 904 may appear to be looking towards the user and / or the environment. In some implementations, once the avatar 904 appears to stop facing the content item and / or category, the avatar 904 will not look back at the content item and / or category until the computer system 900 outputs audio corresponding to the specific item and / or after a predetermined time period.

[0247] like Figure 9D As illustrated, computer system 900 displays a third content item 910 (e.g., "First TV Program") (e.g., TV program recommendation). Additionally, computer system 900 displays category 950 (e.g., a category corresponding to the movie content item). Computer system 900 displays a second content item 908 overlapping with a first content item 906 within category 950. Computer system 900 overlaps the first content item 906 and the second content item 908 because it determines that the first content item 906 and the second content item 908 belong to the same category of items (e.g., "Movie"). In some embodiments, computer system 900 may visually group content without causing it to overlap. For example, computer system 900 may display content items belonging to the same category as (e.g., along a shared edge) adjacent to each other, and / or may specifically arrange content items belonging to the same category to display them without causing them to overlap. Figure 9D At this point, computer system 900 stops outputting the audible description corresponding to the second content item 908 and outputs the audible description corresponding to the third content item 910. For example... Figure 9D As illustrated, when computer system 900 presents third content item 910, it will transform into 904 and display it facing the third content item 910.

[0248] Further explanation of the purpose of the categories may aid understanding. Computer system 900 may display similar content (e.g., as defined by the computer system and / or user input) within categories (e.g., groupings of content). In instances where multiple content items are placed within the same category, computer system 900 may visually group them in a structural (e.g., order and / or organization) (e.g., overlapping) manner to convey similarity. In some embodiments, if a new content item is added to an interaction with a pre-existing category, or if the new content item is identical to a content item in an existing category, computer system 900 may display the new content item within that category (e.g., overlapping with other content items in that category). For example, if a music video category exists and a new music video content item is introduced, computer system 900 may automatically display the new music video content item within the existing music video category (e.g., overlapping with content items already displayed in that category). In some embodiments, if a new content item is added to an interaction with a pre-existing category, or if the new content item is not identical to a content item in an existing category, computer system 900 may display the new content item within a new category. For example, if a music video category exists and a photo content item is introduced, the computer system 900 can automatically create a new photo category and display the new photo content item within that new category. In some implementations, the computer system 900 can display a source indicator corresponding to the content item (e.g., an indicator showing the source of the content item). For example, if a movie content item is only available on a certain streaming service, the computer system 900 can display a source indicator to indicate this information to the user.

[0249] In some implementations, when outputting audio descriptions corresponding to the displayed content items, computer system 900 may display new content items and visually group them based on existing categories. For example, if a movie category exists and a new movie content item is displayed, computer system 900 may automatically categorize the new movie content item within an existing movie category when outputting its audio description.

[0250] In some implementations, computer system 900 may visually group content items after ceasing to output audio descriptions corresponding to the content items. For example, once computer system 900 stops outputting audio descriptions for movie content items, computer system 900 may visually group the movie content item with one or more other movie content items. In some implementations, computer system 900 may visually group content items while outputting audio descriptions corresponding to the content items. For example, while computer system 900 outputs audio descriptions for movie content items, computer system 900 may visually group the movie content item with one or more other movie content items. In some implementations, content items are not visually grouped when interaction begins. In this example, content may be grouped after the user completes the interaction. In another example, content may be grouped while the user is still engaged in the interaction.

[0251] In some implementations, the computer system 900 does not visually group all displayed content items. For example, if two movie content items, three music content items, and one application content item are displayed, the computer system 900 may visually group the movie content items into one category, the music content items into one category, but not the application content items into one category.

[0252] In some implementations, at least one content item may be visually grouped with another content item. For example, if two music content items are displayed one after the other with a video content item, the computer system 900 may display categories for the two music content items. In some implementations, at least one content item will not be visually grouped. For example, if two movie content items exist, the computer system 900 does not necessarily have to group them into one category. This may happen if the computer system 900 groups based on different characteristics other than media type and / or content type (e.g., genre, runtime, time period, and / or audience response). While the previous examples used movies as examples, it should be recognized that this is merely an example, and the techniques described herein can work with other content items and / or content types.

[0253] like Figure 9E As illustrated, computer system 900 displays category 950 (e.g., a category corresponding to movie content items) and category 952 (e.g., a category corresponding to television program content items). Category 950 includes a first content item 906 and a second content item 908. Category 952 includes a third content item 910 and a fourth content item 912 (e.g., television program recommendations). Figure 9E As illustrated, computer system 900 displays avatar 904 in the right center of user interface 902. Figure 9EAt that point, computer system 900 outputs an audible description corresponding to the fourth content item 912. Figure 9E At this point, computer system 900 detects verbal input 905e (e.g., “What should I hear?”), and this verbal input interruption corresponds to the output of the audio description of the fourth content item 912. In some embodiments, the interruption occurs when the user speaks (e.g., directs verbal input to computer system 900) while computer system 900 is outputting the audio description (e.g., the user speaks through computer system 900). In some embodiments, the interruption occurs when the user speaks while computer system 900 “pauses” or when there is a natural pause in the output of the audio description.

[0254] exist Figure 9E In this context, computer system 900 updates avatar 904 so that avatar 904 appears to shift its gaze away from one or more content items to look at the environment and / or the user within it. Computer system 900 in... Figure 9E The computer system 900 updates its avatar 904 to indicate that it is listening to the user in the environment that caused the interruption. In other words, the computer system 900 updates its avatar 904 to indicate that it is listening. In some embodiments, the computer system 900 minimizes the user interface when an interruption is detected. In some embodiments, the computer system 900 may minimize the user interface even when no interruption is detected, such as when the computer system 900 detects that the interaction has been completed. In some embodiments, when an interruption is detected, the computer system 900 may change the user interface in ways other than minimizing the content, such as fading out the content on the user interface, enlarging the content on the user interface, stopping the display of the content on the user interface, changing the color of the content on the user interface, increasing the opacity of the content on the user interface, displaying instructions, and / or moving the content on the user interface. It should be understood that although an interruption is discussed as being detected in response to the detection of verbal input, other types of input can also make an interruption detectable, such as air gestures detected when the computer system is outputting an audio description and / or when it detects that the user has removed their gaze from the computer system 900 for more than a period of time.

[0255] In some implementations, if a response to an interrupt is not required, the computer system 900 can be re-displayed. Figure 9D The user interface (e.g., the user interface displayed before the interruption). For example, if the user creates an incomprehensible interruption (e.g., an interruption that the computer system 900 cannot understand and / or does not understand), such as loud noise and / or incomprehensible verbal input, the computer system 900 may revert to the user interface displayed before the interruption occurred.

[0256] In some implementations, if the computer system 900 detects an interruption corresponding to a content item and / or category for which the computer system 900 is outputting an audio description, the computer system 900 will not display the interaction in a second manner (e.g., a minimized manner). For example, if the computer system 900 outputs an audio description corresponding to a movie content item, and an interruption corresponding to the movie content item is detected, the computer system 900 will not display the interaction in a second manner (e.g., the computer system 900 may continue to display the interaction in a first manner).

[0257] In some implementations, after displaying the interaction in a second manner in response to an interruption, if the computer system 900 detects an interruption corresponding to a first content item and / or category, the computer system 900 may display the interaction in a first manner. For example, if the first content is a music content item, and if the computer system 900 is displaying the interaction, and if the computer system 900 detects an interruption corresponding to the music content item, the computer system 900 will stop displaying the interaction in the second manner and may display the interaction in the first manner. If the computer system 900 is displaying the interaction, and if the computer system 900 detects an interruption not corresponding to a music content item, the computer system 900 may continue displaying the interaction in the second manner. In some implementations, if the computer system 900 is outputting an audio description corresponding to a first content item and detects an interruption, the computer system 900 will begin outputting an audio description corresponding to a second content item. For example, if the computer system 900 is outputting an audio description corresponding to a movie content item and detects an interruption, the computer system 900 will output an audio description corresponding to a music content item. In some implementations, if the computer system 900 is outputting an audio description corresponding to the first content item and no interruption is detected, the computer system 900 will continue outputting the audio description corresponding to the first content item. (See previous section.) Figure 9E The computer system 900 detects verbal input 905e and determines that verbal input 905e is a request that requires a response.

[0258] like Figure 9F As illustrated, in response to the detection of verbal input 905e, computer system 900 visually overlaps movie and television content items within category 954 (e.g., categories containing different content types). Furthermore, computer system 900 displays a fifth content item 914 (e.g., a first song) (e.g., a song recommendation). In some embodiments, computer system 900 stops displaying movie and television content items because it has been determined that the interaction is different from the interaction corresponding to the song content item. For example, in some embodiments, "What should I listen to?" is a different interaction from "What should I watch?".

[0259] like Figure 9GAs illustrated, in response to the detection of verbal input 905e, since it is determined that the movie and television content item originates from an interaction different from the interaction corresponding to the song content item, the computer system 900 stops displaying the movie and television content item (e.g., category 954). In response to the detection of verbal input 905e, the computer system 900 displays a sixth content item 916 (e.g., "Second Song") below the fifth content item 914 (e.g., song recommendation). Figure 9G As illustrated, when computer system 900 presents the sixth content item 916, it will transform into 904 and display it facing the sixth content item 916. It is worth noting that, in Figure 9G In this context, song content items are configured differently from TV and movie content items because different interactions can result in different layouts and / or configurations of content items.

[0260] exist Figure 9G At this point, the computer system 900 detects verbal input 905g (e.g., "What should I look at again?"). For example... Figure 9H As illustrated, in response to the detection of verbal input 905g, computer system 900 creates category 956 (e.g., a category corresponding to a song content item). Category 956 includes a fifth content item 914 and a sixth content item 916. Furthermore, computer system 900 shrinks the display of the fifth content item 914 and the sixth content item 916 to make room for content corresponding to the previously interacted content. Figure 9H As illustrated, in response to the detection of verbal input 905g, computer system 900 invokes and displays content from the previous interaction. (For example, previous content corresponding to the verbal ...

Claims

1. A method, the method comprising: At the computer system communicating with the mobile component: When the computer system is in a first location in the environment via the mobile component, the occurrence of the first interaction is detected; and In response to the detection of the first interaction: Based on the determination that the first interaction is a first type of interaction, the user moves via the mobile component to a second location in the environment that is different from the first location in the environment; as well as Based on the determination that the first interaction is a second type of interaction different from the first type of interaction, the movement to the second location via the moving component is abandoned.

2. The method according to claim 1, wherein, Before detecting the first interaction, a first portion of the computer system faces a first direction, and the system is moved to a second position so that the first portion faces a second direction different from the first direction.

3. The method according to any one of claims 1 to 2, wherein when it is determined that the number of users participating in the first interaction is higher than a threshold, the first interaction is the first type of interaction, and wherein when it is determined that the number of people participating in the first interaction is lower than the threshold, the first interaction is the second type of interaction.

4. The method according to any one of claims 1 to 3, wherein at least a portion of the first interaction is directed to the computer system.

5. The method according to any one of claims 1 to 4, wherein at least a portion of the first interaction does not refer to the computer system.

6. The method according to any one of claims 1 to 5, further comprising: In response to the detection of the first interaction: Based on the determination that the first interaction is the second type of interaction, the user moves via the mobile component to a third location that is different from the first location and the second location.

7. The method according to any one of claims 1 to 6, wherein, Before detecting the first interaction, the second part of the computer system is facing a third party, and the method further includes: In response to the detection of the first interaction: Based on the determination that the first interaction is the second type of interaction, the second part of the computer system continues to be directed towards the third party.

8. The method according to any one of claims 1 to 7, wherein, Before detecting the first interaction, the third part of the computer system faces the fourth direction, and the method further includes: In response to the detection of the first interaction: Based on the determination that the first interaction is the second type of interaction, the system moves via the mobile component to a fourth location different from the first location, while continuing to orient the third part of the computer system toward the fourth direction.

9. The method according to any one of claims 1 to 7, wherein, Before detecting the first interaction, the fourth part of the computer system faces the fifth direction, and the method further includes: In response to the detection of the first interaction: Based on the determination that the first interaction is the second type of interaction, the movement of the computer system via the mobile component is abandoned, while the fourth part of the computer system continues to face the fifth direction.

10. The method according to any one of claims 1 to 9, further comprising: In response to the detection of the first interaction: Based on the determination that the first interaction is a third type of interaction different from the first type of interaction and the second type of interaction, the user moves to the second location in the environment via the mobile component.

11. The method of any one of claims 1 to 10, wherein when it is determined that the first interaction includes a first type of dialogue, the first interaction is an interaction of the first type, and wherein when it is determined that the first interaction includes a second type of dialogue different from the first type of dialogue, the first interaction is an interaction of the second type.

12. The method according to any one of claims 1 to 11, wherein: Before detecting the first interaction and while the computer system is in the first location, the fifth part of the computer system faces the first user who is currently communicating. and After moving to the second location in response to detecting the first interaction and determining that the first interaction is the first type of interaction, the fifth part of the computer system faces a second user different from the first user when the computer system is in the second location.

13. The method of claim 12, wherein the computer system communicating with one or more input devices that detect the occurrence of the first interaction includes receiving input from the first user and the second user via the one or more input devices.

14. The method of any one of claims 1 to 13, wherein detecting the occurrence of the first interaction includes receiving an indication that a third user is not in communication.

15. The method of claim 14, wherein detecting the occurrence of the first interaction includes detecting that a fourth user, different from the third user, is communicating.

16. The method of any one of claims 1 to 15, wherein when the computer system is in the first position, the computer system is in a first tilt position, and wherein moving to the second position in the environment via the moving component includes tilting from the first tilt position to a second tilt position different from the first tilt position via the moving component.

17. The method of any one of claims 1 to 16, wherein when the computer system is in the first position, the computer system is in a first rotational position, and wherein moving to the second position in the environment via the moving component includes rotating from the first rotational position to a second rotational position different from the first rotational position via the moving component.

18. The method of any one of claims 1 to 17, wherein the first positioning includes a first lateral positioning, and wherein the second positioning in the environment via the moving component includes moving from the first lateral positioning to a second lateral positioning different from the first lateral positioning via the moving component.

19. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a mobile component, the one or more programs including instructions for performing the method according to any one of claims 1 to 18.

20. A computer system that communicates with a mobile component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 18.

21. A computer system that communicates with a mobile component, the computer system comprising: Components for performing the method according to any one of claims 1 to 18.

22. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a mobile component, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 18.

23. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a mobile component, the one or more programs including instructions for: When the computer system is in a first location in the environment via the mobile component, the occurrence of the first interaction is detected; and In response to the detection of the first interaction: Based on the determination that the first interaction is a first type of interaction, the user moves via the mobile component to a second location in the environment that is different from the first location in the environment; and Based on the determination that the first interaction is a second type of interaction different from the first type of interaction, the movement to the second location via the moving component is abandoned.

24. A computer system that communicates with a mobile component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: When the computer system is in a first location in the environment via the mobile component, the occurrence of the first interaction is detected; and In response to the detection of the first interaction: Based on the determination that the first interaction is a first type of interaction, the user moves via the mobile component to a second location in the environment that is different from the first location in the environment; as well as Based on the determination that the first interaction is a second type of interaction different from the first type of interaction, the movement to the second location via the moving component is abandoned.

25. A computer system that communicates with a mobile component, the computer system comprising: The component is used to detect the occurrence of a first interaction when the computer system is in a first position in the environment via the mobile component; as well as In response to the detection of the first interaction: The component is used for: moving via the mobile component to a second location in the environment that is different from the first location in the environment, based on determining that the first interaction is a first type of interaction; and The component is used for: abandoning movement to the second location via the moving component based on determining that the first interaction is a second type of interaction different from the first type of interaction.

26. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a mobile component, the one or more programs comprising instructions for: When the computer system is in a first location in the environment via the mobile component, the occurrence of the first interaction is detected; and In response to the detection of the first interaction: Based on the determination that the first interaction is a first type of interaction, the user moves via the mobile component to a second location in the environment that is different from the first location in the environment; and Based on the determination that the first interaction is a second type of interaction different from the first type of interaction, the movement to the second location via the moving component is abandoned.

27. A method, the method comprising: At the computer system that communicates with the display components and microphone: When the user interface is displayed via the display component, the first voice input is detected via the microphone; In response to detecting the first voice input, a first group of one or more words corresponding to the first voice input are displayed via the display component in a first manner; While displaying the first group of one or more words corresponding to the first voice input, the second voice input is detected via the microphone; as well as In response to the detection of the second voice input: Based on the determination that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, the new word corresponding to the second voice input is displayed in the first manner via the display component together with the display of the first group of one or more words; as well as Based on the determination that the second voice input includes the new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, a second group of one or more words including the new word corresponding to the second voice input is displayed via the display component in the first manner, while the display of the first group of one or more words in the first manner is stopped, wherein the second group of one or more words is different from the first group of one or more words.

28. The method of claim 27, wherein the new word is a first new word, and the method further comprises: While displaying one or more words in the second group that include the first new word corresponding to the second voice input, a third voice input is detected via the microphone; as well as In response to the detection of the third voice input: Based on the determination that the third voice input includes a second new word that is different from the first new word, and that the second new word corresponding to the third voice input should be added to the first group of one or more words, the second new word corresponding to the third voice input is displayed via the display component together with the first group of one or more words.

29. The method according to any one of claims 27 to 28, wherein the new word is a third new word, the method further comprising: When displaying one or more words in the second group that include the third new word corresponding to the second voice input, a fourth voice input is detected via the microphone; as well as In response to the detection of the fourth voice input: Based on the determination that the fourth voice input includes a fourth new word that is different from the third new word and that the fourth new word corresponding to the fourth voice input should be added to the second group of one or more words, the fourth new word corresponding to the fourth voice input is displayed via the display component together with the display of the second group of one or more words.

30. The method according to any one of claims 27 to 29, wherein the new word is a fifth new word, and the method further comprises: When displaying one or more words in the second group that include the fifth new word corresponding to the second voice input, the fifth voice input is detected via the microphone; as well as In response to the detection of the fifth voice input: Based on the determination that the fifth voice input includes a sixth new word that is different from the fifth new word and that the sixth new word corresponding to the fifth voice input should not be added to the second group of one or more words, the third group of one or more words including the sixth new word corresponding to the fifth voice input is displayed via the display component, while the display of the second group of one or more words in the first manner is stopped, wherein the third group of one or more words is different from the second group of one or more words.

31. The method according to any one of claims 27 to 30, wherein the new word is a seventh new word, and the method further comprises: When displaying one or more words in the second group that include the seventh new word corresponding to the second voice input, the sixth voice input is detected via the microphone; as well as In response to the detection of the sixth voice input: Based on the determination that the sixth voice input includes an eighth new word that is different from the seventh new word and that the eighth new word corresponding to the sixth voice input should not be added to the corresponding group of one or more words, the display of the seventh new word corresponding to the sixth voice input is abandoned via the display component.

32. The method according to claim 31, further comprising: In response to the detection of the sixth voice input: Based on the determination that the sixth voice input includes the eighth new word and that the eighth new word should not be added to the word list, the second group of one or more words continues to be displayed via the display component in the first manner.

33. The method according to any one of claims 27 to 32, wherein the second speech input comprises a phrase containing the new word.

34. The method according to any one of claims 27 to 33, wherein the new word is a ninth new word, and the method further comprises: In response to the detection of the second voice input: Based on the determination that the second voice input includes a tenth new word that is different from the ninth new word, the tenth new word corresponding to the second voice input should be added to the first group of one or more words. The second voice input includes the ninth new word, and the ninth new word corresponding to the second voice input should be added to the first group of one or more words. The ninth new word corresponding to the second voice input and the tenth new word corresponding to the second voice input are displayed concurrently in the first manner via the display component together with the display of the first group of one or more words.

35. The method of claim 34, wherein the second speech input includes an eleventh new word between the ninth and tenth new words in the second speech input, the method further comprising: In response to detecting the second voice input, the display of the eleventh new word corresponding to the second voice input is abandoned via the display component.

36. The method according to any one of claims 27 to 35, wherein the new word is the twelfth new word, and the method further comprises: In response to the detection of the second voice input: Based on the determination that the second voice input includes a thirteenth new word that is different from the twelfth new word, and that the thirteenth new word corresponding to the second voice input should not be added to the first group of one or more words, and that the twelfth new word corresponding to the second voice input should not be added to the first group of one or more words, the thirteenth new word corresponding to the second voice input and the twelfth new word corresponding to the second voice input are concurrently displayed as part of the second group of one or more words via the display component.

37. The method of claim 36, wherein the second speech input includes a fourteenth new word different from the thirteenth and twelfth new words in the second speech input, the method further comprising: In response to detecting the second voice input, the display of the fourteenth new word corresponding to the second voice input is abandoned via the display component.

38. The method of any one of claims 27 to 37, wherein the second voice input does not include an explicit instruction to add the new word to a particular set of one or more words.

39. The method according to any one of claims 27 to 38, further comprising: In response to the detection of the second voice input: Based on the determination that the second voice input includes a fifteenth new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, a first group of one or more instructions corresponding to the first group of one or more words are displayed, and the fifteenth new word corresponding to the second voice input is displayed in the first manner together with the display of the first group of one or more words; as well as Based on the determination that the second voice input includes the fifteenth new word corresponding to the second voice input and that the fifteenth new word corresponding to the second voice input should not be added to the first group of one or more words, a second group of instructions, different from the first group of instructions, is displayed corresponding to the second group of one or more words, while the new word corresponding to the second voice input is displayed in the first manner.

40. The method according to any one of claims 27 to 39, wherein the first group of one or more words is displayed in a first arrangement, and wherein the second group of one or more words is displayed in a second arrangement different from the first arrangement.

41. The method according to any one of claims 27 to 40, wherein: Displaying the first group of one or more words includes displaying a first group of one or more media representations corresponding to the first group of one or more words; and Displaying the second group of one or more words includes displaying a second group of one or more media representations corresponding to the second group of one or more words, wherein the second group of one or more media representations is different from the first group of one or more media representations.

42. The method according to any one of claims 27 to 41, wherein determining that the new word corresponding to the second speech input should be added to the first group of one or more words includes determining that the new word is a keyword in the second speech input.

43. The method of claim 42, wherein determining whether the new word is a keyword in the second speech input comprises: Based on the determination that the current context is the first context, the new word is determined to be the keyword. as well as Based on the determination that the current context is a second context different from the first context, it is determined that the new word is not the keyword.

44. The method of any one of claims 27 to 43, wherein determining whether the new word corresponding to the second speech input should be added to the first group of one or more words comprises: Based on the determination that the new word is contextually relevant to the first group of one or more words, it is determined that the new word should be added to the first group of one or more words; as well as Based on the determination that the new word is not relevant to the context of the first group of one or more words, it is determined that the new word should not be added to the first group of one or more words. When detecting the second voice input: At the first moment, detect the first part of the second voice input; In response to detecting the first portion of the second voice input, words corresponding to the first portion of the second voice input are displayed via the display component together with the first group of one or more words; At the second time, a second part of the second voice input that is different from the first part is detected; as well as In response to detecting the second part of the second voice input, a word corresponding to the second part of the second voice input is displayed via a display component together with one or more words from the first group.

45. The method of claim 45, wherein: Based on the determination that the second voice input has a first speed, the first time and the second time are separated by a first time interval; and It is determined that the second voice input has a second speed different from the first speed, and the first time and the second time are separated by a second time interval different from the first time interval.

46. ​​The method according to any one of claims 45 to 46, wherein: Based on the determination that the first part of the second voice input has a first group of one or more characteristics, the word corresponding to the first part of the second voice input is a first size; and Based on the determination that the first part of the second speech input has a second group of one or more characteristics that are different from the first group of one or more characteristics, the word corresponding to the first part of the second speech input is a second size that is different from the first size.

47. The method according to any one of claims 45 to 47, wherein: Based on determining that the word corresponding to the first part of the second speech input has a first relevance score relative to the first group of one or more words, the word corresponding to the first part of the second speech input is displayed at a first location relative to the first group of one or more words; and Based on the determination that the word corresponding to the first part of the second speech input has a second relevance score that is different from the first relevance score relative to the first group of one or more words, the word corresponding to the first part of the second speech input is displayed at a second location that is different from the first location relative to the first group of one or more words.

48. The method of any one of claims 27 to 48, wherein stopping the display of the first group of one or more words in the first manner includes removing the display of the first group of one or more words.

49. The method of any one of claims 27 to 48, wherein stopping the display of the first group of one or more words in the first manner includes displaying the first group of words in a second manner different from the first manner via the display component.

50. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and a microphone, the one or more programs including instructions for performing the method according to any one of claims 27 to 49.

51. A computer system communicating with a display component and a microphone, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 27 to 49.

52. A computer system communicating with a display component and a microphone, the computer system comprising: Components for performing the method according to any one of claims 27 to 49.

53. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and a microphone, the one or more programs comprising instructions for performing the method according to any one of claims 27 to 49.

54. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and a microphone, the one or more programs comprising instructions for: When the user interface is displayed via the display component, the first voice input is detected via the microphone; In response to detecting the first voice input, a first group of one or more words corresponding to the first voice input are displayed via the display component in a first manner; While displaying the first group of one or more words corresponding to the first voice input, the second voice input is detected via the microphone; as well as In response to the detection of the second voice input: Based on the determination that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, the new word corresponding to the second voice input is displayed in the first manner via the display component together with the display of the first group of one or more words; as well as Based on the determination that the second voice input includes the new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, a second group of one or more words including the new word corresponding to the second voice input is displayed via the display component in the first manner, while the display of the first group of one or more words in the first manner is stopped, wherein the second group of one or more words is different from the first group of one or more words.

55. A computer system communicating with a display component and a microphone, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: When the user interface is displayed via the display component, the first voice input is detected via the microphone; In response to detecting the first voice input, a first group of one or more words corresponding to the first voice input are displayed via the display component in a first manner; While displaying the first group of one or more words corresponding to the first voice input, the second voice input is detected via the microphone; as well as In response to the detection of the second voice input: Based on the determination that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, the new word corresponding to the second voice input is displayed in the first manner via the display component together with the display of the first group of one or more words; as well as Based on the determination that the second voice input includes the new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, a second group of one or more words including the new word corresponding to the second voice input is displayed via the display component in the first manner, while the display of the first group of one or more words in the first manner is stopped, wherein the second group of one or more words is different from the first group of one or more words.

56. A computer system communicating with a display component and a microphone, the computer system comprising: The component is used for detecting first voice input via the microphone when the user interface is displayed via the display component; The component is used for: in response to detecting the first voice input, displaying, via the display component, a first group of one or more words corresponding to the first voice input in a first manner; The component is used for: detecting a second voice input via the microphone when displaying one or more words corresponding to the first voice input; and In response to the detection of the second voice input: The component is used for: determining that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, and displaying the new word corresponding to the second voice input in the first manner via the display component together with the display of the first group of one or more words; and The component is used to: based on determining that the second voice input includes the new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, display a second group of one or more words including the new word corresponding to the second voice input in the first manner via the display component, while stopping the display of the first group of one or more words in the first manner, wherein the second group of one or more words is different from the first group of one or more words.

57. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display component and a microphone, the one or more programs comprising instructions for: When the user interface is displayed via the display component, the first voice input is detected via the microphone; In response to detecting the first voice input, a first group of one or more words corresponding to the first voice input are displayed via the display component in a first manner; While displaying the first group of one or more words corresponding to the first voice input, the second voice input is detected via the microphone; as well as In response to the detection of the second voice input: Based on the determination that the second voice input includes a new word and that the new word corresponding to the second voice input should be added to the first group of one or more words, the new word corresponding to the second voice input is displayed in the first manner via the display component together with the display of the first group of one or more words; as well as Based on the determination that the second voice input includes the new word corresponding to the second voice input and that the new word corresponding to the second voice input should not be added to the first group of one or more words, a second group of one or more words including the new word corresponding to the second voice input is displayed via the display component in the first manner, while the display of the first group of one or more words in the first manner is stopped, wherein the second group of one or more words is different from the first group of one or more words.

58. A method, the method comprising: At a computer system that communicates with a display component and one or more input devices: Input corresponding to the user is detected via the one or more input devices; as well as In conjunction with detecting input corresponding to the user, displaying a representation of a first portion of content related to the input and a representation of a second portion of content related to the input via the display component includes: Based on the determination that the first part of the content is in a first category of the content and the second part of the content is in the first category of the content, the representation of the first part of the content and the representation of the second part of the content are visually grouped. as well as Based on the determination that the first part of the content is in the first category of the content and the second part of the content is in a second category of the content that is different from the first category of the content, the visual grouping of the representation of the first part of the content and the representation of the second part of the content is abandoned.

59. The method of claim 58, wherein the input is verbal input.

60. The method according to any one of claims 58 to 59, the method further comprising: When displaying the representation of the first portion of the content and the representation of the second portion of the content, a representation of the third portion of the content is displayed via the display component, wherein the representation of the third portion of the content differs from the representation of the first portion of the content and the representation of the second portion of the content, wherein displaying the representation of the third portion of the content includes: Based on the determination that the first part of the content is in the first category of the content, the second part of the content is in the first category of the content, and the third part of the content is in the first category of the content, the representation of the first part of the content, the representation of the second part of the content, and the representation of the third part of the content are visually grouped. Based on the determination that the first part of the content belongs to the first category of the content, the second part of the content belongs to the second category of the content, and the third part of the content belongs to the first category of the content, the representation of the first part of the content and the representation of the third part of the content are visually grouped, while the representation of the second part of the content and the representation of the third part of the content are not visually grouped; and the representation of the first part of the content and the representation of the second part of the content are not visually grouped; Based on the determination that the first part of the content is in the first category of the content, the second part of the content is in the second category of the content, and the third part of the content is in the second category of the content, the representation of the second part of the content and the representation of the third part of the content are visually grouped, but not visually grouped, and not visually grouped, the representation of the first part of the content and the representation of the third part of the content.

61. The method according to any one of claims 58 to 60, the method further comprising: While displaying the representation of the first portion of the content and the representation of the second portion of the content, and when the representation of the first portion of the content is not visually grouped with the representation of the second portion of the content, a representation of the fourth portion of the content is displayed via the display component, wherein the representation of the fourth portion of the content differs from the representation of the first portion of the content and the representation of the second portion of the content, wherein displaying the representation of the fourth portion of the content includes: Based on the fact that the fourth part of the content and the first part of the content are in the same category of the content, the representation of the fourth part of the content and the representation of the first part of the content are visually grouped. Based on the fact that the fourth part of the content and the second category of the content belong to the same category of content, the representation of the fourth part of the content and the representation of the second part of the content are visually grouped; and Based on the fact that the fourth part of the content is determined to be in a different category from the first part and the second part of the content: Abandoning the visual grouping of the representation of the fourth part of the content and the representation of the first part of the content; and The visual grouping of the fourth part of the content and the second part of the content is abandoned.

62. The method according to any one of claims 58 to 61, wherein the representation of the second part is a first representation of the second part, the method further comprising: Before visually grouping the representation of the first part of the content and the first representation of the second part of the content, a second representation of the second part of the content that is not visually grouped with the representation of the first part of the content is displayed via the display component.

63. The method of claim 62, wherein the second representation of the second portion of the content displayed which is not visually grouped with the representation of the first portion of the content comprises displaying the second representation of the second portion of the content via the display component which does not overlap with and is not superimposed on the user interface elements.

64. The method according to any one of claims 62 to 63, wherein the second representation of the second portion is a first size, wherein the first representation of the second portion is a second size smaller than the first size, the method further comprising: After the second representation of the second part is initially displayed, an animation is displayed via the display component that transforms the second representation of the second part into the first representation of the second part by shrinking the second representation of the second part.

65. The method of any one of claims 62 to 64, wherein the second representation of the second portion is initially displayed at the first position, and wherein the first representation of the second portion is displayed at a second position different from the first position, the method further comprising: After the second representation of the second part is initially displayed at the first position, an animation is displayed via the display component that transforms the second representation of the second part into the first representation of the second part by moving the second representation of the second part toward the first position.

66. The method according to any one of claims 58 to 65, wherein the input is a first input, wherein the user is a first user, and the method further comprises: When displaying the representation of the first portion of visually grouped content and the representation of the second portion of content, a second input corresponding to the second user is detected via the one or more input devices; as well as In response to detecting the second input corresponding to the second user and determining that the second input meets the first set of criteria: Stop displaying the representation of the first portion of the visually grouped content and the representation of the second portion of the content via the display component; as well as The content corresponding to the second input is displayed via the display component.

67. The method of claim 66, further comprising: Based on the fact that the first part of the content is in the first category of the content and the second part of the content is in the first category of the content, and after stopping the display of the representation of the first part of the visually grouped content and the representation of the second part of the content, a third input corresponding to the first category of the content is detected via the one or more input devices; as well as In response to the detection of the third input, the representation of the first portion of the visually grouped content and the representation of the second portion of the content are displayed.

68. The method of claim 66, further comprising: Based on the fact that the first part of the content is in the first category of the content and the second part of the content is in the first category of the content, and after stopping the display of the representation of the first part of the content and the representation of the second part of the content, a fourth input corresponding to a third category of content that is different from the first category of the content is detected via the one or more input devices; as well as In response to the detection of the fourth input, the representation of the first portion of the visually grouped content and the representation of the second portion of the content are displayed via the display component.

69. The method according to claims 58 to 68, wherein the user is a second user, and the method further comprises: When the representation of the first part of the visually grouped content overlaps with the representation of the second part of the content, a fifth input corresponding to a third user is detected via the one or more input devices; as well as In response to the detection of the fifth input, content corresponding to the fifth input is displayed via the display component, while simultaneously displaying the representation of the first portion of the visually grouped content and the representation of the second portion of the content.

70. The method according to any one of claims 58 to 69, the method further comprising: When outputting audio content and displaying the representation of the first part of the content and the representation of the second part of the content, displaying a representation of a fifth part of the content that differs from the representation of the first part of the content and the representation of the second part of the content via the display component includes: Based on the determination that the first part of the content is in the first category of the content and the fifth part of the content is in the first category of the content, the representation of the fifth part of the content and the representation of the first part of the content are visually grouped. Based on the determination that the second part of the content is in the first category of the content and the fifth part of the content is in the first category of the content, the representation of the fifth part of the content and the representation of the second part of the content are visually grouped. Based on the determination that the first part of the content belongs to the first category of the content and the fifth part of the content belongs to a third category of the content that is different from the first category of the content, the representation of the fifth part of the content is displayed without visually grouping the representation of the first part of the content and the representation of the fifth part of the content; and Based on the determination that the second part of the content is in the first category of the content and the fifth part of the content is in the third category of the content, the representation of the fifth part of the content is displayed, without visually grouping the representation of the second part of the content and the representation of the fifth part of the content.

71. The method according to any one of claims 58 to 70, wherein the representation of the first portion of the displayed content and the representation of the second portion of the content comprise: Based on the determination that the first part of the content is in the first category of the content and the second part of the content is in the second category of the content, the representation of the first part of the content is displayed via the display component in a way that is not visually grouped with the representation of the second part of the content.

72. The method according to claim 71, further comprising: When displaying the representation of the first part of the content and the representation of the second part of the content: Based on the determination that the first part of the content and the sixth part of the content are in the same category of the content, a representation of the sixth part of the visually grouped content and the representation of the first part of the content are displayed via the display component, wherein the representation of the sixth part of the content is different from the representation of the first part of the content and the representation of the second part of the content; as well as Based on the determination that the second part of the content and the sixth part of the content are in the same category of the content, the representation of the sixth part of the visually grouped content and the representation of the second part of the content are displayed via the display component.

73. The method according to any one of claims 58 to 72, further comprising: Based on the detected input corresponding to the user, a seventh representation of the content is displayed via the display component that is not visually grouped with user interface elements, wherein the seventh representation of the content is different from the representation of the first part and the representation of the second part.

74. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 58 to 73.

75. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 58 to 73.

76. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 58 to 73.

77. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 58 to 73.

78. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for: Input corresponding to the user is detected via the one or more input devices; as well as In conjunction with detecting input corresponding to the user, displaying a representation of a first portion of content related to the input and a representation of a second portion of content related to the input via the display component includes: Based on the determination that the first part of the content is in a first category of the content and the second part of the content is in the first category of the content, the representation of the first part of the content and the representation of the second part of the content are visually grouped. as well as Based on the determination that the first part of the content is in the first category of the content and the second part of the content is in a second category of the content that is different from the first category of the content, the visual grouping of the representation of the first part of the content and the representation of the second part of the content is abandoned.

79. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: Input corresponding to the user is detected via the one or more input devices; as well as In conjunction with detecting input corresponding to the user, displaying a representation of a first portion of content related to the input and a representation of a second portion of content related to the input via the display component includes: Based on the determination that the first part of the content is in a first category of the content and the second part of the content is in the first category of the content, the representation of the first part of the content and the representation of the second part of the content are visually grouped. as well as Based on the determination that the first part of the content is in the first category of the content and the second part of the content is in a second category of the content that is different from the first category of the content, the visual grouping of the representation of the first part of the content and the representation of the second part of the content is abandoned.

80. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for: detecting input corresponding to the user via the one or more input devices; as well as In conjunction with detecting input corresponding to the user, displaying a representation of a first portion of content related to the input and a representation of a second portion of content related to the input via the display component includes: The component is used to: visually group the representation of the first part of the content and the representation of the second part of the content based on determining that the first part of the content is in a first category of the content and the second part of the content is in the first category of the content; and The component is used to: abandon the visual grouping of the representation of the first part of the content and the representation of the second part of the content based on the determination that the first part of the content is in the first category of the content and the second part of the content is in a second category of the content that is different from the first category of the content.

81. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for: Input corresponding to the user is detected via the one or more input devices; as well as In conjunction with detecting input corresponding to the user, displaying a representation of a first portion of content related to the input and a representation of a second portion of content related to the input via the display component includes: Based on the determination that the first part of the content is in a first category of the content and the second part of the content is in the first category of the content, the representation of the first part of the content and the representation of the second part of the content are visually grouped. as well as Based on the determination that the first part of the content is in the first category of the content and the second part of the content is in a second category of the content that is different from the first category of the content, the visual grouping of the representation of the first part of the content and the representation of the second part of the content is abandoned.

82. A method, the method comprising: At a computer system that communicates with a display component and one or more input devices: Detect a request corresponding to a previous interaction via the one or more input devices; as well as In response to detecting the request corresponding to the previous interaction, a user interface is displayed via the display component, the user interface including: A first representation of the first application corresponding to the previous interaction; A first representation of a first response to the request, wherein the first response originates from the previous interaction; and A second representation of a second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response is different from the second representation of the second response.

83. The method of claim 82, wherein the first representation of the first response and the second representation of the second response are visually grouped together.

84. The method according to claim 83, further comprising: While displaying the first representation of the first response and the second representation of the second response, a third representation of a third response to the request is displayed via the display component, wherein the third representation of the third response is not visually grouped with the first representation of the first response and the second representation of the second response, and wherein the third response is different from the first response and the second response.

85. The method according to any one of claims 82 to 84, the method further comprising: If no one or more inputs are detected after the first representation of the first response is displayed and after the representation of the first response is displayed, a fourth representation of a fourth response to the request is displayed via the display component, wherein the fourth response is from the previous interaction and wherein the fourth representation of the fourth response is different from the first representation of the first response.

86. The method according to claim 85, further comprising: In conjunction with the fourth representation of the fourth response, the display of the first representation of the first response is stopped.

87. The method according to any one of claims 85 to 86, wherein: The first response and the second response are included in a set of responses; The first representation of the first response and the second representation of the second response are included in the representation of the set of responses; The representations of the set of responses are visually grouped together before the fourth representation of the fourth response is displayed; and The method further includes: Based on the fourth representation of the fourth response and the determination that the content has been output exceeding a threshold amount for the set of responses, the display of the representation of the set of responses is stopped.

88. The method according to any one of claims 82 to 87, the method further comprising: When displaying the first representation of the first application corresponding to the previous interaction, a first input pointing to the first representation of the first application is detected; as well as In response to detecting the first input pointing to the first representation of the first application, a first application user interface corresponding to the first application is displayed via the display component.

89. The method according to any one of claims 82 to 88, the method further comprising: In response to detecting the request corresponding to the previous interaction, a second representation of a second application corresponding to the previous interaction is displayed via the display component, wherein the second application is different from the first application, and wherein the second representation of the second application is displayed concurrently with the first representation of the first application.

90. The method according to any one of claims 82 to 89, the method further comprising: When displaying the second representation of the second response to the request, a second input pointing to the second representation of the second response to the request is detected; as well as In response to detecting a second input pointing to a second representation of the second response to the request, a fifth representation of the second response to the request is displayed, wherein the fifth representation of the second response to the request is different from the second representation of the second response to the request.

91. The method according to any one of claims 82 to 90, the method further comprising: When displaying the second representation of the second response to the request, a third input pointing to the representation of the second response to the request is detected; as well as In response to detecting the third input pointing to the representation of the second response to the request, audio content corresponding to the second response is output via one or more output devices.

92. The method according to any one of claims 82 to 91, wherein the request corresponding to the previous interaction is an audible request.

93. The method of claim 92, wherein the request corresponding to the previous interaction does not include a first explicit indication of displaying the user interface.

94. The method of claim 92, wherein the request corresponding to the previous interaction includes a second explicit indication of displaying the user interface.

95. The method according to any one of claims 82 to 94, the method further comprising: When displaying the first representation of the first response to the request, output the second content corresponding to the first response; as well as When outputting the content corresponding to the first response: Based on the determination that one or more inputs have been detected when the output corresponds to the second content of the first response, a sixth representation of the first response is displayed via the display component without displaying a corresponding representation of the second response, wherein the sixth representation of the first response is different from the first representation of the first response. If it is determined that one or more of the set of inputs have not been detected when outputting content corresponding to the first response, the sixth representation of the first response is abandoned.

96. The method according to claim 95, further comprising: When the second content corresponding to the first response is output and based on the determination that one or more inputs have been detected while displaying the first representation of the first response, the display of the second representation of the second response is stopped.

97. The method according to any one of claims 95 to 96, further comprising: While outputting the second content corresponding to the first response and based on the determination that one or more of the set of inputs have not been detected when displaying the first representation of the first response to the request, the second representation of the second response continues to be displayed.

98. The method according to any one of claims 82 to 97, the method further comprising: When displaying the first representation of the first response, output the third content corresponding to the first response; as well as When outputting the third content corresponding to the first response: Based on the determination that no second group of one or more inputs has been detected when the third content corresponding to the first response is output, the fourth content corresponding to the second response is output; as well as If it is determined that one or more of the second group of inputs have been detected when the output corresponds to the content of the first response, the output corresponding to the second response is abandoned.

99. The method according to claim 98, wherein: The first response includes a first portion of the first response and a second portion of the first response; The first representation of the first response includes the first portion of the response; The third content corresponding to the first response includes the content displayed in the first representation of the first response and the content related to the sub-response corresponding to the first response, wherein the sub-response is the second part of the first response that is not displayed on the user interface.

100. The method of any one of claims 82 to 99, wherein the first representation showing the first response comprises: Based on the determination that the request corresponding to the previous interaction is a second type of interaction, the first representation of the first response is displayed in a second location different from the first location.

101. The method according to any one of claims 82 to 100, wherein the first representation of the first application is not visually grouped with the first representation of the first response to the request.

102. The method according to any one of claims 82 to 101, wherein the first representation of the first response and the first representation of the first application overlap each other.

103. The method according to any one of claims 82 to 102, the method further comprising: When displaying the first representation of the first response to the request, output the fourth content corresponding to the first response; While displaying the first representation of the first response to the request and outputting the fourth content corresponding to the first response, a fourth input pointing to the first representation of the first response is detected; as well as In response to detecting the fourth input pointing to the first representation of the first response: Stop outputting the fourth content corresponding to the first response; as well as The output corresponds to a fifth content of the first response, wherein the fourth content corresponding to the first response is different from the fifth content corresponding to the first response.

104. The method according to any one of claims 82 to 103, the method further comprising: When displaying the first representation of the first application corresponding to the previous interaction, a fifth input pointing to the first representation of the first application is detected; as well as In response to detecting a fifth input pointing to the first representation of the first application, an operation corresponding to the first application is performed.

105. The method according to claim 104, further comprising: In response to the detection of the fifth input pointing to the first representation of the first application, one or more of the first representation of the first response and the second representation of the second response continue to be displayed.

106. The method according to any one of claims 104 to 105, the method further comprising: In response to detecting a fifth input pointing to the first representation of the first application, the display of one or more of the first representation of the first response and the second representation of the second response is stopped.

107. The method according to claim 106, further comprising: In response to detecting the fifth input pointing to the first representation of the first application, a second application user interface corresponding to the first application is displayed via the display component.

108. The method of claim 107, wherein the second application user interface corresponding to the first application is displayed concurrently with one or more responses of the first representation of the first response.

109. The method according to any one of claims 107 to 108, the method further comprising: When displaying the user interface of the second application corresponding to the first application, a sixth input is detected; as well as In response to the detection of the sixth input: Stop displaying the user interface of the second application corresponding to the first application; as well as Displayed concurrently via the display component: The first representation of the first application corresponding to the previous interaction; The first representation of the first response to the request, wherein the first response is derived from the previous interaction; and The second representation of the second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response is different from the second representation of the second response.

110. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 82 to 109.

111. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 82 to 109.

112. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 82 to 109.

113. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 82 to 109.

114. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for: Detect a request corresponding to a previous interaction via the one or more input devices; as well as In response to detecting the request corresponding to the previous interaction, a user interface is displayed via the display component, the user interface including: A first representation of the first application corresponding to the previous interaction; A first representation of a first response to the request, wherein the first response originates from the previous interaction; and A second representation of a second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response is different from the second representation of the second response.

115. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: Detect a request corresponding to a previous interaction via the one or more input devices; as well as In response to detecting the request corresponding to the previous interaction, a user interface is displayed via the display component, the user interface including: A first representation of the first application corresponding to the previous interaction; A first representation of a first response to the request, wherein the first response originates from the previous interaction; and A second representation of a second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response is different from the second representation of the second response.

116. A computer system communicating with a display component and one or more input devices, the computer system comprising: The component is used to detect a request corresponding to a previous interaction via the one or more input devices; as well as In response to detecting the request corresponding to the previous interaction, a user interface is displayed via the display component, the user interface including: For the following component: a first representation of the first application corresponding to the previously interacted; Components for: a first representation of a first response to the request, wherein the first response originates from the previous interaction; and A component for: a second representation of a second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response is different from the second representation of the second response.

117. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for: Detect a request corresponding to a previous interaction via the one or more input devices; as well as In response to detecting the request corresponding to the previous interaction, a user interface is displayed via the display component, the user interface including: A first representation of the first application corresponding to the previous interaction; A first representation of a first response to the request, wherein the first response originates from the previous interaction; and A second representation of a second response to the request, wherein the second response is derived from the previous interaction, and wherein the first representation of the first response is different from the second representation of the second response.

118. A method, the method comprising: At the computer system communicating with the display components: The detection corresponds to the first request in the previous interaction; as well as In response to detecting the request corresponding to the previous interaction: Based on the determination that the request does not correspond to new content, a first summary of the previous interaction is displayed via the display component, the first summary including one or more representations corresponding to the previous interaction that are in a first orientation relative to one or more representations corresponding to the previous interaction; as well as Based on the determination that the request includes new content, a second summary of the previous interaction is displayed via the display component, the second summary including one or more representations of the first group corresponding to the previous interaction in a second orientation relative to the second group or more representations, wherein the second orientation is different from the first orientation.

119. The method of claim 118, wherein the computer system communicates with one or more output devices, the method further comprising: In response to detecting the first request corresponding to the previous interaction: Based on the determination that the request does not correspond to new content, a first audio corresponding to a portion of the previous interaction is output via the one or more output devices; as well as Based on the determination that the request includes new content, a second audio corresponding to the portion of the previous interaction is output.

120. The method of claim 119, wherein the first audio is the same as the second audio.

121. The method according to claim 119, wherein: The first audio is different from the second audio; The first audio includes content corresponding to the portion of the previous interaction; The second audio includes a second amount of content corresponding to the previous interaction, which is different from the first amount of content corresponding to the portion of the previous interaction; and The first quantity is less than the second quantity.

122. The method according to any one of claims 119 to 121, the method further comprising: In response to detecting the first request corresponding to the previous interaction: If it is determined that the request does not correspond to new content, the output of a third audio corresponding to the new content via the one or more output devices is abandoned; as well as Based on the determination that the request includes new content, a third audio corresponding to the new content is output via the one or more output devices.

123. The method according to any one of claims 118 to 122, wherein the first group of one or more representations includes representations that are visually grouped together.

124. The method according to any one of claims 118 to 123, wherein the first group of one or more representations and the second group of one or more representations are not visually grouped together.

125. The method according to any one of claims 118 to 124, the method further comprising: In response to detecting the request corresponding to the previous interaction: Based on the determination that the request corresponds to new content, a third group or more representations corresponding to the new content are displayed via the display; as well as Based on the determination that the request does not correspond to new content, the display of the third group or more representations corresponding to the new content via the display is abandoned.

126. The method of claim 125, wherein the new content is a first new content, the method further comprising: After displaying the third group of one or more representations corresponding to the first new content, a second request corresponding to the previous interaction is detected; as well as In response to detecting the second request corresponding to the previous interaction: Based on the determination that the second request includes second new content, a fourth group of one or more representations corresponding to the second new content are displayed in a third orientation via the display component; as well as Based on the determination that the second request does not correspond to the second new content, the third group of one or more representations corresponding to the first new content continues to be displayed via the display component, without displaying the fourth group of one or more representations corresponding to the second new content, wherein the third group of one or more representations is in a fourth orientation different from the third orientation.

127. The method of any one of claims 125 to 126, wherein displaying the third group of one or more representations corresponding to the new content includes including the display of the third group of one or more representations in the display of one or more representations in the first representation corresponding to the previous interaction.

128. The method of any one of claims 125 to 127, wherein displaying the third group of one or more representations corresponding to the new content includes visually grouping the third group of one or more representations together with one or more of the first representation corresponding to the previous interaction and the second representation corresponding to the previous interaction.

129. The method of any one of claims 125 to 128, wherein displaying the third group of one or more representations corresponding to the new content does not include including the display of the third group of one or more representations in the display of one or more of the first group of one or more representations corresponding to the previous interaction and the second group of one or more representations corresponding to the previous interaction.

130. The method of any one of claims 125 to 129, wherein displaying the third group of one or more representations corresponding to the new content does not include visually grouping the third group of one or more representations together with one or more of the first representation corresponding to the previous interaction and the second representation corresponding to the previous interaction.

131. The method according to any one of claims 118 to 130, the method further comprising: In response to detecting the first request corresponding to the previous interaction, a third audio corresponding to one or more representations of the first group is output; as well as After outputting the third audio content corresponding to one or more representations of the first group, output the fourth audio corresponding to one or more representations of the second group.

132. The method according to claim 131, further comprising: After outputting the initial portion of the third audio corresponding to one or more representations of the first group and before outputting the final portion of the fourth audio corresponding to one or more representations of the second group, the display of the first group or more representations is stopped; as well as After outputting the fourth audio corresponding to one or more representations of the second group, the display of one or more representations of the second group is stopped.

133. The method according to claim 131, further comprising: After outputting the audio content corresponding to one or more representations of the first group, continue displaying the one or more representations of the first group; as well as After outputting the audio content corresponding to one or more representations of the second group, the display of one or more representations of the second group continues.

134. The method according to any one of claims 118 to 133, the method further comprising: In response to detecting the first request corresponding to the previous interaction and determining that the request includes a new topic, a fifth audio corresponding to the new topic is output; as well as After outputting the audio corresponding to the new topic, output the sixth audio content corresponding to one or more representations of the first group.

135. The method according to any one of claims 118 to 134, wherein: Based on the determination that the previous interaction corresponds to a first topic, the first group of one or more representations has a first number of one or more representations; and Based on the determination that the previous interaction corresponds to a second topic different from the first topic, the first group of one or more representations has a second number of one or more representations different from the first number of one or more representations.

136. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, the one or more programs including instructions for performing the method according to any one of claims 118 to 135.

137. A computer system communicating with a display component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 118 to 135.

138. A computer system communicating with a display component, the computer system comprising: Components for performing the method according to any one of claims 118 to 135.

139. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, the one or more programs comprising instructions for performing the method according to any one of claims 118 to 135.

140. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, the one or more programs comprising instructions for: The detection corresponds to the first request in the previous interaction; and In response to detecting the request corresponding to the previous interaction: Based on the determination that the request does not correspond to new content, a first summary of the previous interaction is displayed via the display component, the first summary including one or more representations corresponding to the previous interaction that are in a first orientation relative to a second or more representations corresponding to the previous interaction; and Based on the determination that the request includes new content, a second summary of the previous interaction is displayed via the display component, the second summary including one or more representations of the first group corresponding to the previous interaction in a second orientation relative to the second group or more representations, wherein the second orientation is different from the first orientation.

141. A computer system communicating with a display component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: The detection corresponds to the first request in the previous interaction; and In response to detecting the request corresponding to the previous interaction: Based on the determination that the request does not correspond to new content, a first summary of the previous interaction is displayed via the display component, the first summary including one or more representations corresponding to the previous interaction that are in a first orientation relative to one or more representations corresponding to the previous interaction; as well as Based on the determination that the request includes new content, a second summary of the previous interaction is displayed via the display component, the second summary including one or more representations of the first group corresponding to the previous interaction in a second orientation relative to the second group or more representations, wherein the second orientation is different from the first orientation.

142. A computer system communicating with a display component, the computer system comprising: Used for the following component: detecting the first request corresponding to the previous interaction; as well as In response to detecting the request corresponding to the previous interaction: The component is used for: displaying a first summary of the previous interaction via the display component based on determining that the request does not correspond to new content, the first summary including one or more representations corresponding to the previous interaction in a first orientation relative to one or more representations corresponding to the previous interaction; and The component is used to: display a second summary of the previous interaction via the display component, based on the determination that the request includes new content, the second summary including one or more representations of the first group corresponding to the previous interaction in a second orientation relative to the second group of one or more representations, wherein the second orientation is different from the first orientation.

143. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display component, the one or more programs comprising instructions for: The detection corresponds to the first request in the previous interaction; and In response to detecting the request corresponding to the previous interaction: Based on the determination that the request does not correspond to new content, a first summary of the previous interaction is displayed via the display component, the first summary including one or more representations corresponding to the previous interaction that are in a first orientation relative to a second or more representations corresponding to the previous interaction; and Based on the determination that the request includes new content, a second summary of the previous interaction is displayed via the display component, the second summary including one or more representations of the first group corresponding to the previous interaction in a second orientation relative to the second group or more representations, wherein the second orientation is different from the first orientation.

144. A method, the method comprising: At a computer system that communicates with one or more output devices, including a display component, and one or more input devices: The display component displays visual content including one or more items in a first group, one or more items in a second group that are different from the items in the first group, and an incarnation of the items in the first group that is closer to the items in the second group than the items in the second group. When displaying visual content comprising the first group of items, the second group of items, and the avatar that is closer to the first group of items than the second group of items, the content corresponding to the first group of items is output via the one or more output devices; When the output corresponds to the content of the first group item and the avatar that is closer to the first group item than the second group item is displayed, the detection will output the content corresponding to the second group item; as well as In response to detecting that the content corresponding to the second group item will be output, the avatar positioned closer to the second group item than the first group item is displayed via the display component.

145. The method according to claim 144, further comprising: In response to detecting that content corresponding to the second group item will be output, the display of the avatar is changed so that the avatar visually points to the second group item.

146. The method according to claim 145, further comprising: When the avatar is displayed such that it visually points to the second group of items, and based on a predetermined time period, the display of the avatar is changed such that it visually moves away from the direction of the second group of items.

147. The method of claim 146, wherein after changing the display of the avatar so that the avatar visually deviates from the direction of the second set of items, the avatar visually points to the first user detected in the first detection field of the computer system.

148. The method of claim 146, wherein after changing the display of the avatar so that the avatar visually deviates from the direction of the second set of items, the avatar points to the first physical environment.

149. The method of claim 146, wherein after changing the display of the avatar so that the avatar visually deviates from the direction pointed to by the second group of items, the avatar no longer points to the first group of items.

150. The method according to any one of claims 145 to 149, the method further comprising: In response to detecting that the content corresponding to the second group item will be output, the display of the avatar is changed so that the avatar changes from visually pointing to the first group item to visually not pointing to the first group item.

151. The method according to any one of claims 144 to 150, the method further comprising: In response to detecting that the content corresponding to the second group item will be output, the display of the avatar is changed so that the avatar changes from visually pointing to the second user detected in the second detection field to visually not pointing to the second user detected in the second detection field.

152. The method according to any one of claims 144 to 151, the method further comprising: In response to detecting that the content corresponding to the second group item will be output, the display of the avatar is changed so that the avatar changes from visually pointing to the second physical environment to visually not pointing to the second physical environment.

153. The method according to any one of claims 144 to 152, the method further comprising: In response to detecting that the content corresponding to the second set of items will be output, the avatar is moved from the first location to a second location different from the first location; as well as When the avatar is moved from the first location to the second location, the avatar is displayed via the display component as if visually pointing to the second group of items.

154. The method according to any one of claims 144 to 153, the method further comprising: In response to detecting that the content corresponding to the second set of items will be output, the avatar is moved from the third location to a fourth location different from the third location; as well as After the avatar is moved from the third location to a fourth location different from the third location, the avatar is displayed via the display component as if visually pointing to the second group of items.

155. The method according to claims 144 to 154, the method further comprising: In response to detecting that the content corresponding to the second set of items will be output, the avatar is moved from the fifth location to a sixth location that is different from the third location; as well as Before moving the avatar from the fifth position to a sixth position different from the third position, the avatar is displayed via the display component as pointing to the second group of items.

156. The method of any one of claims 144 to 155, wherein the visual content includes a third group of items different from the first group of items and the second group of items, and wherein the third group of items is visually grouped together with at least one of the first group of items and the second group of items.

157. The method according to any one of claims 144 to 156, wherein: When the avatar is displayed closer to the first group of items than the second group of items, the avatar is displayed on the first portion of the first group of items; When the avatar is displayed closer to the second group of items than the first group of items, the avatar is displayed on the second portion of the second group of items; and The first part is different from the second part.

158. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system including one or more output devices comprising a display component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 144 to 157.

159. A computer system communicating with one or more output devices including a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 144 to 157.

160. A computer system communicating with one or more output devices including a display component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 144 to 157.

161. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system including one or more output devices including display components and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 144 to 157.

162. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system including one or more output devices comprising a display component and one or more input devices, the one or more programs comprising instructions for: The display component displays visual content including one or more items in a first group, one or more items in a second group that are different from the items in the first group, and an incarnation of the items in the first group that is closer to the items in the second group than the items in the second group. When displaying visual content comprising the first group of items, the second group of items, and the avatar that is closer to the first group of items than the second group of items, the content corresponding to the first group of items is output via the one or more output devices; When the output corresponds to the content of the first group item and the avatar that is closer to the first group item than the second group item is displayed, the detection will output the content corresponding to the second group item; as well as In response to detecting that the content corresponding to the second group item will be output, the avatar positioned closer to the second group item than the first group item is displayed via the display component.

163. A computer system communicating with one or more output devices including a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: The display component displays visual content including one or more items in a first group, one or more items in a second group that are different from the items in the first group, and an incarnation of the items in the first group that is closer to the items in the second group than the items in the second group. When displaying visual content comprising the first group of items, the second group of items, and the avatar that is closer to the first group of items than the second group of items, the content corresponding to the first group of items is output via the one or more output devices; When the output corresponds to the content of the first group item and the avatar that is closer to the first group item than the second group item is displayed, the detection will output the content corresponding to the second group item; as well as In response to detecting that the content corresponding to the second group item will be output, the avatar positioned closer to the second group item than the first group item is displayed via the display component.

164. A computer system communicating with one or more output devices including a display component and one or more input devices, the computer system comprising: The component is used to display visual content comprising one or more items in a first group, one or more items in a second group that are different from the items in the first group, and an incarnation of the items in the first group that is closer to the items in the second group than the items in the second group. The component is used to output content corresponding to the first group of items via the one or more output devices when displaying visual content including the first group of items, the second group of items, and the avatar that is closer to the first group of items than the second group of items; The component is used for the following: when outputting the content corresponding to the first group item and displaying the avatar that is closer to the first group item than the second group item, it detects that the content corresponding to the second group item will be output; and The component is used to: in response to detecting that content corresponding to the second group item will be output, display the avatar located closer to the second group item than the first group item via the display component.

165. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system including one or more output devices comprising a display component and one or more input devices, the one or more programs comprising instructions for: The display component displays visual content including one or more items in a first group, one or more items in a second group that are different from the items in the first group, and an incarnation of the items in the first group that is closer to the items in the second group than the items in the second group. When displaying visual content comprising the first group of items, the second group of items, and the avatar that is closer to the first group of items than the second group of items, the content corresponding to the first group of items is output via the one or more output devices; When the output corresponds to the content of the first group item and the avatar that is closer to the first group item than the second group item is displayed, the detection will output the content corresponding to the second group item; as well as In response to detecting that the content corresponding to the second group item will be output, the avatar positioned closer to the second group item than the first group item is displayed via the display component.

166. A method, the method comprising: At a computer system that communicates with a display component and one or more input devices: When a first user interface object is displayed via the display component, input corresponding to the topic is detected via the one or more input devices; as well as In response to detecting the input corresponding to the topic: Based on the determination that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input, the decision is not to increase the size of the first user interface object; as well as The size of the first user interface object is increased based on the determination that the corresponding portion of the input is associated with a confidence level higher than the threshold corresponding to the input.

167. The method of claim 166, wherein the input is an audible input.

168. The method according to any one of claims 166 to 167, the method further comprising: In response to detecting the input corresponding to the topic: The size of the first user interface object is reduced based on the determination that the corresponding portion of the input is associated with a confidence level below the threshold corresponding to the input.

169. The method according to any one of claims 166 to 167, wherein the first user interface object is displayed at a first size, the method further comprising: In response to detecting the input corresponding to the topic: Based on the determination that the corresponding portion of the input is associated with a confidence level below the threshold corresponding to the input, the first user interface object continues to be displayed at the first size.

170. The method of any one of claims 166 to 169, wherein a second user interface object, different from the first user interface object, is displayed at a third size before detecting the input corresponding to the user, the method further comprising: In response to detecting the input corresponding to the topic: Based on determining that the corresponding portion of the input is associated with a confidence level higher than the threshold corresponding to the input, the size of the second user interface object is increased from a fourth size greater than the third size.

171. The method according to any one of claims 166 to 169, wherein a third user interface object, different from the first user interface object, is displayed at a fifth size, the method further comprising: In response to detecting the input corresponding to the topic: Based on the determination that the corresponding portion of the input is associated with a confidence level higher than the threshold corresponding to the input, the third user interface object continues to be displayed at the fifth size.

172. The method according to any one of claims 166 to 171, wherein the input is a first input, the method further comprising: Detect the second input corresponding to the topic; as well as In response to detecting a second input corresponding to the topic: Based on the determination that a corresponding portion of the second input is associated with a confidence level higher than the threshold corresponding to the input, the increase in the size of the first user interface object is abandoned; as well as Based on the determination that the corresponding portion of the second input is associated with a confidence level below the threshold corresponding to the input, the increase in the size of the first user interface object is abandoned.

173. The method according to any one of claims 166 to 172, the method further comprising: When the fourth interface object is displayed via the display component, a third input corresponding to the second theme is detected via the one or more input devices; as well as In response to the detection of the third input corresponding to the second topic: The size of the fourth user interface object is increased based on the determination that the third input corresponds to the fourth user interface object and that a corresponding portion of the third input corresponding to the fourth user interface object is associated with a confidence level higher than a second threshold for the portion corresponding to the third input; as well as Based on the determination that the third input does not correspond to the fourth user interface object and that the corresponding portion of the third input corresponding to the fourth user interface object is associated with a confidence level higher than the second threshold corresponding to the portion of the third input, the increase in the size of the fourth user interface object is abandoned.

174. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 166 to 173.

175. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 166 to 173.

176. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 166 to 173.

177. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 166 to 173.

178. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for: When a first user interface object is displayed via the display component, input corresponding to the topic is detected via the one or more input devices; and In response to detecting the input corresponding to the topic: Based on the determination that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input, the increase in the size of the first user interface object is abandoned; and The size of the first user interface object is increased based on the determination that the corresponding portion of the input is associated with a confidence level higher than the threshold corresponding to the input.

179. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: When a first user interface object is displayed via the display component, input corresponding to the topic is detected via the one or more input devices; and In response to detecting the input corresponding to the topic: Based on the determination that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input, the decision is not to increase the size of the first user interface object; as well as The size of the first user interface object is increased based on the determination that the corresponding portion of the input is associated with a confidence level higher than the threshold corresponding to the input.

180. A computer system communicating with a display component and one or more input devices, the computer system comprising: The component is used for: detecting input corresponding to the theme via the one or more input devices when a first user interface object is displayed via the display component; as well as In response to detecting the input corresponding to the topic: The component is used for: abandoning the increase of the size of the first user interface object based on determining that a corresponding part of the input is associated with a confidence level below a threshold corresponding to the input; and The component is used to increase the size of the first user interface object based on determining that the corresponding portion of the input is associated with a confidence level higher than the threshold corresponding to the input.

181. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for: When a first user interface object is displayed via the display component, input corresponding to the topic is detected via the one or more input devices; and In response to detecting the input corresponding to the topic: Based on the determination that a corresponding portion of the input is associated with a confidence level below a threshold corresponding to the input, the increase in the size of the first user interface object is abandoned; and The size of the first user interface object is increased based on the determination that the corresponding portion of the input is associated with a confidence level higher than the threshold corresponding to the input.

182. A method, the method comprising: At a computer system that communicates with one or more input devices and one or more output devices: Detect input corresponding to a request to view one or more previous interactions with the agent via the one or more input devices; and In response to detecting the input corresponding to the request to view one or more previous interactions with the agent: Based on the determination that one or more criteria of a first set are met, a first representation of the first previous interaction with the agent is output via the one or more output devices; as well as Based on the determination that one or more criteria of the second set are met, the first representation of the first previous interaction with the agent is abandoned when output via the one or more output devices.

183. The method of claim 182, wherein the agent is a virtual assistant.

184. The method according to any one of claims 182 to 183, wherein the first portion of the agent is executed on the computer system, wherein the second portion of the agent is executed on another computer system different from the computer system, and wherein the second portion is different from the first portion.

185. The method of claim 183, wherein the agent is configured to provide output using a large language model (LLM).

186. The method according to any one of claims 182 to 185, the method further comprising: In response to detecting the input corresponding to a request to view one or more previous interactions with the agent, and based on determining that the first set of one or more criteria are met, a second representation of a second previous interaction with the agent is output via the one or more output devices, wherein the second representation is different from the first representation.

187. The method of claim 186, wherein the first represents content of a first type, and wherein the second represents content of a second type different from content of the first type.

188. The method according to any one of claims 182 to 187, wherein the first representation corresponds to a first application, and wherein the second representation corresponds to a second application different from the first application.

189. The method according to any one of claims 182 to 188, wherein the first representation corresponds to a first media item, and wherein the second representation corresponds to a second media item different from the first media item.

190. The method according to any one of claims 182 to 189, the method further comprising: In response to detecting the input corresponding to a request to view one or more previous interactions with the agent, and based on determining that one or more of the first set of criteria are met, a third representation of a third previous interaction with the agent is output via the one or more output devices, wherein the third representation is different from the first representation and the second representation, wherein the second representation is visually grouped with the first representation, wherein the second representation is visually not grouped with the third representation, and wherein the third representation is visually not grouped with the first representation.

191. The method of any one of claims 182 to 190, wherein at least a portion of the content of the first representation was not included in the first prior interaction.

192. The method of any one of claims 182 to 191, wherein the first prior interaction arises from a conversation with the agent.

193. The method according to any one of claims 182 to 192, wherein the first representation includes a suggestion provided by the agent during the first prior interaction.

194. The method of any one of claims 182 to 193, wherein the first prior interaction includes natural language input from the user.

195. The method of any one of claims 182 to 194, wherein the first representation includes visual input provided during the first prior interaction.

196. The method according to any one of claims 182 to 195, wherein the first representation includes a graphic image.

197. The method of any one of claims 182 to 196, wherein the first representation includes text from the first prior interaction.

198. The method of any one of claims 182 to 197, wherein the first representation includes a summary of the first prior interaction, and wherein the summary was not provided during the first prior interaction.

199. The method according to any one of claims 182 to 198, the method further comprising: When the first representation of the first previous interaction is output via the one or more output devices, an input corresponding to the selection of the first representation is detected via the input device; as well as In response to detecting the input corresponding to the selection of the first representation, additional content corresponding to the first previous interaction is output via the one or more output devices.

200. The method of any one of claims 182 to 199, wherein the input corresponding to the request to view one or more previous interactions with the agent is an implicit request to view one or more previous interactions with the agent.

201. The method of any one of claims 182 to 200, wherein the input corresponding to the request to view one or more previous interactions with the agent is an explicit request to view one or more previous interactions with the agent.

202. The method of any one of claims 182 to 201, wherein the input corresponding to the request to view one or more previous interactions with the agent includes an indication of time, and wherein the first set of one or more criteria includes criteria satisfied when the first previous interaction corresponding to the indication of time is performed.

203. The method of any one of claims 182 to 202, wherein the input corresponding to the request to view one or more previous interactions with the agent includes an indication of a topic, and wherein the first set of one or more criteria includes criteria satisfied when the first previous interaction includes content corresponding to the topic.

204. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices, the one or more programs including instructions for performing the method according to any one of claims 182 to 203.

205. A computer system that communicates with one or more input devices and one or more output devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 182 to 203.

206. A computer system that communicates with one or more input devices and one or more output devices, the computer system comprising: Components for performing the method according to any one of claims 182 to 203.

207. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices, the one or more programs comprising instructions for performing the method according to any one of claims 182 to 203.

208. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices, the one or more programs comprising instructions for: Detect input corresponding to a request to view one or more previous interactions with the agent via the one or more input devices; and In response to detecting the input corresponding to the request to view one or more previous interactions with the agent: Based on the determination that one or more criteria of a first set are met, a first representation of the first previous interaction with the agent is output via the one or more output devices; and Based on the determination that one or more criteria of the second set are met, the first representation of the first previous interaction with the agent is abandoned when output via the one or more output devices.

209. A computer system that communicates with one or more input devices and one or more output devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: Detect input corresponding to a request to view one or more previous interactions with the agent via the one or more input devices; and In response to detecting the input corresponding to the request to view one or more previous interactions with the agent: Based on the determination that one or more criteria of a first set are met, a first representation of the first previous interaction with the agent is output via the one or more output devices; as well as Based on the determination that one or more criteria of the second set are met, the first representation of the first previous interaction with the agent is abandoned when output via the one or more output devices.

210. A computer system that communicates with one or more input devices and one or more output devices, the computer system comprising: Components for: detecting input corresponding to a request to view one or more previous interactions with the agent via the one or more input devices; as well as In response to detecting the input corresponding to the request to view one or more previous interactions with the agent: Based on the determination that one or more criteria of the first set are met, the component is used to: output a first representation of a first previous interaction with the agent via the one or more output devices; and Based on the determination that one or more criteria of the second set are met, the component is used to: abandon the first representation of the first previous interaction with the agent output via the one or more output devices.

211. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices, the one or more programs comprising instructions for: Detect input corresponding to a request to view one or more previous interactions with the agent via the one or more input devices; and In response to detecting the input corresponding to the request to view one or more previous interactions with the agent: Based on the determination that one or more criteria of a first set are met, a first representation of the first previous interaction with the agent is output via the one or more output devices; and Based on the determination that one or more criteria of the second set are met, the first representation of the first previous interaction with the agent is abandoned when output via the one or more output devices.

212. A method, the method comprising: At a computer system that communicates with a display component and one or more input devices: Detect a request to display an animation via the one or more input devices; In response to the detected request to display the animation, playback of the animation is initiated via the display component; When the animation is played back and the overlay is displayed at the first position via the display component, it is detected that an object in the animation will be within the distance of the first position when the first frame of the animation is displayed; as well as In response to detecting that the object in the animation will be within the distance of the first position when the first frame of the animation is displayed, the overlay is displayed via the display component at a second position different from the first position, wherein the second position is selected after the playback of the animation is initiated.

213. The method of claim 212, wherein the overlay includes a representation of a face.

214. The method according to any one of claims 212 to 213, further comprising: The overlay is displayed via the display component before the playback of the animation is initiated.

215. The method according to any one of claims 212 to 214, the method further comprising: When the animation is played back, the overlay is displayed at a third position via the display component: Based on the determination that the second frame of the animation includes the first content, the overlay is displayed with a first appearance; as well as The overlay is displayed with a second appearance that is different from the first appearance, based on the determination that the second frame of the animation includes second content that is different from the first content.

216. The method according to any one of claims 212 to 215, the method further comprising: When the animation is played back, the overlay is displayed at the fourth position via the display component: Based on the determination that the user in the first environment is in a first state, the overlay is displayed with a third appearance; as well as Based on the determination that the user in the first environment is in a second state different from the first state, the overlay is displayed with a fourth appearance different from the third appearance.

217. The method according to any one of claims 212 to 216, the method further comprising: When the animation is played back, the overlay is displayed at the fifth position via the display component: Based on the determination that the second environment is in the first state, the stack is displayed with a fifth appearance; as well as Based on the determination that the second environment is in a second state different from the first state, the stack is displayed in a sixth appearance different from the fifth appearance.

218. The method according to any one of claims 212 to 217, wherein the distance is a first distance, the method further comprising: When playing back the animation, after the overlay is displayed at the second location and when the overlay is displayed at the sixth location, it is detected that the second object in the animation will be displayed within the second distance at the sixth location when the third frame of the animation, which is different from the first frame and the second frame, is displayed; as well as In response to detecting that the second object in the animation will be within the second distance of the sixth position when the third frame of the animation is displayed, the overlay is displayed at a seventh position different from the sixth position via the display component.

219. The method according to any one of claims 212 to 218, the method further comprising: When the animation is played back and the overlay is displayed: Based on the determination that the animation includes third content, perform one or more operations of the first group to move the overlay to the eleventh position; as well as Based on the determination that the animation includes a fourth content different from the third content, a second group of one or more operations are performed to move the overlay to the eleventh position, wherein the second group of one or more operations is different from the first group of one or more operations.

220. The method according to any one of claims 212 to 219, wherein the animation includes video.

221. The method according to any one of claims 212 to 220, wherein the animation includes previously recorded content.

222. The method of claim 221, wherein the animation is generated prior to detecting the request to display the animation.

223. The method according to any one of claims 212 to 222, wherein the animation is a first animation, the method further comprising: A request to display a second animation different from the first animation is detected via the one or more input devices; In response to the detection of the request to display the second animation, playback of the second animation is initiated via one or more output devices; as well as When the second animation is played back and the overlay is displayed via the display component: Based on the determination that the second animation is a first type of animation, the overlay is moved to a new position via the display component; as well as Based on the determination that the second animation is a second type of animation different from the first type of animation, the movement of the overlay to the new position via the display component is abandoned.

224. The method according to any one of claims 212 to 223, wherein the computer system does not detect input when playing back the animation.

225. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 212 to 224.

226. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 212 to 224.

227. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 212 to 224.

228. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 212 to 224.

229. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for: Detect a request to display an animation via the one or more input devices; In response to the detected request to display the animation, playback of the animation is initiated via the display component; When the animation is played back and the overlay is displayed at the first position via the display component, it is detected that an object in the animation will be within the distance of the first position when the first frame of the animation is displayed; as well as In response to detecting that the object in the animation will be within the distance of the first position when the first frame of the animation is displayed, the overlay is displayed via the display component at a second position different from the first position, wherein the second position is selected after the playback of the animation is initiated.

230. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: Detect a request to display an animation via the one or more input devices; In response to the detected request to display the animation, playback of the animation is initiated via the display component; When the animation is played back and the overlay is displayed at the first position via the display component, it is detected that an object in the animation will be within the distance of the first position when the first frame of the animation is displayed; as well as In response to detecting that the object in the animation will be within the distance of the first position when the first frame of the animation is displayed, the overlay is displayed via the display component at a second position different from the first position, wherein the second position is selected after the playback of the animation is initiated.

231. A computer system communicating with a display component and one or more input devices, the computer system comprising: The component is used for: detecting a request to display an animation via the one or more input devices; The component is used for: in response to detecting the request to display the animation, initiating playback of the animation via the display component; The component is used for: when playing back the animation and displaying the overlay at a first position via the display component, detecting that an object in the animation will be within a distance of the first position when the first frame of the animation is displayed; and The component is used to: in response to detecting that the object in the animation will be within the distance of the first position when the first frame of the animation is displayed, display the overlay at a second position different from the first position via the display component, wherein the second position is selected after the playback of the animation is initiated.

232. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for: Detect a request to display an animation via the one or more input devices; In response to the detected request to display the animation, playback of the animation is initiated via the display component; When the animation is played back and the overlay is displayed at the first position via the display component, it is detected that an object in the animation will be within the distance of the first position when the first frame of the animation is displayed; as well as In response to detecting that the object in the animation will be within the distance of the first position when the first frame of the animation is displayed, the overlay is displayed via the display component at a second position different from the first position, wherein the second position is selected after the playback of the animation is initiated.