User interfaces and techniques for performing operations based on learned characteristics
By detecting changes in user movement across different contexts, the system automatically adjusts its operation methods and optimizes the user interface, solving the complex and inefficient interface problems of existing technologies and improving operational efficiency and battery life.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2024-09-25
- Publication Date
- 2026-07-31
AI Technical Summary
In the prior art, user interface operation methods are complex and inefficient, especially in battery-powered devices, wasting time and energy and resulting in a poor user experience.
By detecting changes in user movement across different contexts, the system automatically adjusts its operation, reduces unnecessary input actions, and optimizes the user interface to improve efficiency and save battery power.
It enables faster and more efficient user interface operation, reduces the cognitive burden on users, and extends battery life in battery-powered devices.
Smart Images

Figure CN122488945A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on September 25, 2024, with application number 202480062264.1 and invention title "User interface and technology for performing operations based on learned characteristics". Cross-reference of related applications
[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 541,826, filed September 30, 2023; U.S. Provisional Patent Application Serial No. 63 / 541,822, filed September 30, 2023; U.S. Provisional Patent Application Serial No. 63 / 541,825, filed September 30, 2023; and U.S. Provisional Patent Application Serial No. 63 / 541,830, filed September 30, 2023, the entire contents of which are incorporated herein by reference for all purposes. Background Technology
[0003] Users typically interact with the user interface of a computer system, providing commands through the interface to initiate various operations on the computer system. The computer system typically outputs content based on specific input. The computer system is generally used to perform operations by the user. Such operations can be performed in response to user-provided input. Electronic devices typically issue user confirmation. This confirmation indicates that a user has been detected. Summary of the Invention
[0004] Existing technologies for using electronic devices to perform operations based on learned characteristics are often cumbersome and inefficient. For example, some existing technologies use complex and time-consuming user interfaces that may include multiple buttons or keystrokes. Some existing technologies require more time than necessary, wasting user time and device energy. This latter consideration is particularly important in battery-powered devices.
[0005] Therefore, this technology provides electronic devices with faster and more efficient methods and interfaces for performing operations. Such methods and interfaces can optionally complement or replace other methods for performing operations. These methods and interfaces reduce the cognitive burden on users and result in more efficient human-machine interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging. Such methods and interfaces can complement or replace other methods for performing operations based on learned characteristics.
[0006] In some embodiments, a method is described that is executed at a computer system communicating with one or more input devices. In some embodiments, the method includes: detecting movement of a first user relative to a first context via one or more input devices; abandoning the execution of a representation of the first user's movement in response to detecting movement of the first user relative to the first context; detecting a transition to a second context after detecting movement of the first user relative to the first context; and executing the representation of the first user's movement in response to detecting the transition to the second context and based on determining that the second context corresponds to the first context.
[0007] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices. In some embodiments, the one or more programs include instructions for: detecting movement of a first user relative to a first context via one or more input devices; in response to detecting movement of the first user relative to the first context, relinquishing an indication of executing the first user's movement; after detecting movement of the first user relative to the first context, detecting a transition to a second context; and in response to detecting a transition to the second context and based on determining that the second context corresponds to the first context, executing an indication of the first user's movement.
[0008] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices. In some embodiments, the one or more programs include instructions for: detecting movement of a first user relative to a first context via one or more input devices; in response to detecting movement of the first user relative to the first context, relinquishing an indication to execute the movement of the first user; after detecting movement of the first user relative to the first context, detecting a transition to a second context; and in response to detecting a transition to the second context and based on determining that the second context corresponds to the first context, executing an indication to execute the movement of the first user.
[0009] In some embodiments, a computer system communicating with one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting movement of a first user relative to a first context via one or more input devices; abandoning an indication of executing the first user's movement in response to detecting movement of the first user relative to the first context; detecting a transition to a second context after detecting movement of the first user relative to the first context; and executing an indication of the first user's movement in response to detecting the transition to the second context and based on determining that the second context corresponds to the first context.
[0010] In some embodiments, a computer system communicating with one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting movement of a first user relative to a first context via one or more input devices; abandoning the execution of a representation of the first user's movement in response to detecting movement of the first user relative to the first context; detecting a transition to a second context after detecting movement of the first user relative to the first context; and executing the representation of the first user's movement in response to detecting the transition to the second context and based on determining that the second context corresponds to the first context.
[0011] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices. In some embodiments, the one or more programs include instructions for: detecting movement of a first user relative to a first context via one or more input devices; in response to detecting movement of the first user relative to the first context, relinquishing an indication to execute the movement of the first user; after detecting movement of the first user relative to the first context, detecting a transition to a second context; and in response to detecting a transition to the second context and based on determining that the second context corresponds to the first context, executing an indication to execute the movement of the first user.
[0012] In some embodiments, a method is described that is performed at a computer system communicating with one or more input devices. In some embodiments, the method includes: detecting a first gesture along with detecting a first input via one or more input devices; in response to detecting the first gesture along with detecting the first input: performing a first operation; and configuring the first operation to be performed even when the first input is not detected; after performing the first operation and after configuring the first operation to be performed even when the first input is not detected, detecting the first gesture without detecting the first input via one or more input devices; and in response to detecting the first gesture but not detecting the first input, performing the first operation again.
[0013] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first gesture along with detecting a first input via one or more input devices; performing a first operation in response to detecting the first gesture along with detecting the first input; configuring the first operation to be performed even when the first input is not detected; after performing the first operation and after configuring the first operation to be performed even when the first input is not detected, detecting the first gesture without detecting the first input via one or more input devices; and performing the first operation in response to detecting the first gesture but not detecting the first input.
[0014] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first gesture along with detecting a first input via one or more input devices; performing a first operation in response to detecting the first gesture along with detecting the first input; configuring the first operation to be performed even when the first input is not detected; after performing the first operation and after configuring the first operation to be performed even when the first input is not detected, detecting the first gesture without detecting the first input via one or more input devices; and performing the first operation in response to detecting the first gesture but not detecting the first input.
[0015] In some embodiments, a computer system that communicates with one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a first gesture via one or more input devices in conjunction with detecting a first input; in response to detecting the first gesture in conjunction with detecting the first input: performing a first operation; and configuring the first operation to be performed even when the first input is not detected; after performing the first operation and after configuring the first operation to be performed even when the first input is not detected, detecting the first gesture via one or more input devices without detecting the first input; and in response to detecting the first gesture but not detecting the first input, performing the first operation.
[0016] In some embodiments, a computer system that communicates with one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting a first gesture via one or more input devices in conjunction with detecting a first input; in response to detecting the first gesture in conjunction with detecting the first input: performing a first operation; and configuring the first operation to be performed even when the first input is not detected; after performing the first operation and after configuring the first operation to be performed even when the first input is not detected, detecting the first gesture via one or more input devices without detecting the first input; and in response to detecting the first gesture but not detecting the first input, performing the first operation.
[0017] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first gesture along with detecting a first input via one or more input devices; in response to detecting the first gesture along with detecting the first input: performing a first operation; and configuring the first operation to be performed even when the first input is not detected; after performing the first operation and after configuring the first operation to be performed even when the first input is not detected, detecting the first gesture without detecting the first input via one or more input devices; and in response to detecting the first gesture but not detecting the first input, performing the first operation.
[0018] In some embodiments, a method is described that is performed at a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the method includes: detecting input corresponding to a request from a user via one or more input devices; and in response to detecting input corresponding to a request from a user: based on determining that the computer system is operating in a first context and that the first context corresponds to a first set of one or more learned characteristics corresponding to the user, outputting audio content having the first set of one or more audio characteristics via one or more output devices; and based on determining that the computer system is operating in a second context different from the first context and that the second context corresponds to a second set of one or more learned characteristics different from the first set of one or more learned characteristics corresponding to the user, outputting audio content having a second set of one or more audio characteristics different from the first set of one or more audio characteristics via one or more output devices.
[0019] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a request from a user via one or more input devices; and in response to detecting input corresponding to a request from a user: outputting audio content having the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a first context and that the first context corresponds to a first set of one or more learned characteristics corresponding to the user; and outputting audio content having a second set of one or more audio characteristics different from the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a second context different from the first context and that the second context corresponds to a second set of one or more learned characteristics different from the first set of one or more learned characteristics corresponding to the user.
[0020] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a request from a user via one or more input devices; and in response to detecting input corresponding to a request from a user: outputting audio content having the first set of one or more audio characteristics via one or more output devices, based on determining that the computer system is operating in a first context and that the first context corresponds to a first set of one or more learned characteristics corresponding to the user; and outputting audio content having a second set of one or more audio characteristics different from the first set of one or more audio characteristics via one or more output devices, based on determining that the computer system is operating in a second context different from the first context and that the second context corresponds to a second set of one or more learned characteristics different from the first set of one or more learned characteristics corresponding to the user.
[0021] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a request from a user via one or more input devices; and in response to detecting input corresponding to a request from a user: outputting audio content having the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a first context and the first context corresponds to a first set of one or more learned characteristics corresponding to the user; and outputting audio content having a second set of one or more audio characteristics different from the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a second context different from the first context and the second context corresponds to a second set of one or more learned characteristics different from the first set of one or more learned characteristics corresponding to the user.
[0022] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting input corresponding to a request from a user via one or more input devices; and in response to detecting input corresponding to a request from a user: outputting audio content having the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a first context and the first context corresponds to a first set of one or more learned characteristics corresponding to the user; and outputting audio content having a second set of one or more audio characteristics different from the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a second context different from the first context and the second context corresponds to a second set of one or more learned characteristics different from the first set of one or more learned characteristics corresponding to the user.
[0023] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input corresponding to a request from a user via one or more input devices; and in response to detecting input corresponding to a request from a user: outputting audio content having the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a first context and that the first context corresponds to a first set of one or more learned characteristics corresponding to the user; and outputting audio content having a second set of one or more audio characteristics different from the first set of one or more audio characteristics via one or more output devices based on determining that the computer system is operating in a second context different from the first context and that the second context corresponds to a second set of one or more learned characteristics different from the first set of one or more learned characteristics corresponding to the user.
[0024] In some embodiments, a method is described that is executed at a computer system communicating with an audio generation component and one or more input devices. In some embodiments, the method includes: detecting a first input via one or more input devices; and after detecting the first input: based on determining that the first input corresponds to a first group of one or more audio characteristics, outputting first audio content having a second group of one or more audio characteristics via the audio generation component; and based on determining that the first input corresponds to a third group of one or more audio characteristics different from the first group of one or more audio characteristics, outputting second audio content having a fourth group of one or more audio characteristics different from the second group of one or more audio characteristics via the audio generation component.
[0025] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system in communication with an audio generation component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first input via one or more input devices; and after detecting the first input: outputting first audio content having a second set of one or more audio characteristics via the audio generation component based on determining that the first input corresponds to a first set of one or more audio characteristics; and outputting second audio content having a fourth set of one or more audio characteristics different from the second set of one or more audio characteristics via the audio generation component based on determining that the first input corresponds to a third set of one or more audio characteristics different from the first set of one or more audio characteristics.
[0026] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with an audio generation component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first input via one or more input devices; and, after detecting the first input: based on determining that the first input corresponds to a first set of one or more audio characteristics, outputting first audio content having a second set of one or more audio characteristics via the audio generation component; and based on determining that the first input corresponds to a third set of one or more audio characteristics different from the first set of one or more audio characteristics, outputting second audio content having a fourth set of one or more audio characteristics different from the second set of one or more audio characteristics via the audio generation component.
[0027] In some embodiments, a computer system communicating with an audio generation component and one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a first input via one or more input devices; and, after detecting the first input: based on determining that the first input corresponds to a first set of one or more audio characteristics, outputting first audio content having a second set of one or more audio characteristics via the audio generation component; and based on determining that the first input corresponds to a third set of one or more audio characteristics different from the first set of one or more audio characteristics, outputting second audio content having a fourth set of one or more audio characteristics different from the second set of one or more audio characteristics via the audio generation component.
[0028] In some embodiments, a computer system communicating with an audio generation component and one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting a first input via one or more input devices; and after detecting the first input: based on determining that the first input corresponds to a first group of one or more audio characteristics, outputting first audio content having a second group of one or more audio characteristics via the audio generation component; and based on determining that the first input corresponds to a third group of one or more audio characteristics different from the first group of one or more audio characteristics, outputting second audio content having a fourth group of one or more audio characteristics different from the second group of one or more audio characteristics via the audio generation component.
[0029] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with an audio generation component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first input via one or more input devices; and, after detecting the first input: based on determining that the first input corresponds to a first set of one or more audio characteristics, outputting first audio content having a second set of one or more audio characteristics via the audio generation component; and based on determining that the first input corresponds to a third set of one or more audio characteristics different from the first set of one or more audio characteristics, outputting second audio content having a fourth set of one or more audio characteristics different from the second set of one or more audio characteristics via the audio generation component.
[0030] In some embodiments, a method is described that is performed at a computer system having a mobile component and one or more input devices. In some embodiments, the method includes: detecting a request for performing an operation via one or more input devices; and in response to detecting the request for performing the operation: moving via the mobile component in a first manner prior to determining that a user in the environment should perform a first action in the environment before performing the operation; and abandoning the first manner of movement via the mobile component prior to determining that a user in the environment should perform a second action in the environment different from the first action before performing the operation.
[0031] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system having a movable component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request for performing an operation via one or more input devices; and in response to detecting the request for performing the operation: moving in a first manner via the movable component before performing the operation, based on determining that a user in the environment should perform a first action in the environment before performing the operation; and abandoning the first manner of movement via the movable component, based on determining that a user in the environment should perform a second action in the environment different from the first action before performing the operation.
[0032] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system having a movable component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request for performing an operation via one or more input devices; and in response to detecting the request for performing the operation: moving in a first manner via the movable component before performing the operation, based on determining that a user in the environment should perform a first action in the environment before performing the operation; and abandoning the first manner of movement via the movable component, based on determining that a user in the environment should perform a second action in the environment different from the first action before performing the operation.
[0033] In some embodiments, a computer system having a movable component and one or more input devices is described. In some embodiments, the computer system having the movable component and one or more input devices includes one or more processors and memory configured to execute one or more programs by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a request to perform an operation via the one or more input devices; and in response to detecting the request to perform the operation: moving via the movable component in a first manner before performing the operation, based on determining that a user in the environment should perform a first action in the environment before performing the operation; and abandoning the first manner of movement via the movable component, based on determining that a user in the environment should perform a second action in the environment different from the first action before performing the operation.
[0034] In some embodiments, a computer system having a mobile component and one or more input devices is described. In some embodiments, the computer system having the mobile component and one or more input devices includes components for performing each of the following steps: detecting a request to perform an operation via one or more input devices; and in response to detecting the request to perform an operation: moving via the mobile component in a first manner prior to performing the operation, based on determining that a user in the environment should perform a first action in the environment before performing the operation; and abandoning the first manner of movement via the mobile component prior to performing the operation, based on determining that a user in the environment should perform a second action in the environment different from the first action.
[0035] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system having a movable component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a request for performing an operation via one or more input devices; and in response to detecting the request for performing the operation: moving in a first manner via the movable component before performing the operation, based on determining that a user in the environment should perform a first action in the environment before performing the operation; and abandoning the first manner of movement via the movable component, based on determining that a user in the environment should perform a second action in the environment different from the first action before performing the operation.
[0036] In some embodiments, a method is described that is performed at a computer system having one or more input devices and one or more output devices. In some embodiments, the method includes: detecting a user's intent via one or more input devices without detecting an explicit instruction for performing an operation determined to be the user's intent; and in response to detecting the user's intent but not detecting an explicit instruction for performing an operation determined to be the detected intent: outputting a first set of one or more suggested actions for performing the operation via one or more output devices based on determining that the context corresponding to the intent is a first type of context; and abandoning the output of the first set of one or more suggested actions for performing the operation based on determining that the context corresponding to the intent is a second type of context different from the first type of context.
[0037] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system having one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting a user's intent via one or more input devices without detecting explicit instructions for performing an operation determined to be the user's intent; and in response to detecting the user's intent but not detecting the explicit instructions for performing an operation determined to be the detected intent: outputting a first set of one or more suggested actions for performing the operation via one or more output devices based on determining that the context corresponding to the intent is a first type of context; and abandoning the output of the first set of one or more suggested actions for performing the operation based on determining that the context corresponding to the intent is a second type of context different from the first type of context.
[0038] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system having one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting a user's intent via one or more input devices without detecting an explicit instruction for performing an operation determined to be the user's intent; and in response to detecting a user's intent but not detecting an explicit instruction for performing an operation determined to be the detected intent: outputting a first set of one or more suggested actions for performing the operation via one or more output devices based on determining that the context corresponding to the intent is a first type of context; and abandoning the output of the first set of one or more suggested actions for performing the operation based on determining that the context corresponding to the intent is a second type of context different from the first type of context.
[0039] In some embodiments, a computer system having one or more input devices and one or more output devices is described. In some embodiments, the computer system having one or more input devices and one or more output devices includes one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for performing the following operations: detecting a user's intent via one or more input devices without detecting explicit instructions for performing an operation determined to be the user's intent; and in response to detecting a user's intent but not detecting an explicit instruction for performing an operation determined to be the detected intent: outputting a first set of one or more suggested actions for performing the operation via one or more output devices based on determining that the context corresponding to the intent is a first type of context; and abandoning the output of the first set of one or more suggested actions for performing the operation based on determining that the context corresponding to the intent is a second type of context different from the first type of context.
[0040] In some embodiments, a computer system having one or more input devices and one or more output devices is described. In some embodiments, the computer system having one or more input devices and one or more output devices includes components for performing each of the following steps: detecting a user's intent via one or more input devices without detecting an explicit instruction to perform an operation determined to be the user's intent; and in response to detecting the user's intent but not detecting an explicit instruction to perform an operation determined to be the detected intent: outputting a first set of one or more suggested actions for performing the operation via one or more output devices based on determining that the context corresponding to the intent is a first type of context; and abandoning the output of the first set of one or more suggested actions for performing the operation based on determining that the context corresponding to the intent is a second type of context different from the first type of context.
[0041] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system having one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following: detecting a user's intent via one or more input devices without detecting an explicit instruction for performing an operation determined to be the user's intent; and in response to detecting the user's intent but not detecting the explicit instruction for performing an operation determined to be the detected intent: outputting a first set of one or more suggested actions for performing the operation via one or more output devices based on determining that the context corresponding to the intent is a first type of context; and abandoning the output of the first set of one or more suggested actions for performing the operation based on determining that the context corresponding to the intent is a second type of context different from the first type of context.
[0042] In some embodiments, a method is described that is performed at a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the method includes: detecting input from a first user via one or more input devices, wherein the input includes identification information of a set of one or more users; after detecting the input, detecting a second user different from the first user via one or more input devices; and in response to detecting the second user: outputting a first confirmation to the second user via one or more output devices based on determining that the second user corresponds to identification information of the set of one or more users; and outputting a second confirmation to the second user via one or more output devices based on determining that the second user does not correspond to identification information corresponding to the set of one or more users, wherein the second confirmation is different from the first confirmation.
[0043] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting input from a first user via one or more input devices, wherein the input includes identification information of a set of one or more users; detecting a second user different from the first user via one or more input devices after detecting the input; and in response to detecting the second user: outputting a first confirmation to the second user via one or more output devices based on determining that the second user corresponds to the identification information of the set of one or more users; and outputting a second confirmation to the second user via one or more output devices based on determining that the second user does not correspond to the identification information corresponding to the set of one or more users, wherein the second confirmation is different from the first confirmation.
[0044] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs executed by one or more processors of a computer system configured to communicate with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting input from a first user via one or more input devices, wherein the input includes identification information of a set of one or more users; detecting a second user different from the first user via one or more input devices after detecting the input; and in response to detecting the second user: outputting a first confirmation to the second user via one or more output devices based on determining that the second user corresponds to identification information of a set of one or more users; and outputting a second confirmation to the second user via one or more output devices based on determining that the second user does not correspond to identification information corresponding to a set of one or more users, wherein the second confirmation is different from the first confirmation.
[0045] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system communicating with one or more input devices and one or more output devices includes one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for performing the following operations: detecting input from a first user via one or more input devices, wherein the input includes identification information of a set of one or more users; detecting a second user different from the first user via one or more input devices after detecting the input; and in response to detecting the second user: outputting a first confirmation to the second user via one or more output devices based on determining that the second user corresponds to identification information of a set of one or more users; and outputting a second confirmation to the second user via one or more output devices based on determining that the second user does not correspond to identification information corresponding to a set of one or more users, wherein the second confirmation is different from the first confirmation.
[0046] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system communicating with one or more input devices and one or more output devices includes components for performing each of the following steps: detecting input from a first user via one or more input devices, wherein the input includes identification information of a set of one or more users; after detecting the input, detecting a second user different from the first user via one or more input devices; and in response to detecting the second user: outputting a first confirmation to the second user via one or more output devices based on determining that the second user corresponds to identification information of a set of one or more users; and outputting a second confirmation to the second user via one or more output devices based on determining that the second user does not correspond to identification information corresponding to a set of one or more users, wherein the second confirmation is different from the first confirmation.
[0047] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting input from a first user via one or more input devices, wherein the input includes identification information of a set of one or more users; detecting a second user different from the first user via one or more input devices after detecting the input; and in response to detecting the second user: outputting a first confirmation to the second user via one or more output devices based on determining that the second user corresponds to identification information of a set of one or more users; and outputting a second confirmation to the second user via one or more output devices based on determining that the second user does not correspond to identification information corresponding to a set of one or more users, wherein the second confirmation is different from the first confirmation.
[0048] In some embodiments, a method is described that is performed at a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the method includes: detecting a first set of one or more inputs via one or more input devices, the first set of one or more inputs including indications of a first context and indications of one or more movements in the first set; after detecting the first set of one or more inputs, detecting the presence of a corresponding context via one or more input devices; and in response to detecting the presence of a corresponding context: based on determining that the corresponding context is a first context, outputting a representation of the first set of one or more movements via one or more output devices; and based on determining that the corresponding context is a second context different from the first context, abandoning the output of the representation of the first set of one or more movements via one or more output devices.
[0049] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting a first set of one or more inputs via one or more input devices, the first set of one or more inputs including indications of a first context and indications of one or more movements in the first set; detecting the presence of a corresponding context via one or more input devices after detecting the first set of one or more inputs; and in response to detecting the presence of a corresponding context: outputting a representation of the first set of one or more movements via one or more output devices based on determining that the corresponding context is a first context; and abandoning the output of the representation of the first set of one or more movements via one or more output devices based on determining that the corresponding context is a second context different from the first context.
[0050] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting a first set of one or more inputs via one or more input devices, the first set of one or more inputs including indications of a first context and indications of one or more movements in the first set; detecting the presence of a corresponding context via one or more input devices after detecting the first set of one or more inputs; and in response to detecting the presence of a corresponding context: outputting a representation of the first set of one or more movements via one or more output devices based on determining that the corresponding context is a first context; and abandoning the output of the representation of the first set of one or more movements via one or more output devices based on determining that the corresponding context is a second context different from the first context.
[0051] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system communicating with one or more input devices and one or more output devices includes one or more processors and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for performing the following operations: detecting a first set of one or more inputs via one or more input devices, the first set of one or more inputs including indications to a first context and indications to the first set of one or more movements; after detecting the first set of one or more inputs, detecting the presence of a corresponding context via one or more input devices; and in response to detecting the presence of a corresponding context: based on determining that the corresponding context is a first context, outputting an indication of the first set of one or more movements via one or more output devices; and based on determining that the corresponding context is a second context different from the first context, abandoning the output of the indication of the first set of one or more movements via one or more output devices.
[0052] In some embodiments, a computer system communicating with one or more input devices and one or more output devices is described. In some embodiments, the computer system communicating with one or more input devices and one or more output devices includes components for performing each of the following steps: detecting a first set of one or more inputs via one or more input devices, the first set of one or more inputs including an indication of a first context and an indication of one or more movements in the first set; after detecting the first set of one or more inputs, detecting the presence of a corresponding context via one or more input devices; and in response to detecting the presence of a corresponding context: based on determining that the corresponding context is a first context, outputting a representation of one or more movements in the first set via one or more output devices; and based on determining that the corresponding context is a second context different from the first context, abandoning the output of the representation of one or more movements in the first set via one or more output devices.
[0053] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for performing the following operations: detecting a first set of one or more inputs via one or more input devices, the first set of one or more inputs including an indication of a first context and an indication of one or more movements; after detecting the first set of one or more inputs, detecting the presence of a corresponding context via one or more input devices; and in response to detecting the presence of a corresponding context: based on determining that the corresponding context is a first context, outputting a representation of the first set of one or more movements via one or more output devices; and based on determining that the corresponding context is a second context different from the first context, abandoning the output of the representation of the first set of one or more movements via one or more output devices.
[0054] In some embodiments, a method is described that is performed at a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the method includes: detecting input from a user via one or more input devices; and, after detecting input from the user and if no request to provide a suggestion is detected: outputting a first suggestion via one or more output devices, wherein the first suggestion is based on the input from the user, based on determining that the current context is a first context; and abandoning the output of the first suggestion via one or more output devices, based on determining that the current context is not the first context.
[0055] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input from a user via one or more input devices; and, upon detecting input from the user and in the absence of a request to provide a suggestion, outputting a first suggestion via one or more output devices, wherein the first suggestion is based on the input from the user, based on determining that the current context is a first context; and abandoning the output of the first suggestion via one or more output devices, based on determining that the current context is not the first context.
[0056] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input from a user via one or more input devices; and, upon detecting input from the user and in the absence of a request to provide a suggestion, outputting a first suggestion via one or more output devices, wherein the first suggestion is based on the input from the user, based on determining that the current context is a first context; and abandoning the output of the first suggestion via one or more output devices, based on determining that the current context is not the first context.
[0057] In some embodiments, a computer system that communicates with one or more input devices and one or more output devices is described. In some embodiments, the computer system includes: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting input from a user via one or more input devices; and, upon detecting input from the user and in the absence of a request to provide a suggestion, outputting a first suggestion via one or more output devices, wherein the first suggestion is based on the input from the user, based on determining that the current context is a first context; and abandoning the output of the first suggestion via one or more output devices, based on determining that the current context is not the first context.
[0058] In some embodiments, a computer system is described that communicates with one or more input devices and one or more output devices. In some embodiments, the computer system includes components for performing each of the following steps: detecting input from a user via one or more input devices; and, after detecting input from the user and in the absence of a request to provide a suggestion: outputting a first suggestion via one or more output devices, wherein the first suggestion is based on the input from the user, based on determining that the current context is a first context; and abandoning the output of the first suggestion via one or more output devices, based on determining that the current context is not the first context.
[0059] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with one or more input devices and one or more output devices. In some embodiments, the one or more programs include instructions for: detecting input from a user via one or more input devices; and, upon detecting input from the user and in the absence of a request to provide a suggestion: outputting a first suggestion via one or more output devices, wherein the first suggestion is based on the input from the user, based on determining that the current context is a first context; and abandoning the output of the first suggestion via one or more output devices, based on determining that the current context is not the first context.
[0060] Executable instructions for performing these functions may optionally be included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Attached Figure Description
[0061] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals in all the drawings indicate the corresponding parts.
[0062] Figure 1 This is a block diagram illustrating a computer system according to some implementation schemes.
[0063] Figures 2A to 2C These are illustrations of exemplary components and user interfaces of an electronic device according to some implementation schemes.
[0064] Figure 3 This is a block diagram illustrating exemplary components of a device according to some implementation schemes.
[0065] Figure 4 This is a functional diagram of an exemplary actuator device according to some implementation schemes.
[0066] Figure 5 This is a functional diagram of an exemplary agent system based on some implementation schemes.
[0067] Figures 6A to 6E An exemplary user interface for performing movement based on learned characteristics, according to some implementation schemes, is illustrated.
[0068] Figure 7 This is a flowchart illustrating a method for performing a movement representation based on learned characteristics, according to some implementation schemes.
[0069] Figures 8A to 8E An exemplary user interface for configuring operations to be performed based on learned characteristics, according to some implementation schemes, is illustrated.
[0070] Figure 9 This is a flowchart illustrating a method for configuring operations to be performed based on learned characteristics, according to some implementation schemes.
[0071] Figures 10A to 10C An exemplary user interface for automatically outputting content based on context and / or input characteristics, according to some implementation schemes, is illustrated.
[0072] Figure 11 This is a flowchart illustrating a method for automatically outputting audio content based on context, according to some implementation schemes.
[0073] Figure 12 This is a flowchart illustrating a method for automatically outputting audio with specific characteristics according to some implementation schemes.
[0074] Figures 13A to 13D An exemplary user interface for performing operations is illustrated according to some implementation schemes.
[0075] Figure 14 This is a flowchart illustrating a method for performing an operation according to some implementation schemes.
[0076] Figures 15A to 15C An exemplary user interface for predicting demand is illustrated according to some implementation schemes.
[0077] Figure 16 This is a flowchart illustrating a method for predicting demand based on some implementation schemes.
[0078] Figures 17A to 17D An exemplary user interface for performing a custom greeting is illustrated according to some implementation schemes.
[0079] Figure 18 This is a flowchart illustrating a method for performing a customized greeting according to some implementation schemes.
[0080] Figure 19 This is a flowchart illustrating a method for performing a customized greeting according to some implementation schemes. Detailed Implementation
[0081] The following description illustrates exemplary methods, components, parameters, etc. While specific examples are described below, it should be understood that such examples should not be construed as limiting the scope of this disclosure to the explicit descriptions of the examples set forth herein, but rather as providing illustrative examples.
[0082] Each of the modules and applications identified herein corresponds to a set of executable instructions for performing one or more functions described above and methods described in this application (e.g., computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) may optionally not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and therefore various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. For example, a video player module may optionally be combined with a music player module into a single module. In some embodiments, memory may optionally store a subset of the modules and data structures identified above. Furthermore, memory may optionally store additional modules and data structures not described above.
[0083] One or more steps of the method described herein may depend on satisfying one or more conditions. In some embodiments, the method is performed through multiple iterative processes. In some embodiments, the conditional steps may be satisfied in different iterations of the same process and still remain within the scope of the method described herein. For example, for a given method comprising two steps depending on different conditions, those skilled in the art will understand that the given method should be considered performed even if the process is repeated multiple times until the conditional step is satisfied. In some embodiments, multiple iterations of the process are not required to practice the claims as set forth herein. For example, the claims of an electronic device, system, or computer-readable medium may be performed without iteratively repeating the process. In some embodiments, the claims of an electronic device, system, or computer-readable medium include instructions for performing one or more steps depending on satisfying one or more conditions. Because such instructions are stored in one or more processors and / or one or more memory locations, the claims of an electronic device, system, or computer-readable medium may include logic for determining whether one or more conditions have been satisfied without requiring the steps of the process to be repeated.
[0084] Although numerical descriptors such as "first" and / or "second" are used below to describe elements, these elements do not correspond to sequential or different representations and should not be limited to the stated numerical terms. In some embodiments, these terms are used only as prefixes to distinguish references to one element from references to another. For example, "first" device and "second" device can be two separate references to the same device. Conversely, for example, "first" device and "second" device can be references to two different devices (e.g., not the same device and / or not the same type of device). For example, a first computer system and a second computer system do not correspond to first and second in time and are merely used to distinguish the two computer systems. Therefore, without departing from the scope of the various described embodiments, a first computer system may be referred to as a second computer system, and a second computer system may be referred to as a first computer system.
[0085] In the description of various elements and examples, the use of certain terms is intended to provide a productive description of the following topics and should not be construed as restrictive. As used in describing the various examples herein, the singular forms “a,” “an,” and “the” should not be construed as excluding or precluding the plural forms unless the context clearly indicates otherwise. Similarly, “and / or” is used to cover any and all possible combinations of one or more associated listed items. For example, “x and / or y” should be interpreted as including “x” or “y” as well as “x and y” as a possible permutation. Furthermore, the terms “includes,” “including,” “comprises,” and / or “comprising” used in this specification specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0086] When describing choices and / or logical possibilities, the term "if" may optionally be interpreted, depending on the context, as meaning "when," "in response to determination," "in response to detection," or "according to determination." Similarly, depending on the context, the phrases "if determination..." or "if [the stated condition or event] is optionally interpreted as meaning "in response to determination," "in response to detection," "in response to detection," or "according to determination."
[0087] The processes described below enhance device operability and make user-device interactions more efficient through various technologies (e.g., by helping users provide correct input and reducing user errors during operation / interaction with the device). These technologies include: providing users with improved feedback (e.g., visual, tactile, audible, and / or haptic feedback); reducing the amount of input required to perform an operation; providing additional control options without cluttering the user interface with additional displayed controls; performing an operation without further input (e.g., user input) when a set of conditions are met; and / or other technologies (such as improving the security and / or privacy of the computer system and reducing the aging of one or more parts of the display user interface). These technologies also reduce power consumption and extend device battery life by enabling users to use the device more quickly and efficiently.
[0088] under, Figure 1 , Figures 2A to 2C and Figures 3 to 5 A description of an exemplary device for performing the techniques described herein is provided. Figures 6A to 6E An exemplary user interface for performing movement based on learned characteristics, according to some implementation schemes, is illustrated. Figure 7 This is a flowchart illustrating a method for performing a movement representation based on learned characteristics, according to some implementation schemes. Figures 6A to 6E The user interface in the document is used to illustrate the processes described below, including Figure 7 The process in. Figures 8A to 8E An exemplary user interface for configuring operations to be performed based on learned characteristics, according to some implementation schemes, is illustrated. Figure 9 This is a flowchart illustrating a method for configuring operations to be performed based on learned characteristics, according to some implementation schemes. Figures 8A to 8E The user interface in the document is used to illustrate the processes described below, including Figure 9 The process in. Figures 10A to 10C An exemplary user interface for automatically outputting content based on context and / or input characteristics, according to some implementation schemes, is illustrated. Figure 11 This is a flowchart illustrating a method for automatically outputting audio content based on context, according to some implementation schemes. Figure 12 This is a flowchart illustrating a method for automatically outputting audio with specific characteristics according to some implementation schemes. Figures 10A to 10C The user interface in the document is used to illustrate the processes described below, including Figure 11 and Figure 12 The process in. Figures 13A to 13D An exemplary user interface for performing operations is illustrated according to some implementation schemes. Figure 14 This is a flowchart illustrating a method for performing an operation according to some implementation schemes. Figures 13A to 13DThe user interface in the document is used to illustrate the processes described below, including Figure 14 The process in. Figures 15A to 15C An exemplary user interface for predicting demand is illustrated according to some implementation schemes. Figure 16 This is a flowchart illustrating a method for predicting demand based on some implementation schemes. Figures 15A to 15C The user interface in the document is used to illustrate the processes described below, including Figure 16 The process in. Figures 17A to 17D An exemplary user interface for performing custom confirmations is illustrated according to some implementation schemes. Figure 18 This is a flowchart illustrating a method for performing customized verification according to some implementation schemes. Figure 19 This is a flowchart illustrating a method for performing customized verification according to some implementation schemes. Figures 17A to 17D The user interface in the document is used to illustrate the processes described below, including Figure 18 and Figure 19 The process in.
[0089] Figure 1 A block diagram depicts a computer system 100 (e.g., an electronic device and / or electronic system) comprising a set of electronic components communicating (e.g., connected) with each other (e.g., wired or wireless). It should be understood that computer system 100 is merely one example of a computer system that can be used to perform the functions described below, and one or more other computer systems can be used to perform the functions described below. Furthermore, although... Figure 1 The computer architecture of computer system 100 is described, but other computer architectures of computer systems (e.g., including more components, similar components and / or fewer components) may be used to perform the functionality described herein.
[0090] In some implementations, computer system 100 may correspond to (e.g., is and / or includes) a system-on-a-chip, a server system, a personal computer system, a smartphone, a smartwatch, a wearable device, a tablet computer, a laptop computer, a fitness tracker, a head-mounted display (HMD) device, a desktop computer, public equipment (e.g., smart speakers, connected thermostats and / or additional home-based computer systems), accessories (e.g., switches, lights, speakers, air conditioners, heaters, window covers, fans, locks, media playback devices, televisions, etc.), controllers, hubs and / or sensors.
[0091] In some embodiments, the sensor includes one or more hardware components capable of detecting (e.g., sensing, generating, and / or processing) information about the physical environment near the sensor. For example, the sensor may be configured to detect information around the sensor, detect information in one or more directions extending outward from the sensor, and / or detect information based on contact between the sensor and elements of the physical environment. In some embodiments, the hardware components of the sensor include sensing components (e.g., temperature and / or image sensors), transmitting components (e.g., radio and / or laser transmitters), and / or receiving components (e.g., laser and / or radio receivers). In some implementations, the sensors include angle sensors, breakage sensors, flow sensors, force sensors, gas sensors, humidity or moisture sensors, glass breakage sensors, chemical sensors, contact sensors, non-contact sensors, image sensors (e.g., RGB cameras and / or infrared sensors), particle sensors, photoelectric sensors (e.g., ambient light and / or sunlight), positioning sensors (e.g., GPS), precipitation sensors, pressure sensors, proximity sensors, radiation sensors, inertial measurement units, leak sensors, liquid level sensors, metal sensors, microphones, motion sensors, distance or depth sensors (e.g., RADAR, LiDAR), speed sensors, temperature sensors, time-of-flight sensors, torque sensors, ultrasonic sensors, vacancy sensors, presence sensors, voltage and / or current sensors, conductivity sensors, resistivity sensors, capacitance sensors, and / or water sensors. Although in Figure 1 Only a single computer system is depicted, but the functionality described below can be implemented using two or more computer systems operating together. Additionally, in some embodiments, computer system 100 includes one or more sensors as described above, and captures information about the physical environment by combining data from one sensor with data from one or more additional sensors (e.g., which are part of the computer and / or one or more additional computer systems).
[0092] like Figure 1As illustrated, computer system 100 comprises a processor subsystem 110, memory 120, and I / O interface 130. Memory 120 corresponds to system memory that communicates with processor subsystem 110. Electronic components constituting computer system 100 are electrically connected via interconnects 150, which allow communication between components of computer system 100. For example, interconnect 150 may be a system bus, one or more memory locations, and / or additional electrical channels for connecting multiple components of computer system 100. Additionally, I / O interface 130 is connected to I / O device 140 via wired and / or wireless connections. In some embodiments, computer system 100 includes a component comprising I / O interface 130 and I / O device 140, such that the functionality of each component is included within that component. Furthermore, it should be understood that computer system 100 may include one or more I / O interfaces that communicate with one or more I / O devices. In some embodiments, computer system 100 comprises multiple processor subsystems 100s, each processor subsystem being electrically connected via interconnect 150.
[0093] In some embodiments, processor subsystem 110 includes one or more processors or separate processing units capable of executing instructions (e.g., programs, systems, and / or interrupts) to perform the functionality described herein. For example, operating system-level and / or application-level instructions executed by processor subsystem 110. In some embodiments, processor subsystem 110 includes one or more components (e.g., implemented as hardware, software, and / or combinations thereof) capable of supporting, interpreting, and / or executing machine learning instructions and / or operations. For example, computer system 100 may perform operations locally based on a machine learning model. Alternatively or additionally, computer system 100 may communicate with (e.g., perform computations thereto and / or execute corresponding instructions) a remote interactive knowledge base (e.g., processing resources implementing machine learning models, artificial intelligence models, and / or large language models) to perform operations that may otherwise be outside the set of capabilities of computer system 100. For example, computer system 100 may determine a set of inputs (e.g., instructions, data, and / or parameters) to an interactive knowledge base for performing desired machine learning operations.
[0094] The memory 120, which communicates with the processor subsystem 110, can be implemented using a variety of different physical, non-transitory memory media. In some embodiments, the computer system 100 includes multiple memory components and / or various types of memory components, each of which is directly and / or connected to the processor subsystem 110 via interconnect 150. For example, the memory 120 can be implemented using removable flash drives, storage arrays, storage area networks (e.g., SANs), flash memory, hard disk storage devices, optical drive storage devices, floppy disk storage devices, removable disk storage devices, random access memory (e.g., SDRAM, DDR SDRAM, RAM-SRAM, EDO RAM, and / or RAMBUS RAM), and / or read-only memory (e.g., PROM and / or EEPROM). Additionally, in some embodiments, the processor subsystem 110 and / or interconnect 150 are connected to a memory controller, which is electrically connected to the memory 120.
[0095] In some embodiments, the instructions may be executed by processor subsystem 110. In this example, memory 120 may include a computer-readable medium (e.g., a non-transitory or transient computer-readable medium) that can be used to store (e.g., configured to store, assigned to store, and / or store) instructions executable by processor subsystem 110. In some embodiments, each instruction stored by memory 120 and executed by processor subsystem 110 corresponds to an operation for performing the functionality described herein. For example, memory 120 may store program instructions to implement the methods described below (including methods 700, 900, 1100, 1200, 1400, 1600, 1800, and / or 1900). Figure 7 , Figure 9 , Figure 11 , Figure 12 , Figure 14 , Figure 16 , Figure 18 and / or Figure 19 Related functionality.
[0096] As mentioned above, I / O interface 130 may be one or more types of interfaces that enable computer system 100 to communicate with other devices. In some embodiments, I / O interface 130 includes a bridge chip (e.g., a southbridge) connecting a front-side bus to one or more back-side buses. In some embodiments, I / O interface 130 enables communication with one or more I / O devices (exemplified as I / O device 140) via one or more corresponding buses or other interfaces. For example, I / O devices may include one or more of the following: physical user interface devices (e.g., physical keyboard, mouse, and / or joystick), storage devices (e.g., as described above with respect to memory 120), network interface devices (e.g., to a local area network or wide area network), sensor devices (e.g., as described above with respect to sensors), and / or auditory and / or visual output devices (e.g., screens, speakers, lamps, and / or projectors). In some embodiments, the visual output device is referred to as a display component. For example, a display component may be configured to provide visual output, such as displaying images on a physical visual medium via an LED display or image projection. As used herein, “display” content includes content that is displayed by sending data (e.g., image data and / or video data) to an integrated or external display component via a wired or wireless connection to visually generate content (e.g., video data rendered and / or decoded by a display controller).
[0097] In some embodiments, computer system 100 includes a component that integrates I / O device 140 with other components (e.g., a component including I / O interface 130 and I / O device 140). In some embodiments, I / O device 140 is separate from other components of computer system 100 (e.g., it is a discrete component). In some embodiments, I / O device 140 includes a network interface device that allows computer system 100 to connect to a network or other computer system (e.g., communicate with it) via wired or wireless means. In some embodiments, the network interface device may include Wi-Fi, Bluetooth, NFC, USB, Thunderbolt, Ethernet, etc. For example, computer system 100 may utilize NFC connectivity to facilitate banking, credit, financial, token (e.g., fungible or non-fungible tokens), and / or cryptocurrency transactions between computer system 100 and another nearby computer system.
[0098] In some embodiments, I / O device 140 includes components for detecting users (e.g., individuals, people, animals, another computer system different from the computer system, and / or objects) and / or input from the detected users (e.g., tap input and / or non-tap input (e.g., verbal input, audible requests, audible commands, audible statements, swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)). In some embodiments, I / O device 140 enables computer system 100 to identify users associated with and / or not having accounts within the environment. For example, computer system 100 may detect known users (e.g., users corresponding to accounts) and access information about the users using the known users' accounts. In some embodiments, as part of computer system 100's user detection, computer system 100 detects that a user's account is associated with a group of users (e.g., included in and / or identified relative to that group of users). For example, computer system 100 may access information associated with accounts in a family defined as a group of accounts in response to detecting a member of that family. In some implementations, the user's account may be connected to additional accounts and / or additional computer systems. For example, computer system 100 may detect such additional computer systems and / or detect such computer systems used to detect users. In some implementations, computer system 100 may detect unknown users and enable guest accounts of unknown users to utilize computer system 100.
[0099] In some embodiments, I / O device 140 includes one or more cameras. In some embodiments, the camera includes an image sensor (e.g., one or more optical sensors and / or one or more depth camera sensors) that provides computer system 100 with the ability to detect user and / or user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, air gestures are gestures detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)). In some implementations, one or more cameras enable computer system 100 to send image and / or video information to an application. For example, image data captured by a camera can enable computer system 100 to complete a video call by sending video data to an application used to perform the video call.
[0100] In some embodiments, I / O device 140 includes one or more microphones. For example, the microphone may be used by 100 to obtain data and / or information from a user without contact input. In some embodiments, the microphone enables computer system 100 to detect verbal and / or voice input from a user. In some embodiments, computer system 100 utilizes voice input to enable personal assistant functionality. For example, a user makes a request to computer system 100 to perform an action and / or obtain information from the user. In some embodiments, computer system 100 utilizes voice input (e.g., in conjunction with one or more other input and / or output technologies) to request and / or detect information from a user without requiring physical contact between the user and computer system 100.
[0101] In some embodiments, I / O device 140 includes physical input media for a user to interact directly with computer system 100. In some embodiments, the physical input media includes one or more physical buttons (e.g., tactilely pressable buttons and / or touch-sensitive non-pressable components) on and / or connected to computer system 100, mouse and keyboard input methods (e.g., connected to computer system 100 together with and / or separately from one or more I / O interfaces), and / or touch-sensitive display components.
[0102] In some embodiments, I / O device 140 includes one or more components for outputting information (e.g., display components, audio generation components, speakers, haptic output devices, displays, projectors, and / or touch-sensitive displays). In some embodiments, computer system 100 uses I / O device 140 to transmit information and / or the state of computer system 100. In some embodiments, I / O device 140 includes haptic output components. For example, the haptic output component may be a haptic generation component that enables computer system 100 to transmit information to a user who is in contact with computer system 100 (e.g., holding, touching, and / or near it). In some embodiments, I / O device 140 includes one or more components for outputting visual output (e.g., video, images, animations, 3D rendering, augmented reality overlay, motion graphics, data visualization, digital art, etc.). For example, displaying content from one or more applications and / or system applications, and / or displaying widgets corresponding to one or more applications (e.g., controls displaying real-time information and / or data).
[0103] In some implementations, I / O device 140 includes one or more components for outputting audio (e.g., smart speaker, home theater system, soundbar, headphones, earphones, earbuds, speaker, TV speaker, augmented reality headphone speaker, audio jack, optical audio output, Bluetooth audio output, HDMI audio output, audio sensor, etc.). In some implementations, computer system 100 is capable of outputting audio through one or more speakers. For example, computer system 100 outputs audio-based content and / or information to a user. In some implementations, one or more speakers enable spatial audio (e.g., audio output corresponding to the environment (e.g., computer system 100 detects materials and / or objects in the environment and / or computer system 100 changes audio modes, intensities, and / or waveforms to compensate for changing environmental characteristics)).
[0104] Figure 2 to Figure 5Exemplary components and user interfaces of an electronic device 200 according to some embodiments are illustrated. The electronic device 200 (sometimes referred to herein as device 200) may include one or more features of the computer system 100. (Referring to Figures 2 to...) Figure 5 In the described example, device 200 is a laptop computer. In some embodiments, device 200 is not limited to a laptop computer, and those skilled in the art will recognize that device 200 can be one or more other devices (e.g., one or more of the components and / or functions described herein with respect to device 200). For example, device 200 can be a public device (such as a smart display, smart speaker, and / or television) and / or a personal device (such as a smartphone, smartwatch, tablet, desktop computer, fitness tracker, and / or head-mounted display). In some embodiments, the public device is configured to provide functionality to multiple users (e.g., simultaneously and / or at different times). In such embodiments, the public device can be managed and / or set by a single user. In some embodiments, the personal device is configured to provide functionality to a single user (e.g., once, such as when a single user logs into the personal device).
[0105] Figures 2A to 2C An example is shown of a device 200 located in three different physical locations. For example... Figure 2A As illustrated, device 200 is a laptop computer (also referred to herein as a "laptop"), which includes a base portion 200-2 (e.g., as shown in the image). Figure 2A The device 200 is horizontally placed on a surface such as a table and connected to a base portion 200-2 at a connection 200-3 (e.g., one or more connection points, motor arms, hinges, and / or joints). This connection allows the display portion 200-1 to pivot and / or change orientation relative to the base portion 200-2. For example, the device 200 may pivot at the connection 200-3 to rotate the display portion 200-1 and / or the device 200 to one or more positions corresponding to the “closed” internal state (e.g., as described below regarding...). Figure 2C(Further description). In some embodiments, the positioning corresponding to the "off" internal state is the positioning of the device 200 in a predetermined pose. For example, the predetermined pose may include a display portion 200-1 positioned parallel to the base portion 200-2 or forming a predetermined angle (e.g., 60 degrees) with respect to the base portion 200-2. In some embodiments, in the "off" internal state, the area of the device 200 in which content is displayed is positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., an area facing downwards, not visible, and / or obscuring the displayed content). In some embodiments, in the "off" internal state, the area of the device 200 in which content is displayed is not positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., instead positioned in a manner corresponding to the "on" internal state). For example, when not in a "closed" internal state, device 200 can be positioned within a range of different open positions (e.g., where display portion 200-1 is not parallel to base portion 200-2, and where the area where the content displayed by device 200 is visible and / or unobstructed). It should be recognized that display portion 200-1 being parallel to base portion 200-2 is an example of positioning corresponding to a "closed" internal state of device 200 (e.g., closed positioning). In some embodiments, another configuration may set another orientation of display portion 200-1 relative to base portion 200-2 as a closed positioning of device 200, such as... Figure 2C exemplified.
[0106] Figure 2A The left side illustrates display screen 200-4 (representing the area where device 200 displays content), and the right side illustrates device 200 in the corresponding pose. For example... Figure 2A As illustrated, device 200 is in a first position (e.g., display portion 200-1 is perpendicular to base portion 200-2, forming a 90-degree angle). Figure 2A In this context, display screen 200-4 represents the content currently being displayed (e.g., via a display component) when device 200 is first activated. Figure 2A In this embodiment, display screen 200-4 illustrates the device 200 in an "on" internal state (e.g., operable, powered, awake, higher power and / or more resource-intensive than the "off" state, and / or activated). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces (e.g., user interface objects, windows, application user interfaces, system user interfaces, controls, and / or other visual content). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces while in an "on" internal state. For example, in Figure 2AIn this configuration, device 200 is in an "on" internal state, and display screen 200-4 shows a desktop user interface 200-5, including an application window. In some embodiments, the user interface includes (and / or) one or more user interface objects (e.g., windows, icons, and / or other graphical objects). For example, the user interface (e.g., 200-5) may include one or more graphical objects that are different from and / or the same as the application window.
[0107] Figure 2B Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2B As illustrated, device 200 is in a second position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming an angle of 120 degrees (e.g., more than). Figure 2A (at a larger angle). Figure 2B In the diagram, display screen 200-4 represents the content being displayed when device 200 is in the second position. Display screen 200-4 illustrates the internal state of device 200 being "on" (e.g., with...). Figure 2A (The top diagram shows the same internal state). Figure 2B In the process, device 200 displays (e.g., via display screen 200-4) a desktop user interface 200-5 (e.g., with...). Figure 2A (The same as shown in the image). In some implementations, device 200 displays a different user interface (e.g., different from desktop user interface 200-5). For example, although... Figure 2B Example of device 200 in a state of being with Figure 2A Different positioning displays and Figure 2A The same desktop user interface 200-5 exists, but device 200 may display different user interfaces. In some embodiments, device 200 displays a user interface corresponding to (e.g., based on, due to, caused by, involved in, and / or configured to accompany) a physical state (e.g., positioning, location, and / or orientation), including content specific to a particular angle or specific to the current context.
[0108] Figure 2C Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2C As illustrated, device 200 is in a third position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming a 60-degree angle (e.g., compared to...). Figure 2A and Figure 2B (smaller angles)). Figure 2C In the diagram, display screen 200-4 shows the content being displayed when device 200 is in the third position. Figure 2CIn the diagram, displays 200-4 illustrate an internal state in which device 200 is "off" (e.g., not operating, not powered, not woken up, not activated, powered off, asleep, hibernating, inactive, and / or disabled). In some embodiments, device 200 does not display (e.g., via displays 200-4) one or more user interfaces (e.g., no visual content is displayed) when it is in the "off" internal state. In some embodiments, device 200 displays (e.g., via displays 200-4) one or more user interfaces (e.g., the same as and / or different from one or more user interfaces displayed when it is in the "on" internal state) (e.g., a user interface specific to the "off" state and / or a way of displaying a user interface not specific to the "off" internal state). Figure 2C In this case, display screen 200-4 is blank because nothing is displayed on the monitor of device 200 (e.g., display screen 200-4 is off and / or does not display the user interface) (e.g., desktop user interface 200-5 is not displayed on display screen 200-4).
[0109] In some embodiments, device 200 includes one or more components (referred herein also as “moving components”) that enable device 200 to perform (e.g., cause and / or control) movement (and / or be moved). For example, performing movement may include a portion of mobile device 200 (e.g., less or all components of the device moving), all of mobile device 200 (e.g., the entire device (including all its components) moving, such as by changing position), and / or moving one or more other devices and / or components (e.g., communicating with device 200 and / or the moving components of device 200). For example, device 200 may move automatically (e.g., pivot), cause and / or control movement of display portion 200-1 relative to base portion 200-2, such as moving to... Figures 2A to 2CAny of the illustrated locations. In some embodiments, device 200 performs movement based on its internal state. Performing movement based on internal state enables device 200 to perform new (e.g., otherwise unavailable) interactions. For example, such new interactions of device 200 can be configured using special features, functions, patterns, and / or procedures that leverage device 200's ability to perform movement. Examples of such interactions include using movement to (e.g., to a user) convey the device's internal state (e.g., on, off, sleep, and / or hibernate) to assist user input (e.g., shorten the distance to the user) and / or enhance the device's interactive behavior (e.g., moving in a specific manner during interaction with the user, conveying information such as importance and / or direction of attention). In some embodiments, the performed movement corresponds to (e.g., caused by, responded to, and / or determined and / or performed based on) one or more of the following: detected input, detected context (e.g., environmental context and / or user context), and / or the device 200's internal state (e.g., internal state and / or a set of multiple internal states). For example, device 200 can move the display portion, causing device 200 to move from a position where... Figure 2A The illustrated first positioning moves to the position where Figure 2B The illustrated second positioning. In this example, device 200 can detect that the user has repositioned relative to device 200 (e.g., the user stands up), and in response, device 200 can perform a movement to the second positioning such that the display is at an optimized viewing angle based on the height and / or angle of the user's eye relative to the display of device 200. As another example, device 200 can perform a movement such that device 200 moves from a position... Figure 2A The illustrated first positioning moves to the position where Figure 2C The illustrated third location. In this example, device 200 may perform a movement to a third location in response to detecting an internal state with reduced activity (e.g., an "off" internal state as described above). In this way, movement of device 200 to one or more locations can indicate the internal state of device 200.
[0110] Figures 2A to 2C An example is illustrated of a device 200 having a display portion capable of moving with one degree of freedom via a connection 200-3 (e.g., a hinge) connecting the display portion 200-1 to a base portion 200-2. In some embodiments, the device 200 includes one or more components having one or more degrees of freedom. For example, a moving component of the device 200 (e.g., an output component that causes and / or allows movement) (e.g., Figure 5Device 200-26C may include multiple degrees of freedom (e.g., six degrees of freedom including three translational components and three rotational components). For example, device 200 may be implemented to move the display portion by telescopic forward or backward movement (e.g., display portion 200-1 moves forward in space relative to the base portion while the base portion 200-2 remains stationary (e.g., to shorten and / or lengthen the user's viewing distance)). As yet another example, device 200 may be implemented to move the display portion to rotate about an axis perpendicular to the hinge, such that the display portion can rotate to position the display to follow the user as the user walks around device 200. Although Figures 2A to 2C The example shown illustrates a hinge, but other moving components may be included in device 200, such as actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), movable bases, rotatable components, and / or rotatable bases. In some embodiments, one or more moving components may enable device 200 to move in different ways, such as rotation (e.g., 0 to 360 degrees), lateral movement (e.g., to the right, left, down, up, and / or any combination thereof), and / or tilting (e.g., 0 to 360 degrees).
[0111] Figure 3 An exemplary block diagram of device 200 is illustrated. In some embodiments, device 200 includes... Figure 1 A, Figure 1 B. Figure 3 and Figure 5 B describes some or all of the components. For example... Figure 3 As illustrated, device 200 has a bus 200-13 that operatively couples I / O segments 200-12 (also referred to as I / O sub-segments and / or I / O interfaces) to processor 200-11 and memory 200-10. For example... Figure 3 As illustrated, I / O section 200-12 is connected to output device 200-16 (also referred to herein as "output component"). In some embodiments, output device 200-16 includes one or more visual output devices (e.g., display components such as monitors, displays, projectors, and / or touch-sensitive displays), one or more tactile output devices (e.g., devices that cause vibration and / or other tactile outputs), one or more audio output devices (e.g., speakers), and / or one or more moving components (e.g., actuators, motors, mechanical linkages, devices that cause and / or allow movement, and / or one or more moving components as described above). Figure 3As illustrated, output device 200-16 includes two exemplary moving components (e.g., a movement controller 200-17 and an actuator 200-18). Actuator 200-18 can be any component that performs (e.g., partial and / or overall) physical movement of a device (e.g., device 200 and / or devices coupled to and / or in contact with that device). Movement controller 200-17 can be any component (e.g., a control device) that controls actuator 200-18 (e.g., provides control signals to it). For example, movement controller 200-17 can provide control signals that actuate actuator 200-18 (e.g., cause physical movement). In some embodiments, movement controller 200-17 includes one or more logic components (e.g., a processor), one or more feedback components (e.g., sensors), and / or one or more control components (e.g., for applying control signals, such as relays, switches, and / or control lines). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in the same device and / or component (e.g., a dedicated onboard motion controller 200-17 attached to the actuator 200-18). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in different devices and / or components (e.g., one or more processors 200-11 may serve as the motion controller 200-17 for the actuator 200-18). In some embodiments, the motion controller 200-17 and / or the actuator 200-18 are embodied in a device (or one or more devices) other than device 200 (e.g., device 200 is coupled to (e.g., temporarily and / or removably) another device and may instruct the motion controller 200-17 and / or the actuator 200-18 to control the other device). Actuator 200-18 can be used to induce one or more types of mechanical movement (e.g., linear and / or rotary movement) in one or more ways (e.g., using electric, magnetic, hydraulic and / or pneumatic power). Examples of actuator 200-18 may include electromechanical actuators, linear actuators and / or rotary actuators.
[0112] like Figure 3As illustrated, I / O section 200-12 is connected to input device 200-14. In some embodiments, input device 200-14 includes one or more visual input devices (e.g., cameras and / or light sensors), one or more physical input devices (e.g., buttons, sliders, switches, touch-sensitive surfaces, and / or rotatable input mechanisms), one or more audio input devices (e.g., microphones), and / or other input devices (e.g., accelerometers, pressure sensors (e.g., contact strength sensors), distance sensors, temperature sensors, GPS sensors, accelerometers, orientation sensors (e.g., compasses), gyroscopes, motion sensors, and / or biometric sensors). Furthermore, I / O section 200-12 may be connected to communication unit 200-15 for receiving application and operating system data using Wi-Fi, Bluetooth, Near Field Communication (NFC), cellular, and / or other wireless (and / or wired) communication technologies.
[0113] The memory 200-10 of the personal electronic device 200 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions, which, when executed by one or more computer processors 200-11, cause the computer processors to perform, for example, the techniques described below, including processes 700 and 900. Figure 7 and Figure 9 A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device. In some embodiments, the storage medium is a transient computer-readable storage medium. In some embodiments, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, and Blu-ray technologies, and persistent solid-state storage such as flash memory and solid-state drives. Electronic device 200 is not limited to Figure 3 The components and configurations may include, but may include, other and / or additional components in a variety of possible configurations, all of which are intended to fall within the scope of this disclosure.
[0114] Figure 4 A functional diagram of actuator 200-18B according to some embodiments is illustrated. As described above, actuator 200-18B can be any component that performs physical movement. In some embodiments, actuator 200-18B is operated using inputs including control signal 200-18A and / or energy source 200-18B. For example, actuator 200-18B can be a rotary actuator that converts electrical energy into rotational movement. This rotational movement can cause the aforementioned... Figures 2A to 2CThe movement of the display portion of the described device 200 (e.g., the counterclockwise rotation of the actuator moves the device 200 to a position with a large angle) Figure 2B The illustrated second positioning), and the clockwise (e.g., counterclockwise) rotational movement of the actuator moves the device 200 to a positioning with a smaller angle (e.g., Figure 2C (The illustrated third positioning). Control signal 200-18A may indicate one or more start and / or stop commands, movement and / or actuation direction, movement and / or actuation speed, movement and / or actuation time, target positioning (e.g., pose and / or position) of movement and / or actuation, and / or one or more other characteristics of movement and / or actuation. In some embodiments, the control signal and the energy source are the same signal and / or input. In some embodiments, one or more additional components (e.g., mechanical and / or electrical) (e.g., removably or permanently) are coupled to actuator 200-18B to influence movement and / or actuation (e.g., mechanical linkages such as lead screws, gears, and / or other components for changing (e.g., switching) the characteristics of movement and / or actuation). In some embodiments, actuator 200-18B includes one or more feedback components (e.g., a positioning sensor, an encoder, an overcurrent sensor, and / or a force sensor) that form part of a feedback loop for modifying and / or stopping movement and / or actuation (e.g., slowing down actuation upon reaching a target position and / or stopping actuation if physical resistance to actuation is detected via a sensor). In some embodiments, one or more feedback components are included (e.g., partially and / or entirely) in a motion controller (e.g., motion controller 200-13) operatively coupled to the actuator.
[0115] Now turn attention to the functionality (e.g., features and / or capabilities) of one or more devices (e.g., computer system 100 and / or electronic device 200). One such functionality is the implementation of an "intelligent agent," which may alternatively be referred to as a software intelligent agent, intelligent intelligent agent, interactive intelligent agent, virtual assistant, intelligent virtual assistant, interactive virtual assistant, personal assistant, intelligent personal assistant, interactive personal assistant, intelligent interactive personal assistant, and / or artificial intelligence (AI) assistant. In some embodiments, an intelligent agent refers to one or more sets of functions implemented in hardware and / or software (e.g., local and / or remote) on an intelligent agent system (e.g., a single device and / or multiple devices). In some embodiments, the intelligent agent performs operations to perceive the environment, acquire knowledge, retrieve knowledge, learn skills, interact with a user, and / or perform tasks. The intelligent agent may perform these (and / or other) operations, for example, in response to user input and / or automatically (e.g., at an appropriate time determined based on the perceived context). An incomplete list of exemplary operations that an intelligent agent can be used with and / or employed therewith includes: tracking a user’s eyes, face, and / or body (e.g., to move with the user and / or identify the user’s intentions and / or activities); detecting, identifying, and / or classifying users in the environment; detecting and / or responding to input (e.g., verbal input, air gestures, and / or physical input, such as touch input and / or force input to physical hardware components (e.g., buttons, knobs, and / or sliders); detecting context (e.g., user context, operational context, and / or environmental context); moving (e.g., changing pose, orientation, orientation, and / or location); performing one or more operations in response to input, context, and / or stimuli (e.g., objects or events that elicit one or more responsive operations on the device (e.g., outside and / or inside the device)); providing intelligent interaction capabilities (e.g., in part due to one or more machine learning (“ML”) models, such as large language models (“LLM”)) to respond to and / or perform operations; and / or (e.g., automatically and / or intelligently) performing tasks (e.g., a set of operations for achieving a specific goal). In some implementations, the agent performs actions in response to contactless input (e.g., air gestures and / or natural language commands). The foregoing list is intended to exemplify actions that can be performed by an agent, but is not intended to be an exhaustive list. Other actions fall within the expected scope of the agent's capabilities. Furthermore, for the purposes of this disclosure, the agent need not include all the functionalities mentioned herein, but may include fewer or more functionalities (e.g., the agent may be implemented on an agent system that does not have mobile functionality but otherwise includes an intelligent personal assistant capable of interacting with a user).
[0116] In some implementations, a user is one or more individuals, people, objects, and / or animals within an environment (e.g., a device) that is perceived (e.g., detected by the device, one or more other devices, and / or one or more of its components). In some implementations, a user is an entity that is distinguished from surrounding entities (e.g., components of the environment and / or other users) and / or is considered to be a discrete logical construct via one or more components (e.g., a sensing component and / or other components). In some implementations, a user is physical and / or virtual. For example, a physical user may represent a user standing in front of the device and perceived by the device. As another example, a virtual user may represent an avatar in a virtual scene perceived by the device (e.g., an avatar detected in a media stream received by the device and / or captured by the device's camera). Although presented above as an example of “user,” throughout this disclosure, terms and / or concepts referred to as “individual,” “person,” “object,” and / or “animal” are interchangeable with “user” unless otherwise expressly indicated. For example, unless otherwise expressly indicated, the use of the term “individual” is also understood to mean “user.”
[0117] As an example, and to revisit Figures 2A to 2C An agent, at least partially implemented on device 200, can perform operations that cause the display portion 200-1 of device 200 to move relative to the base portion 200-2. For example, agent detection (e.g., perceiving and determining its occurrence) includes contexts such as a user standing up (e.g., based on face detection and tracking); and in response, the agent causes device 200 to open and / or device 200 to open the display portion 200-1 to a greater angle. As another example, the agent can detect verbal input corresponding to (e.g., interpreted as and / or implying) a request to move the display (e.g., “Please move my display” or “Please enter sleep mode”); and in response, the agent causes device 200 to move and / or device 200 to move the display portion 200-1.
[0118] Figure 5 A functional diagram of an exemplary intelligent agent system 200-20A is shown. Figure 5 As illustrated, the agent system 200-20A has a dashed box boundary surrounding the input component 200-22, the agent component 200-24, and the output component 200-26. In some embodiments, the agent system 200-20A includes more than Figure 5The system may contain fewer, more, and / or different components as illustrated herein. In some embodiments, the agent system 200-20 is implemented on a single device (e.g., computer system 100 and / or electronic device 200). In some embodiments, the agent system 200-20 is implemented on multiple devices. In some embodiments, in Figure 5 One or more components of the agent system 200-20 illustrated and / or described with respect to this figure are external to but operatively coupled to the agent system (e.g., accessories, external devices, external sensors, external actuators, external display components, external speakers, and / or external databases). In some embodiments, one or more components of the agent system 200-20 are local to one or more other components of the agent system 200-20. In some embodiments, one or more components of the agent system 200-20 are remote from one or more other components of the agent system 200-20.
[0119] In some implementations, input components 200-22 include components for performing sensing and / or communication functions of the agent system 200-20. For example... Figure 5 As illustrated, input components 200-22 include one or more sensors 200-22A. The one or more sensors 200-22A may include any components for detecting data corresponding to the physical environment. Examples of the one or more sensors 200-22A may include: cameras, light sensors, microphones, accelerometers, positioning sensors, pressure sensors, temperature sensors, olfactory sensors, and / or contact sensors. This list is not intended to be exhaustive, and the one or more sensors 200-22A may include other sensors not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to detect data corresponding to the physical environment. Figure 5 As illustrated, input component 200-22 includes one or more communication components 200-22B. The one or more communication components 200-22B may include any component (e.g., antenna, modem, network interface component, encoder, decoder, and / or communication protocol stack) for transmitting and / or receiving communications internal and / or external to the intelligent agent system 200-20. Communication components 200-22B may be between different devices and / or between components within the same device. Communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, input component 200-22 includes more than Figure 5 The components illustrated herein may be fewer, more, and / or different. In some implementations, input components 200-22 are implemented in hardware and / or software.
[0120] In some implementations, agent components 200-24 include components that manage and / or perform the functions of the agents in agent system 200-20. For example... Figure 5 As illustrated, agent components 200-24 include the following functional components: task flow, coordination and / or orchestration component 200-24A, management component 200-24B, perception component 200-24C, evaluation component 200-24D, interaction component 200-24E, policy and decision-making component 200-24F, knowledge component 200-24G, learning component 200-24H, model component 200-24I, and API component 200-24J. Each of these components is briefly described below. It is important to note that this list of agent components 200-24 is not intended to be exhaustive, and agent components 200-24 may include other functional components not explicitly identified herein, which may be used (e.g., to process, store, and / or transform) any function of the agent, such as those described herein. In some embodiments, agent components 200-24 include more than Figure 5 The components illustrated herein may be fewer, more, and / or different. In some implementations, the agent components 200-24 are implemented in hardware and / or software.
[0121] In some implementations, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to manage coordination between various components. For example, operations may include managing a data processing task flow to move from perception components 200-24C (e.g., those detecting speech input) to model components 200-24I (e.g., for processing the detected speech input using a large language model to determine the content and / or intent of the speech input). In some implementations, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to manage coordination between one or more external components (e.g., resources). For example, Figure 5 Examples of external components (such as external databases 200-30) are illustrated. In some embodiments, management component 200-24B includes functionality performed by the operating system of the device implementing the intelligent agent system 200-20. In some embodiments, management component 200-24B includes functionality performed by one or more applications of the device implementing the intelligent agent system 200-20.
[0122] In some embodiments, management components 200-24B perform operations that enable the agent system to handle management tasks, such as managing system and / or component updates, managing user accounts, and managing system settings and / or component settings. In some embodiments, management components 200-24B include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, management components 200-24B include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0123] In some embodiments, the sensing components 200-24C perform operations that enable the agent to perceive environmental input. For example, operations may include detecting that context and / or environmental conditions have occurred, detecting the presence of a user (e.g., an individual, person, object, and / or animal in the environment), detecting input including voice, detecting input including air gestures, detecting facial expressions, detecting user characteristics (e.g., visible and / or invisible), and / or detecting verbal and / or physical cues. In some embodiments, the sensing components 200-24C include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the sensing components 200-24C include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0124] In some embodiments, the evaluation component 200-24D performs operations that enable the agent to process evaluation data (e.g., to determine context, such as user context, environmental context, and / or operational context). For example, the operations may include evaluating data collected from the perception component 200-24C, the knowledge component 200-24G, the external database 200-30, and / or the remote processing resource 200-32. In some embodiments, the evaluation component 200-24D includes functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the evaluation component 200-24D includes functionality performed by one or more applications of the device implementing the agent system 200-20.
[0125] This document refers to an environmental context (also referred to herein as "context of the environment" and / or "context corresponding to the environment"). In some embodiments, an environmental context is a context based on one or more characteristics of the environment (e.g., user, location, time, weather, and / or lighting). For example, an environmental context may include rain outside, daytime, and / or the device currently being in a park. In some embodiments, the device (e.g., using an agent) uses one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device) to determine the environmental context (e.g., currently true, happening, and / or applicable).
[0126] This document refers to user context (also referred to herein as "user context" and / or "context corresponding to a user") (and / or user context). In some embodiments, user context is a context based on one or more characteristics of a user (and / or a user). For example, user context may include a user's appearance and / or clothing, personality, actions, behaviors, movement, location, and / or pose. In some embodiments, a device (e.g., using an agent) determines user context (e.g., currently true, happening, and / or applicable) using one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, a device determines user context based on historical context and / or learned user characteristics, wherein one or more user characteristics are learned and / or stored by the device over a period of time.
[0127] This document refers to an operational context (also referred to herein as "the context of operation" and / or "operational context"). In some embodiments, an operational context is a context based on one or more characteristics of the device's operation (e.g., the device and / or one or more other devices that determine and / or access the operational context). For example, an operational context may include the internal state of the device (and / or one or more components of the device), the device's internal dialogue (e.g., the device's understanding of the context), the operations performed by the device, and applications and / or processes executed on the device (e.g., running and / or opening). In some embodiments, the device (e.g., using an agent) uses one or more of the following to determine the operational context (e.g., currently true, happening, and / or applicable): detected input (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, the device (e.g., using an agent) uses one or more internal states (e.g., accessed, retrieved, and / or queried by the device's processes) to determine the operational context (e.g., currently true, happening, and / or applicable).
[0128] In some embodiments, the interaction components 200-24E perform operations that enable the agent to manage and / or perform interactions with a user. For example, operations may include determining an appropriate interaction model for a specific context and / or in response to specific inputs. In some embodiments, the interaction components 200-24E include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the interaction components 200-24E include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0129] In some embodiments, the policy and decision components 200-24F perform operations that enable the agent to take actions based on available data. For example, operations may include determining which operations to perform and / or which functional components to utilize in response to detected context. In some embodiments, the policy and decision components 200-24F include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the policy and decision components 200-24F include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0130] In some embodiments, knowledge components 200-24G perform operations that enable the agent to access and use the stored knowledge. For example, operations may include indexing, storing, and / or retrieving data from a data repository, database, and / or other resource. In some embodiments, knowledge components 200-24G include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, knowledge components 200-24G include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0131] In some embodiments, the learning components 200-24H perform operations that enable the agent to learn through experience. For example, operations may include observing and / or tracking data, including preferences, routines, user characteristics, and / or environmental characteristics, in a way that allows the data to inform future actions of the agent and / or its components (e.g., when performing tasks and / or interacting with a user). In some embodiments, the learning components 200-24H include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the learning components 200-24H include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0132] In some embodiments, model components 200-24I perform operations that enable the agent to apply an ML model (e.g., a large language model (LLM)) to process data. For example, operations may include storing the ML model, executing the ML model, training and / or retraining the ML model, and / or otherwise managing aspects of implementing the ML model. In some embodiments, model components 200-24I include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, model components 200-24I include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0133] In some implementations, the agent system 200 responds to natural language input. For example, the agent system 200 responds to natural language input in the form of statements, questions, commands, and / or requests. In some implementations, the agent system 200 outputs text and / or speech output provided in natural language or mimicking the style of natural language. For example, the agent system 200 may use a speech response indicating the current outside temperature at the user's location (e.g., "It's 18 degrees outside") to handle the natural language question "How hot is it outside?". In some implementations, the agent system 200 responds to natural language input by providing information (e.g., weather, travel, and / or calendar information) and / or performing tasks (e.g., opening a document, searching a database, and / or opening an application).
[0134] In some embodiments, the intelligent agent system 200 includes and / or relies on one or more data models to process inputs (e.g., natural language input, gesture input, visual input, and / or other data input) and / or provide outputs (e.g., information output via natural language output, visual output, audio output, and / or text output). Such data models may include user data (e.g., data based on a specific interaction and / or from the user with whom the interaction takes place) and / or global data (e.g., general data based on the interaction and / or data from many users) and / or be trained using user data and / or global data. For example, user data (e.g., preferences, previous use of language and / or phrases, calendar entries, contact lists, and / or activity data) can be used to better infer user intent and / or provide responses more likely to resolve user requests. In some embodiments, the data models used by the intelligent agent system 200 include one or more machine learning components (e.g., hardware and / or software) (e.g., one or more neural networks), which are used by and / or implemented using them. Such machine learning components can be used to process speech input to determine words and / or phrases therein, one or more contexts corresponding to the words, user intent corresponding to the words, one or more confidence scores, and / or a set of one or more actions to be taken in response to the speech input. Similar operations can be performed to process other types of input, such as visual input, data input, and / or text input. Such data models may include machine learning and / or data processing models, including but not limited to natural language processing models, language models, speech recognition models, object recognition models, visual processing models, ontology, task flow models, and / or intent recognition models (e.g., for determining user intent).
[0135] In some implementations, application programming interface (API) components 200-24J perform operations that enable agents to interface with services, devices, and / or components. For example, operations may include relaying data (e.g., requests, responses, and / or other messages) between data interfaces (e.g., between software programs, between system processes and application processes, between system processes, between application processes, between communication protocols, between clients and servers, between file systems, and / or between components on different sides of a trust boundary). In some implementations, the data interfaces served by API components 200-24J are local (e.g., for a device, such as two application processes exchanging data) and / or remote (e.g., from a device, such as interfacing with a web service via a remote server). In some implementations, API components 200-24J include functionality performed by the operating system of the device implementing agent system 200-20. In some implementations, API components 200-24J include functionality performed by one or more applications of the device implementing agent system 200-20.
[0136] In some implementations, output components 200-26 include components for performing the output functions of the agent system 200-20. A brief description follows. Figure 5 The exemplary output components are illustrated herein. In some embodiments, output components 200-26 include... Figure 5 The components illustrated may include fewer components, more components, and / or different components. In some implementations, output components 200-26 are implemented in hardware and / or software.
[0137] like Figure 5 As illustrated, output components 200-26 include one or more visual output components 200-26A. One or more visual output components 200-26A may include any component used for outputting (e.g., generating, creating, and / or displaying) and / or causing visual output (e.g., visually perceptible output, such as a graphical user interface, playback of visual media content, and / or lighting). Examples of one or more visual output components 200-26A may include: display components, projectors, head-mounted display (HMD) devices, light-emitting diodes (“LEDs”), and / or components that create visually perceptible effects (e.g., movement). This list is not intended to be exhaustive, and one or more visual output components 200-26A may include other visual output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) the output visual output.
[0138] like Figure 5 As illustrated, output components 200-26 include one or more audio output components 200-26B. One or more audio output components 200-26B may include any component for outputting (e.g., generating and / or creating) and / or causing audio output (e.g., audibly perceptible output, such as sound, music, speech, and / or audio media content). Examples of one or more audio output components 200-26B may include: speakers, audio amplifiers, tone generators, and / or components that produce audibly perceptible effects (e.g., movement, such as vibration). This list is not intended to be exhaustive, and one or more audio output components 200-26B may include other audio output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) the output audio output.
[0139] like Figure 5As illustrated, output components 200-26 include one or more motion output components 200-26C (also referred to herein as "motion components"). One or more motion output components 200-26C may include any component for outputting (e.g., generating and / or creating) and / or causing motion output (e.g., output including physical movement of a device and / or another device / component). Examples of one or more motion output components 200-26C may include: motion controllers, actuators, mechanical linkages, electromechanical devices, and / or components that generate physical movement. This list is not intended to be exhaustive, and one or more motion output components 200-26C may include other motion output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output motion output. Figure 5 As illustrated, output components 200-26 include one or more haptic output components 200-26D. One or more haptic output components 200-26D may include any component for outputting (e.g., generating, creating, and / or displaying) and / or causing haptic output (e.g., output using haptically perceptible means, such as vibration, pressure, texture, and / or shape). Examples of one or more haptic output components 200-26D may include: speakers, components that generate vibrations, components that generate texture changes, components that generate pressure changes, and / or components that create perceptible haptic effects. This list is not intended to be exhaustive, and one or more haptic output components 200-26D may include other haptic output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output haptic output.
[0140] like Figure 5 As illustrated, output components 200-26 include one or more communication components 200-26E. The one or more communication components 200-26E may include any component (e.g., antenna, modem, network interface component, encoder, decoder, and / or communication protocol stack) for transmitting and / or receiving communications internal and / or external to the agent system 200-20. In some embodiments, communication may be between different devices and / or between components within the same device. In some embodiments, communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, the one or more communication components 200-26E include one or more features of one or more communication components 200-22B (e.g., as described above). In some embodiments, the one or more communication components 200-26E are identical to one or more communication components 200-22B (e.g., handling communication inputs and outputs and therefore considered as one or more of input and output components and / or both).
[0141] Throughout this disclosure, reference may be made to moving output (e.g., referred to in various forms such as: movement, device movement, moving output, device motion, motion output, and / or motion output). In some embodiments, output movement (e.g., an output that causes movement) refers to movement of an electronic device (e.g., a portion or component thereof relative to another portion and / or the entire electronic device). For example, refer again to Figure 2B The movable output can refer to the device 200 actuating the movable component 200-3 to move the display part 200-1 to... Figure 2B The illustrated location (e.g., from) Figure 2A (Positioning within). In some embodiments, the motion output is not (e.g., excluding and / or not only including) tactile output (e.g., tactile motion output). In some embodiments, the motion output is not (e.g., excluding and / or not only including) vibration output. In some embodiments, the motion output is not (e.g., excluding and / or not only including) oscillatory motion (e.g., movement of an actuator that causes vibration solely by repeatedly moving a component along a path within the device). In some embodiments, the motion output includes (e.g., requiring and / or causing) a change in the position and / or pose of at least a portion (and / or all) of a component or electronic device. In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device from a first position and / or a first pose to a second position and / or a second pose. For example, relative to Figures 2A to 2C ,exist Figure 2A , Figure 2B and Figure 2C In each of these embodiments, display portion 200-1 is shown in a different position (e.g., in space) and pose (e.g., relative to base portion 200-2). In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device to a third position and / or a third pose (e.g., from a first position and / or a first pose and / or from a second position and / or a second pose). In some embodiments, the third position and / or the third pose is the same as the first position and / or the first pose and / or the second position and / or the second pose. For example, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement is made to return to... Figure 2A The illustrated first positioning. For example, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement continues until it stops at... Figure 2C The third positioning illustrated.
[0142] Throughout this disclosure, electronic devices can be exemplified (and / or described) as being in different positions and / or poses at different times. For example, Figure 2A Example of device 200 in the first position, Figure 2B An example is shown of device 200 in the second position, and Figure 2A A device 200 in a third position is illustrated. In some embodiments, the electronic device moves itself between such positions and / or poses (e.g., using a movement output). For example, device 200 moves from a first position to a second position under its own power (e.g., using a power supply and one or more actuators to induce movement). Specifically, any examples of electronic devices illustrated and / or described herein in different positions and / or poses (e.g., at different times) should be understood to cover scenarios where the device moves itself between such positions and / or poses (e.g., unless otherwise explicitly stated).
[0143] Throughout this disclosure, reference may be made to “performing output,” “causing output,” and / or “output” (e.g., via one or more output generating devices and / or via one or more output generating components) (and / or similar phrases). In some embodiments, the output (e.g., or variations thereof) includes (and / or) output movement (e.g., moving the output as described above).
[0144] Throughout this disclosure, references may be made to “display,” “cause display,” and / or “output visual content” (e.g., via one or more display components) (and / or similar phrases). In some embodiments, display (e.g., or variations thereof) includes displaying visual content in conjunction with output movement (e.g., moving output as described above).
[0145] Throughout this disclosure, reference may be made to "output audio," "output that causes audio," and / or "provide audio output" (e.g., via one or more audio generation components and / or via one or more audio output devices) (and / or similar phrases). In some embodiments, outputting audio (e.g., or variations thereof) includes outputting audio content in conjunction with output movement (e.g., movement output as described above).
[0146] Throughout this disclosure, reference may be made to the movement (and / or similar phrases) of an avatar (e.g., or other representations of a displayed user, agent, and / or role) (e.g., via one or more display components). In some embodiments, moving an avatar (e.g., or a variant thereof) includes movement in conjunction with output movement (e.g., movement output as described above) to display visual content. For example, displaying an avatar nodding in agreement may include an electronic device moving in a manner similar to avatar movement (e.g., simulating a nod). In some embodiments, moving an avatar (e.g., or a variant thereof) includes movement with output movement (e.g., movement output as described above) without displaying visual content. For example, a device may perform a simulated nod without moving the displayed avatar's movement output (e.g., the avatar does not move relative to the display). Figure 5As illustrated, agent system 200-20 may optionally interface with external components such as external database 200-30, remote processing component 200-32, and / or remote management component 200-34. In some embodiments, external database 200-30 represents one or more functions that provide data storage resources accessible to agent system 200-20. In some embodiments, access to data in external database 200-30 is provided directly to agent system 200-20 (e.g., the agent system manages the database) and / or indirectly to agent system 200-20 (e.g., the database is managed by a different system, but the data stored therein can be provided and / or stored for use by agent system 200-20). In some embodiments, external database 200-30 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., is a database of web services accessible to different agent systems), and / or a combination of dedicated and non-dedicated database resources. In some embodiments, remote processing component 200-32 represents one or more components that serve as data processing resources accessible to agent system 200-20. In some embodiments, access to remote processing component 200-32 is provided directly to agent system 200-20 (e.g., the agent system manages the processing resources) and / or indirectly to agent system 200-20 (e.g., processing resources managed by a different system, but which can provide data processing for the benefit of agent system 200-20). In some embodiments, remote processing component 200-32 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., a processing resource of a web service accessible to a different agent system), and / or a combination of both dedicated and non-dedicated processing resources. Examples of data processing include processing image data (e.g., for feature extraction and / or object detection), processing audio data (e.g., for processing natural language speech input via a large language model), and / or training machine learning algorithms and / or models. In some implementations, remote management component 200-34 represents functions including management functions and / or functions related to management functions. For example, such management functions may include providing component updates (e.g., software and / or firmware updates) to agent system 200-30, managing accounts (e.g., associated licenses, access controls, and / or preferences), synchronizing between different agent systems and / or their components (e.g., enabling agents accessible via multiple devices of a user to provide a consistent user experience across such devices), managing collaboration with other services and / or agent systems, error reporting, managing backup resources to maintain agent system reliability and / or agent availability and / or other functions required for agent system 200-20 to perform operations, such as those described herein.
[0147] The above text is about Figure 5 The various components of the described intelligent agent system 200-20 represent functional blocks that represent functionality. This functionality can be implemented on the same and / or different hardware (e.g., physical components) and / or by the same and / or different software. For example, a functional block can be implemented using one or more physical components, devices (e.g., computer system 100 and / or electronic device 200), and / or software programs. In other words, each functional block does not necessarily represent a single, discrete physical component, device, and / or software program, but can be implemented using one or more of these. Furthermore, the intelligent agent system 200-20 may include multiple implementations of the functionality represented by the respective functional blocks. For example, the intelligent agent system 200-20 may include multiple different model components representing ML models used in different contexts, multiple different API components representing different APIs for different services, and / or multiple different visual output components for outputting different types of visual output.
[0148] Now let’s turn our attention to a discussion of the concepts that may arise regarding the operation of intelligent agents.
[0149] As discussed throughout, an intelligent agent may be able to interact with a user. In some implementations, this capability includes the ability to process explicit requests, commands, and / or statements. In some implementations, explicit requests, commands, and / or statements include and / or are interpreted as instructions relating to completing a task (e.g., displaying X, completing task Y, and / or performing operation Z). In some implementations, the intelligent agent includes the ability to process implicit requests, commands, and / or statements. In some implementations, implicit requests, commands, and / or statements do not include explicit requests, commands, and / or statements. For example, “I like to go to Europe” could be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays an itinerary in response to the statement. As another example, “This picture is for my grandmother” could be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays a suggestion to modify the picture. As another example, “I am tired” could be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 initiates a meditation session with a sleep meditation application. As another example, "I miss my grandfather" can be interpreted as an implicit request, command, and / or statement, and upon detection, device 200 can initiate a real-time communication session with the grandfather (e.g., a phone call, video call, and / or text messaging session). In some implementations, implicit requests are more likely to be processed based on one or more current contexts, operational contexts, and / or user contexts, while explicit requests are less likely to be processed based on one or more current contexts, operational contexts, and / or user contexts. For example, the phrase "Call my grandfather" can be an explicit request, and in response to detecting such a request, device 200 will initiate a real-time communication session with the grandfather regardless of one or more current contexts, operational contexts, and / or user contexts. However, the phrase "I miss my grandfather" can be an implicit request, and in response to detecting such a request, device 200 can display a list of gifts to buy for the grandfather if the user has recently been discussing buying gifts, or can call the grandfather in a different context that does not include the user's recent discussions about buying gifts. In some implementations, a request can include one or more explicit requests and one or more implicit requests. In some implementations, implicit requests are responded to independently of explicit requests; in other implementations, responses to implicit requests depend on explicit requests.
[0150] This document may refer to responses of an intelligent agent output by a device. In some embodiments, the response includes an audio component (e.g., audio output, audible output, sound and / or speech) (also referred to herein as a “verbal response,” “audio response,” and / or “audible response”) and / or a visual component (e.g., a display and / or movement of a representation and / or avatar). In some embodiments, the response includes a motion component (e.g., movement of the device). In some embodiments, the response includes a tactile component (e.g., touch and / or vibration).
[0151] This document may refer to internal dialogue, internal context, and / or operational context, which may refer to the dynamic context or dynamic decision-making process of a device, the internal state of device 200, and / or internal data of the device based in part on its decisions. In some embodiments, internal dialogue includes a set of one or more rules, characteristics, detections, and / or observations used by a computer system to generate responses to one or more commands, questions, and / or statements. In some embodiments, the set of one or more rules, characteristics, detections, and / or observations is learned and / or generated via deep learning and / or one or more machine learning algorithms and / or using one or more machine learning and / or system agents. In some embodiments, internal dialogue is generated in real time. In some embodiments, internal dialogue is stored locally and / or via cloud storage. In some embodiments, internal dialogue can be modified, updated, and / or deleted. In some embodiments, internal dialogue is generated based on other internal dialogues.
[0152] This document may refer to (e.g., the personality and / or behavior of an agent, user, and / or role) of a person or entity (or a representation of personality / behavior). In some embodiments, personality and / or behavior refers to one or more characteristics that a device detects, understands, conforms to, applies, and / or tracks. In some embodiments, personality or behavior is used as the basis for performing operations. For example, an agent may detect a user's personality and respond in a personality-based manner (e.g., outputting different responses in response to different user personalities). As another example, an agent may output responses having characteristics corresponding to one or more characteristics corresponding to personality and / or behavior (e.g., outputting responses in different ways depending on the agent's personality). In some embodiments, such characteristics represent and / or simulate a user's personality, such as how a user acts and / or speaks. In some embodiments, such characteristics approximate a user's personality.
[0153] In some embodiments, the intelligent agent is a system intelligent agent. In some embodiments, the system intelligent agent is an intelligent agent corresponding to an operating system originating from the device (e.g., the device implementing the intelligent agent) and / or a process controlled by the device's operating system. In some embodiments, the intelligent agent is an application intelligent agent. In some embodiments, the application intelligent agent is an intelligent agent corresponding to an application originating from the device (e.g., the device implementing the intelligent agent) (e.g., installed on and / or executed by the device) and / or a process controlled by the device's application.
[0154] This document may refer to representations (e.g., avatars and / or avatar representations) of agents (e.g., and / or users (e.g., people, objects, and / or animals) and / or user interface objects (e.g., animated characters)). In some embodiments, an agent's representation refers to a set of output characteristics (e.g., visual and / or audio) of the agent (and / or user and / or user interface object). For example, an agent's representation may include (and / or correspond to) a set of one or more visual characteristics (e.g., facial features of an animated face) and / or one or more audio characteristics (e.g., language and speech characteristics of audio output). In some embodiments, (e.g., the agent's) representation is used to represent the agent's output. For example, a device implementing an interactive agent outputs audio in the agent's voice and displays an animated face of the agent moving in a manner that simulates the agent speaking the audio output. In this way, the user can feel that they are having a normal conversation with the agent. In some embodiments, an agent's representation includes (or does not include) personality and / or behavioral characteristics (e.g., as described above). For example, an agent's representation may include (and / or correspond to) a set of visual characteristics (e.g., facial features of an animated face) and a set of personality characteristics. In some implementations, the representation of the intelligent agent includes a set of user characteristics corresponding to the user's visual representation (e.g., representations of the user's appearance, voice, and / or personality used as avatars that appear to move and / or speak). In some implementations, the representation is a facial representation (e.g., a user interface object outputting features that simulate a human face and / or facial expressions (e.g., for conveying information to a viewer)).
[0155] In some implementations, a role (e.g., the role of an agent and / or avatar) refers to a specific set of characteristics represented. For example, an avatar may embody the characteristics (e.g., use, application, interaction with, and / or output) of fictional and / or non-fictional characters (e.g., from movies, shows, books, TV series, and / or popular culture).
[0156] In some implementations, (e.g., agent and / or avatar) speech refers to one or more characteristics corresponding to sound outputs that are similar to (e.g., represent, imitate, and / or reproduce) spoken utterances (e.g., attributable to and / or simulated as output by an agent and / or avatar). For example, device 200 may output sentences that sound different depending on the speech used. In some implementations, a particular character and / or avatar may be configured to use a particular speech (e.g., have a corresponding speech). In some implementations, the particular speech may mimic a user's speech.
[0157] In some implementations, the appearance (e.g., of an agent and / or avatar) refers to a set of one or more characteristics corresponding to the visual output representing the avatar (and / or agent). For example, device 200 may output an avatar having a set of facial features that form an appearance similar to a specific character from a movie.
[0158] In some implementations, an avatar's expression refers to one or more characteristics corresponding to a specific visual appearance of the user, avatar, and / or agent. For example, device 200 may output an avatar having a set of facial features arranged in a specific manner to give the appearance of a facial expression (e.g., which can be used as a form of nonverbal communication to the user) (e.g., a frown is an expression of sadness, a smile is an expression of happiness, and / or wide eyes are an expression of surprise). As another example, device 200 may output an avatar having a set of body features (e.g., arms and / or legs) arranged in a specific manner to give the appearance of a body expression (e.g., which can be used as a form of nonverbal communication to the user) (e.g., a gesture is an expression of approval, covering the eyes is an expression of fear, and / or shrugging is an expression of lack of awareness). In some implementations, expressions include avatar movement (e.g., a nod is an expression of agreement and / or disagreement). In some implementations, device 200 may be movable via a movement component to indicate expressions with or without avatar movement. In some implementations, the agent performs one or more operations that depend on the user's facial expressions (e.g., detecting whether a person is sad and responding with a benevolent statement or question). In some implementations, facial expressions (e.g., whether and / or how to use and / or how to output) depend on personality. For example, a first-person sex may use more specific facial expressions than a second-person sex. As another example, a first-person sex's expressions (e.g., frowning, smiling, and / or widening eyes) may look different from a second-person sex's expressions (and / or similar and / or equivalent expressions) (e.g., a first-person sex smiles with their teeth showing, but a second-person sex smiles without showing their teeth).
[0159] In some implementations, an agent (e.g., an avatar of the agent and / or an agent system implementing the agent (e.g., hardware and / or software)) mimics the characteristics of another user, agent, and / or role (e.g., in terms of personality, behavior, facial expressions, and / or voice). In some implementations, mimicry includes mirroring the user (e.g., copying the use of phrases and / or movements detected from a user interacting with the agent). In some implementations, simulating user characteristics includes attempting to reproduce the user's characteristics (e.g., in exactly the same way and / or in a way that is similar to the characteristics but not an exact reproduction of the characteristics). For example, an agent mimicking voice and / or facial expressions does not require the agent to have exactly the same voice and / or facial expressions as the user being mimicked (e.g., simply to resemble the user's voice and / or facial expressions).
[0160] In some implementations, components and / or devices use (e.g., performing actions, making decisions, and / or determining context based on them) learned characteristics (e.g., characteristics of the context, user, and / or environment learned by the device over time (e.g., via detection, prior experience, and / or feedback (e.g., from one or more users))). For example, characteristics learned over time may include user routines. In such an example, if a particular user requests a summary of any new messages for that user from the agent at the same time every day, the agent may learn to automate actions based on the characteristics of the learned routines (e.g., what data is needed, when data is needed, and / or for which user). In some implementations, the learned characteristics enable the agent (and / or device) to improve its understanding (and / or response to) of the context, user, and / or environment, and / or its understanding of context, user, and / or environment that is otherwise not (and / or will not) understood (e.g., not responded to or responded to incorrectly). In some implementations, the learned characteristics are formed using reinforcement learning (e.g., by and / or for the agent). In some implementations, the learned features correspond to one or more confidence levels, determinisms, and / or rewards (e.g., shaped by one or more reward functions). In some implementations, the learned features (and / or how they are used to influence the output of the agent and / or device) can change over time (e.g., confidence levels, determinisms, and / or rewards change over time). For example, the output of a device before learning a set of learned features may differ from the output of a device after learning a set of learned features. In some implementations, components and / or devices use the learned knowledge. For example, similar to what is described above regarding learned features, learned knowledge may refer to information used to update (e.g., enhance, add to, and / or expand) the device's knowledge base (e.g., for use by agents implemented thereon). In some implementations, multiple sets of learned features for a user may be stored and / or used. In some implementations, different sets of learned features for different users may be stored and / or used.
[0161] This document may refer to interactions with an agent (and / or device). In some implementations, an interaction refers to a set of one or more inputs and / or outputs from a device implementing the agent and one or more users. For example, an interaction may be a user input (e.g., “Please turn on the light”) and a corresponding output (e.g., turning on the light and / or the device’s response “OK”). In some implementations, an interaction may include multiple inputs / outputs performed by one or more parties to the interaction (e.g., a device and / or a user). For example, an interaction may include a first user input (e.g., “Please turn on the light”) and a corresponding first output (e.g., “Which lights?”), and also include a second user input (e.g., “Kitchen light”) and a second output from the device (e.g., “OK”). In some implementations, which inputs and / or outputs are considered together as an interaction is based on logical and / or contextual grouping (e.g., interactions within the previous thirty (30) seconds and / or interactions related to turning on the light). As those skilled in the art will understand, interactions may be considered in an implementation-dependent manner (e.g., determining when an interaction is complete may involve determining whether the user is still present (e.g., is still talking) and / or whether the user is still talking about the light or has moved on to a different topic). In some implementations, the interaction is the current interaction (e.g., ongoing, currently occurring, and / or active). In some implementations, the interaction is a previous interaction. The examples above describe a device that engages in dialogue with a user. In some implementations, the dialogue is between two or more users (e.g., users in an environment). For example, the device may detect dialogue between users (e.g., users directing their voice and responses to each other, rather than to the device).
[0162] In some implementations, the agent (and / or device) determines and / or performs an action based on an intent corresponding to the user. For example, the device detects user input and outputs a response that depends on the intent of the user input. For example, the device detects user input including a pointing gesture detected along with a verbal command to “turn on the light,” and in response, the device turns on the light determined to correspond to the intent of the input (e.g., the light the pointing gesture is pointing to). In some implementations, one or more of the following are used to determine the intent (e.g., determined by the device that detects the input and / or by one or more other devices): one or more inputs, knowledge (e.g., knowledge about the user learned based on observed behavior, personality, and history of interactions), learned characteristics, and / or context. In some implementations, the intent is determined based on one or more types of input (e.g., verbal input, visual input via a camera, and / or contextual input).
[0163] Now turn our attention to the implementation of user interfaces (“UIs”) and associated processes on electronic devices (such as computer system 100 and / or electronic device 200).
[0164] Figures 6A to 6E Exemplary user interfaces for performing movement based on learned characteristics, according to some embodiments, are illustrated. The user interfaces in these figures are used to illustrate the processes described below, including... Figure 7 The process in.
[0165] Figures 6A to 6E The computer system 600 is illustrated as a tablet computer displaying different user interfaces. It should be understood that the computer system 600 can be other types of computer systems, such as smartphones, smartwatches, laptops, public facilities, smart speakers, accessories, personal gaming systems, desktop computers, fitness trackers, and / or head-mounted display (HMD) devices. In some embodiments, the computer system 600 includes one or more sensors (e.g., one or more cameras, one or more LiDAR detectors, one or more motion sensors, one or more infrared sensors, and / or one or more microphones) and / or communicates with them. In some embodiments, the computer system 600 includes one or more output devices (e.g., displays, projectors, touch-sensitive displays, and / or speakers) and / or communicates with them. In some embodiments, the computer system 600 includes one or more moving components (e.g., actuators, movable bases, rotatable components, and / or rotatable bases) and / or communicates with them. In some embodiments, the computer system 600 includes one or more components and / or features described above with respect to devices 100 and / or 200.
[0166] Figures 6A to 6E An example is illustrated of a scenario in which computer system 600 learns user behaviors (e.g., a set of movements (e.g., facial and / or body) and / or sounds) and characteristics within a learning context. When in the same or similar context, computer system 600 outputs a representation of the learned behaviors (e.g., movements and / or sounds). Movements included in the representation are movements learned by the user that have similar speeds, expressions, and / or directions and / or sets of orientations.
[0167] In the examples described below, the representation of user movement output by computer system 600 includes computer system 600 moving UI elements (e.g., avatars) via a display. In some embodiments, the representation of user movement output by computer system 600 includes computer system 600 moving a portion of computer system 600 via one or more moving components. In some embodiments, the representation of user movement output by computer system 600 includes computer system 600 moving UI elements and portions of computer system 600 (e.g., display portions and / or hardware components, such as buttons and / or rotatable input mechanisms) in a coordinated manner (e.g., moving at the same speed, in the same direction, and / or at the same rhythm). In some embodiments, references to moving at the same speed, direction, rhythm, beat, and / or tempo include movement at similar speed, direction, rhythm, beat, and / or tempo and / or movement generated using a common speed, direction, rhythm, beat, and / or tempo.
[0168] like Figure 6A As illustrated, computer system 600 displays a virtual assistant user interface 602. The virtual assistant user interface 602 includes a virtual assistant avatar 604, which computer system 600... Figure 6A The virtual assistant avatar is displayed in the center of the virtual assistant user interface 602. In this example, the computer system 600 displays the virtual assistant avatar 604 occupying most of the virtual assistant user interface 602. In some embodiments, the computer system 600 displays the virtual assistant avatar 604 at different sizes and / or different locations within the virtual assistant user interface 602. In some embodiments, the virtual assistant avatar 604 is a human-like visual representation of a virtual assistant application and / or artificial intelligence application that visually changes based on content output and / or input and reacts to and / or responds to input. In some embodiments, the virtual assistant avatar 604 corresponds to a system process of the computer system 600 (e.g., managed, controlled, output, created, and / or requested by it). In this example, the virtual assistant avatar 604 is a visual representation of a human and / or animal face. In some implementations, the virtual assistant avatar 604 has different appearances (e.g., different colors (e.g., color groups, skin tones, red, orange, yellow, green, blue and / or purple), textures (e.g., skin, hair, fur, scales, plastic, glass, feathers and / or wood), accessories (e.g., hats, glasses, monocles, wands, books, collars, bows, wings, halos and / or crowns) and / or facial types (e.g., human, animal, anthropomorphic object, alien, ordinary face, fantasy creature and / or a collection of similar faces)). Figure 6A A time indicator 608 is also shown to indicate the passage of time in the diagram described below.
[0169] Figure 6A Within a first context, computer system 600 and user 610 are illustrated (e.g., the user responds to something, the user performs an action, and / or computer system 600 outputs something). In this example, the first context includes computer system 600 detecting a user (e.g., 610) within the environment and outputting a greeting in response to detecting the user. In some embodiments, the first context includes computer system 600 outputting a notification. For example, the first context may include computer system 600 outputting the current score of a game being played by the user's favorite team. In some embodiments, the first context includes computer system 600 detecting a user reacting to a detectable situation. For example, the first context may include computer system 600 detecting a user dancing while certain music is playing. In some embodiments, the first context includes computer system 600 detecting a user performing an action. For example, the first action may include computer system 600 detecting a user encouraging themselves before a meeting.
[0170] exist Figure 6A In the context of computer system 600, user 610 is detected within the environment. For example... Figure 6A As illustrated, in response to the detection of user 610 within the environment, computer system 600 outputs a first action. Figure 6AIn response to detecting that user 610 and computer system 600 output a first action (e.g., a first set of movements and / or sounds), it is determined that computer system 600 is operating within a first context. In some embodiments, the first action is a default action corresponding to the first context. In some embodiments, the default action is a behavior preset (e.g., previously configured) by the publisher of the virtual assistant user interface 602 (e.g., or otherwise defaulted to, rather than based on user input and / or the content being output). In some embodiments, the default action is a user-preset action. In some embodiments, the first action is a previously learned characteristic corresponding to the first context. In this example, the first action corresponds to a greeting to the detected user. In this example, the first action includes computer system 600 outputting a first audio, which includes audio output 606 (e.g., "Hello"). In some embodiments, the first audio is a default audio corresponding to the first context. In some embodiments, the first audio is a previously learned audio corresponding to the first context. In some embodiments, the first action includes computer system 600 outputting a first movement (e.g., user interface elements moving via a display and / or a portion of computer system 600 moving via one or more moving components). For example, the first movement may include the computer system 600 moving a portion of itself in an up-and-down motion similar to a bow. In some embodiments, the first movement is a default movement corresponding to a first context. In some embodiments, the first movement is a previously learned movement corresponding to a first context. In some embodiments, the first action includes a first audio and a first movement. In some embodiments, the first action includes the computer system 600 moving the virtual assistant avatar 604 in sync with the first audio to achieve the appearance of the virtual assistant avatar 604 giving a greeting.
[0171] exist Figure 6B In this example, computer system 600 detects a second action (e.g., a second set of movements and / or sounds) from user 610 corresponding to a response to the first action. In this example, the second action includes second audio and second movement. Figure 6B As illustrated, the second audio input includes audio input 605B1 (e.g., "What's wrong?"). Similarly, as... Figure 6BAs illustrated, the second movement includes movement input 605B2, whereby user 610 raises and then nods in a greeting manner (e.g., a nod). In some embodiments, computer system 600 detects the second audio before detecting the second movement. In some embodiments, computer system 600 detects the second audio after detecting the second movement. In some embodiments, computer system 600 detects both the second audio and the second movement simultaneously. In some embodiments, the second action includes only the second movement. In some embodiments, the second action includes only the second audio. In this example, the second movement is a movement of the user's head. In some embodiments, the second movement is a movement of the user's face. For example, the second movement could be smiling and blinking. In some embodiments, the second movement is a movement of the user's body. For example, the second movement could be the user tilting their shoulder at an angle. Figure 6B In response to determining that the computer system 600 is operating in a first context and detecting a second action from the user 610, the computer system 600 learns a representation of the second action but does not execute the second action.
[0172] exist Figure 6C In the middle, the time after the second behavior is detected, such as Figure 6C The middle shows with Figure 6B The time indicator 608 indicates different times, and the computer system 600 detects users 610 and 616. Figure 6C In response to detecting user 610 and / or user 616, it is determined that computer system 600 is operating in a second context. Figure 6C In this context, the second context is determined to correspond to the first context. For example... Figure 6C As illustrated, in response to determining that computer system 600 is operating in a second context and that the second context corresponds to the first context, computer system 600 outputs a third action (e.g., a third set of movements and / or sounds), which is derived from... Figure 6B The representation of the second behavior of computer system 600 detected by user 610.
[0173] In some implementations, the first context and the second context are identical, leading to the determination that the second context corresponds to the first context. For example, the computer system 600 detecting and greeting user 610 in both the first and second contexts may lead to the determination that the second context corresponds to the first context. In some implementations, the first context and the second context are different, leading to the determination that the second context does not correspond to the first context. For example, the computer system 600 including a first context that outputs a calendar notification and a second context that includes the computer system 600 outputting a social media feed may lead to the determination that the second context does not correspond to the first context. In some implementations, the first context and the second context are different, but sufficiently similar in one or more respects to lead to the determination that the second context corresponds to the first context. For example, the computer system 600 including a first context that outputs a phone call from the user's parents and a second context that includes the computer system 600 outputting a text notification from the user's parents may lead to the determination that the second context corresponds to the first context. In some implementations, the first context and the second context include the same input. For example, the computer system 600 may detect the user performing the same air gesture in both the first and second contexts. In some implementations, the first context and the second context include similar types of input. For example, computer system 600 can detect audio input corresponding to a question about an upcoming birthday in a first context and audio input corresponding to a question about an upcoming anniversary in a second context.
[0174] In some implementations, in response to determining that a second context does not correspond to a first context, computer system 600 outputs a fourth action (e.g., a fourth set of movements and / or sounds). For example, in response to determining that a second context does not correspond to a first context, computer system 600 may output a fourth action that is a user's indication when the user is confused. In some implementations, the fourth action corresponds to the second context. For example, in response to outputting a notification of the winning score of the user's favorite team (e.g., a second context that does not correspond to detecting and greeting the user (e.g., the first context)), the second context computer system 600 may output audio as an indication that the user is saying "yes," while moving part of computer system 600 in the form of movement as an indication that the user is performing a celebratory dance. In some implementations, in response to determining that a second context does not correspond to a first context, computer system 600 does not output a set of movements and / or sounds.
[0175] In some implementations, the user in the first context is different from the user in the second context. In some implementations, the difference between the user in the first context and the user in the second context leads to the determination that the second context does not correspond to the first context. For example, in response to detecting user 616 (e.g., instead of user 610) while in the second context, it can be determined that the second context corresponds to a third context, which can cause computer system 600 to output a fifth action (e.g., a fifth set of movements and / or sounds) corresponding to user 616 (e.g., a representation of user 616's learned behavior). In some implementations, the second context is determined to correspond to the first context even if the user in the first context is different from the user in the second context. For example, in response to detecting user 616 (e.g., instead of user 610) while in the second context, it can be determined that the second context corresponds to the first context because both contexts involve detecting and greeting the user, which can cause computer system 600 to output a third action.
[0176] In this example, the third line includes the computer system 600 outputting a third audio and a third motion. For example... Figure 6C As illustrated, the third audio input includes an audio output 614, which is a representation of the audio input 605B1 (e.g., "What's wrong?"). Figure 6C As illustrated, computer system 600 outputs audio (e.g., first audio) in a second context that is the same as the audio output by computer system 600 in a first context. Figure 6A (As shown) different audio (e.g., the third audio).
[0177] Figures 6D to 6E An example is illustrated where computer system 600 outputs a third movement. In this example, computer system 600 outputting a third movement includes animates virtual assistant avatar 604 via a display in a manner representing movement input 605B2. In some embodiments, computer system 600 outputting a third movement includes moving a portion of computer system 600. For example, computer system 600 outputting a third movement may include moving a portion of computer system 600 up and then down in a movement representing movement input 605b.
[0178] like Figure 6D As illustrated, after the computer system 600 outputs audio output 614, the computer system 600 moves the virtual assistant avatar 604 upward within the virtual assistant user interface 602, which is a representation of the user 610 raising their chin in the movement input 605B2 (e.g., as shown). Figure 6B(As shown). In some embodiments, the computer system 600 outputs audio output 614 when it moves the virtual assistant avatar 604 upward within the virtual assistant user interface 602. In some embodiments, the computer system 600 outputs audio output 614 after it has moved the virtual assistant avatar 604 upward within the virtual assistant user interface 602.
[0179] like Figure 6E As illustrated, after computer system 600 moves virtual assistant avatar 604 upward within virtual assistant user interface 602, computer system 600 moves virtual assistant avatar 604 downward within virtual assistant user interface 602. Figure 6C The position within the virtual assistant user interface 602, which is the indication of the user 610 nodding in the movement input 605B2 (e.g. Figure 6B (As shown). In some embodiments, the computer system 600 outputs audio output 614 when the virtual assistant avatar 604 is moved down within the virtual assistant user interface 602. In some embodiments, the computer system 600 outputs audio output 614 after the virtual assistant avatar 604 has been moved down within the virtual assistant user interface 602.
[0180] In some embodiments, in response to receiving communication (e.g., voice call, video call, and / or text) from another user, computer system 600 executes a sixth action (e.g., a sixth set of movements and / or sounds) corresponding to the other user (e.g., a representation of the other user's learned behavior). In some embodiments, the sixth action is included in the user profile of another user on another computer system. In some embodiments, computer system 600 remembers the sixth action from previous interactions with other users. In some embodiments, the sixth action is configured by the master user of computer system 600 (e.g., the user with the most privileges) to correspond to the settings of the other user.
[0181] Figure 7 This is a flowchart illustrating a method using a computer system to perform a representation of a movement based on learned characteristics, according to some embodiments. Method 700 is performed at a computer system (e.g., 100, 200, and / or 600). Some operations in method 700 may be combined, some operations may be ordered differently, and some operations may be omitted.
[0182] As described below, Method 700 provides an intuitive way to perform movement representations based on learned characteristics. This method reduces the cognitive burden on users interacting with computer systems, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to interact with computer systems faster and more efficiently, saving power and increasing the time between battery charging.
[0183] In some implementations, method 700 is used with one or more input devices (e.g., as described above regarding...). Figures 6A to 6B This is performed at a computer system (e.g., 600) that communicates with (e.g., a touch-sensitive display, a rotatable input mechanism, a camera (e.g., a telephoto, wide-angle, and / or ultra-wide-angle camera) and / or a sensor (e.g., a gyroscope and / or a heart rate sensor)). In some embodiments, the computer system communicates with display components (e.g., a display screen, a projector, and / or a touch-sensitive display) and / or one or more output devices (e.g., a speaker, a haptic output device, a display screen, a projector, and / or a touch-sensitive display). In some embodiments, the computer system is a watch, a phone, a tablet, a fitness tracker, a processor, a head-mounted display (HMD) device, a public utility, a media device, a speaker, a television, and / or a personal computing device. In some embodiments, the computer system communicates with moving components (e.g., an actuator, a motor, an electro-arm, a lift, and / or a lever). In some embodiments, the computer system communicates with one or more cameras (e.g., a telephoto, wide-angle, and / or ultra-wide-angle camera).
[0184] The computer system detects (702) a first user's (e.g., 610) relative to a first context (e.g., when operating in the first context, when being in the first context, and / or when the first context is detected) via one or more input devices (e.g., as...). Figures 6A to 6B Movements described herein (e.g., body movements (e.g., head, shoulder, arm and / or finger) and / or facial movements (e.g., lips, eyes, eyelids, mouth and / or eyebrows)) (e.g., user responses to stimuli, user location (e.g., outdoors, indoors, home, school, car, concert, museum, theater, place of religion and / or gym), computer system output content, topics discussed by the user (e.g., sports, travel, current events, school, work and / or hobbies) and / or user actions).
[0185] In response to detecting movement of the first user (e.g., 610) relative to the first context, the computer system abandons (704) performing (e.g., projection, display, animation, rasterization, and / or rendering) a representation of the first user's movement (e.g., the actual movement of the first user and / or one or more characteristics that approximate the same characteristics as the first user's movement) (e.g., as described above regarding...). Figures 6A to 6B (As described).
[0186] After detecting the movement of the first user relative to the first context, the computer system detects (706) the transition to the second context (e.g., as described above regarding...). Figure 6C(as described) (e.g., user responses to stimuli, user location (e.g., outdoors, indoors, at home, at school, in a car, at a concert, in a museum, at a theater, at a place of religion, and / or at a gym), computer system output, topics discussed by the user (e.g., sports, travel, current events, school, work, and / or hobbies), and / or actions performed by the user).
[0187] In response to detecting a shift to a second context and based on determining that the second context corresponds to the first context (and / or satisfies a predetermined confidence threshold for semantic meaning, syntactic meaning, similarity, relation, connection, and / or relevance), the computer system executes (708) a representation of the first user's movement (e.g., the first user's actual movement and / or one or more characteristics that are similar to the first user's movement) (e.g., as described above regarding...). Figures 6C to 6E (As described). In some implementations, the computer system does not render the first user's movement based on the determination that the second context does not correspond to the first content (and / or does not meet a predetermined confidence threshold for semantic meaning, syntactic meaning, similarity, relation, connection, and / or relevance). Representing the first user's movement based on satisfying predetermined conditions enables the computer system to replicate the user's movement according to the current situation and / or context, thereby providing the user with improved visual feedback and performing actions without further user input when a set of conditions are met.
[0188] In some implementations, in response to detecting a shift to a second context and determining that the second context does not correspond to the first context (and / or does not meet a predetermined confidence threshold for semantic meaning, syntactic meaning, similarity, relation, association, connection, and / or relevance), the computer system abandons the execution of the first user's movement representation (e.g., as described above regarding...). Figures 6A to 6E (As described). The representation that the first user's movement is not performed based on the fulfillment of specified conditions prevents the computer system from replicating the user's movement according to the current situation and / or context, thereby providing the user with improved visual feedback and performing an operation without further user input when a set of conditions are met.
[0189] In some implementations, the second context is the same as the first context (e.g., as described above regarding...). Figures 6A to 6E (As described). Making the second context identical to the first context enables the computer system to learn and replicate the user's movements relative to the same situation and / or context, thereby providing the user with improved visual feedback and / or performing actions when a set of conditions are met without requiring further user input.
[0190] In some implementations, the second context is different from the first context (e.g., as mentioned above regarding...). Figures 6A to 6E(As described). Distinguishing the second context from the first context enables the computer system to learn and replicate the user's movements relative to different situations and / or contexts, thereby providing the user with improved visual feedback and / or performing actions when a set of conditions are met without requiring further user input.
[0191] In some embodiments, the computer system (e.g., 600) communicates with a microphone. In some embodiments, detecting movement of a first user (e.g., 610) relative to a first context includes receiving speech input from the first user via the microphone (e.g., 605B1). In some embodiments, detecting a shift to a second context includes receiving speech input from a second user (e.g., 616) different from the first user (e.g., as described above regarding...). Figures 6A to 6E (As described). In some embodiments, the verbal input from the second user differs from the verbal input from the first user. In some embodiments, detecting a transition to a second context does not include detecting the movement of the second user and / or the movement of the first user. A representation of performing the first user's movement based on satisfying predetermined conditions enables the computer system to replicate user movement when responding to different users, thereby providing the user with improved visual feedback and performing actions when a set of conditions are met without requiring further user input.
[0192] In some embodiments, detecting movement of a first user (e.g., 610) relative to a first context includes receiving a first input via at least one of one or more input devices (e.g., 605B1). In some embodiments, detecting a transition to a second context includes receiving a second input different from the first input via at least one of one or more input devices (e.g., 605B1) (e.g., as described above regarding...). Figures 6A to 6E (As described above). In some implementations, the second input and the first input are the same type of input (e.g., as described above regarding...). Figures 6A to 6E (As described) (e.g., air gestures, mouse clicks, taps, voice input, swipes, keyboard input, and / or rotation of a rotatable input mechanism). A representation of performing a first user's movement based on satisfying predetermined conditions enables the computer system to replicate the user's movement when responding to the same input, thereby providing the user with improved visual feedback and performing an operation without further user input when a set of conditions are met.
[0193] In some implementations, detecting movement of a first user (e.g., 610) relative to a first context includes detecting third input via one or more input devices (e.g., 605B1) (e.g., as described above regarding...). Figures 6A to 6E(As described above). In some implementations, detecting a transition to a second context includes detecting a fourth input different from the third input via one or more input devices (e.g., 605B2) (e.g., as described above regarding...). Figures 6A to 6E (As described above). In some implementations, the third and fourth inputs are different types of inputs corresponding to the operation (and / or the same operation) (e.g., air gestures, mouse clicks, taps, voice input, swipes, keyboard input, and / or rotation of a rotatable input mechanism) (e.g., as described above regarding...). Figures 6A to 6E (As described). In some implementations, the third input is the input that causes the computer system to perform an operation. In some implementations, the fourth input is another type of input that causes the computer system to perform the same operation. The representation of performing a first user's movement based on satisfying predetermined conditions enables the computer system to replicate the user's movement based on the detection of different types of input, thereby providing the user with improved visual feedback and performing an operation when a set of conditions are met without requiring further user input.
[0194] In some implementations, based on determining that the movement of a first user (e.g., 610) relative to a first context has a first set of one or more movement characteristics, the representation of the first user's movement is a representation of a second set of one or more movement characteristics (e.g., as described above regarding...). Figures 6A to 6E (As described above). In some embodiments, based on determining that the movement of the first user (e.g., 610) relative to the first context has a third set of one or more movement characteristics that are different from the first set of one or more movement characteristics, the representation of the first user's movement is a representation of a fourth set of one or more movement characteristics (e.g., as described above regarding...). Figures 6A to 6E(As described). In some embodiments, the second set of one or more motion characteristics is the same as and / or generated based on the first set of one or more motion characteristics. In some embodiments, the second set of one or more motion characteristics is different from and / or not generated based on the third set of one or more motion characteristics. In some embodiments, the second set of one or more motion characteristics is different from and / or not generated based on the fourth set of one or more motion characteristics. In some embodiments, the fourth set of one or more motion characteristics is different from and / or not generated based on the first set of one or more motion characteristics. In some embodiments, the fourth set of one or more motion characteristics is different from and / or not generated based on the second set of one or more motion characteristics. In some embodiments, the fourth set of one or more motion characteristics is the same as and / or generated based on the third set of one or more motion characteristics. In some embodiments, the computer system displays (and / or projects) a representation of a first user's movement, which includes one or more characteristics of the first user's movement relative to a first context. Making the representation of the first user's movement a representation of the second set of one or more motion characteristics or a representation of the fourth set of one or more motion characteristics based on the satisfaction of predetermined conditions enables the computer system to replicate a particular user movement, thereby providing the user with improved visual feedback and performing an operation without further user input when a set of conditions is satisfied.
[0195] In some embodiments, the first group of one or more motion characteristics and the second group of one or more motion characteristics have a first acceleration (and / or deceleration) characteristic. In some embodiments, the third group of one or more motion characteristics and the fourth group of one or more motion characteristics have a second acceleration characteristic that differs from the first acceleration characteristic (e.g., as described above regarding...). Figures 6A to 6E (As described). In some embodiments, the representation of the first user's movement is faster than the first user's movement relative to the first context and has the same constant time-scale invariance between events as the first user's movement relative to the first content. In some embodiments, the representation of the first user's movement is slower than the first user's movement relative to the first context and has the same constant time-scale invariance between events in the first user's movement relative to the first content. Imposing a first set of one or more movement characteristics and a second set of one or more movement characteristics with a first acceleration, and imposing a third set of one or more movement characteristics and a fourth set of one or more movement characteristics with a second acceleration, enables the computer system to replicate a specific speed of user movement, thereby providing the user with improved visual feedback and performing an action when a set of conditions is met without requiring further user input.
[0196] In some embodiments, the first group of one or more movement characteristics and the second group of one or more movement characteristics have a first directional characteristic (e.g., north, south, east, west, left, right, up, and / or down) (e.g., and / or movement in a specific set of directions (e.g., left, then right, then up, and / or then down)). In some embodiments, the third group of one or more movement characteristics and the fourth group of one or more movement characteristics have a second directional characteristic different from the first directional characteristic (e.g., north, south, east, west, left, right, up, and / or down) (e.g., and / or movement in a specific set of directions (e.g., left, then right, then up, and / or then down)) (e.g., as described above regarding...). Figures 6A to 6E (As described). By giving a first group of one or more movement characteristics and a second group of one or more movement characteristics a first directional characteristic, and by giving a third group of one or more movement characteristics and a fourth group of one or more movement characteristics a second directional characteristic, the computer system is able to replicate a specific speed of user movement, thereby providing the user with improved visual feedback and performing an operation when a set of conditions are met without requiring further user input.
[0197] In some embodiments, one or more movement features in the first group and one or more movement features in the second group have first facial expression features (e.g., smiling, frowning, grinning, open mouth, shy expression, happy expression, sad expression, nodding, expression of agreement, expression of disagreement, and / or expression of disdain). In some embodiments, one or more movement features in the third group and one or more movement features in the fourth group have second facial expression features that differ from the first facial expression features (e.g., smiling, frowning, grinning, open mouth, shy expression, happy expression, sad expression, nodding, expression of agreement, expression of disagreement, and / or expression of disdain) (e.g., as described above regarding...). Figures 6A to 6E (As described). By giving a first group of one or more movement characteristics and a second group of one or more movement characteristics a first facial expression characteristic, and by giving a third group of one or more movement characteristics and a fourth group of one or more movement characteristics a second facial expression characteristic, the computer system is able to reproduce the user's facial expressions via movement, thereby providing the user with improved visual feedback and performing an operation when a set of conditions are met without requiring further user input.
[0198] In some embodiments, the computer system (e.g., 600) communicates with a first moving component (e.g., an actuator (e.g., a pneumatic actuator, a hydraulic actuator, and / or an electric actuator), a movable base, a rotatable component, and / or a rotatable base). In some embodiments, the representation of movement by a first user (e.g., 610) includes movement via the first moving component (e.g., tilting, rotating, moving right, left, up, down, vertically, horizontally, inward, and / or outward) of a first part of the computer system (e.g., a physical part, a part of the display component, the center of the display, and / or another part of the display, and / or hardware components (e.g., hardware buttons and / or rotatable input mechanisms)) (e.g., as described above regarding...). Figures 6A to 6E (As described).
[0199] In some implementations, the computer system (e.g., 600) communicates with a display component. In some implementations, representing the movement of a first user (e.g., 610) includes displaying animations via the display component (e.g., as described above regarding...). Figures 6C to 6E As described above). In some embodiments, the animation showing the movement of a portion of a user interface element (e.g., 602) includes: at a first moment, moving a portion (e.g., face, eyes, eyebrows, mouth, lips, torso, and / or shoulders) of a user interface element (e.g., avatar, optional user interface object, text, symbol, button, and / or content) in a first direction (e.g., face, eyes, eyebrows, mouth, lips, torso, and / or shoulders, as described above). Figures 6C to 6E As described above); and at a second time different from the first time, in a second direction different from the first direction, a portion of the user interface element is moved (e.g., as described above regarding...). Figures 6C to 6E (As described). The display animation includes moving a portion of the user interface element in a first direction and moving that portion of the user interface element in a second direction when predetermined conditions are met, enabling the computer system to replicate the user's facial expressions via movement, thereby providing the user with improved visual feedback and performing an operation without further user input when a set of conditions are met.
[0200] In some implementations, user interface elements (e.g., 602) include representations of faces (e.g., 604) (e.g., a face with a nose, eyes, mouth, and / or eyebrows) (e.g., as described above regarding...). Figures 6A to 6E (As described). Including facial representations in user interface elements enables computer systems to replicate facial expressions made and changed by the user, thereby providing the user with improved visual feedback and performing actions when a set of conditions are met without requiring further user input.
[0201] In some embodiments, the computer system communicates with a second moving component (e.g., an actuator (e.g., a pneumatic actuator, a hydraulic actuator, and / or an electric actuator), a movable base, a rotatable component, and / or a rotatable base). In some embodiments, the representation of movement by a first user (e.g., 610) includes moving (e.g., tilting, rotating, right, left, up, down, vertically, horizontally, inward, and / or outward) a second part of the computer system (e.g., a physical part, a portion of the display component, the center of the display, and / or another portion of the display, and / or hardware components (e.g., hardware buttons and / or rotatable input mechanisms)) via the first moving component and simultaneously displaying animations of user interface elements (e.g., as described above regarding...). Figures 6A to 6E (As described).
[0202] In some implementations, portions of the computer system move and portions of the user interface elements are displayed as moving at the same speed (e.g., as described above regarding...). Figures 6A to 6E (As described). In some implementations, directionality is determined based on the speed of movement of the first user.
[0203] In some implementations, portions of the computer system move and portions of the user interface elements are displayed as moving in the same direction (e.g., as described above regarding...). Figures 6A to 6E (As described). In some implementations, directionality is determined based on the direction of movement of the first user.
[0204] In some implementations, portions of the computer system move and portions of the user interface elements are displayed at the same rhythm (e.g., as mentioned above). Figures 6A to 6E The movement is described (e.g., beat, jitter, jerking, and / or movement pattern). In some implementations, the rhythm is determined based on the rhythm of the first user's movement.
[0205] In some implementations, after detecting a movement of a first user (e.g., 610) relative to a first context, the computer system detects a transition to a third context different from the first and second contexts. In some implementations, in response to detecting a transition to the third context, the computer system abandons the representation of the first user's (e.g., 610) movement (e.g., as described above regarding...). Figures 6A to 6E (As described). In some implementations, a representation of the first user's movement is executed in response to the detection of a transition to a second context and based on the determination that the second context corresponds to a third context. Executing a representation of the first user's movement without responding to the detection of a transition to a third context enables the computer system to provide improved visual feedback to the user and to perform an operation without requiring further user input when a set of conditions are met, in certain situations, without replicating the user's movement.
[0206] In some implementations, in response to detecting a transition to a third context, the computer system provides an output that differs from the representation of movement performed by the first user (e.g., 610) (e.g., as described above regarding...). Figures 6A to 6E (As described). Providing an output that differs from the representation of the movement performed by the first user in response to a detected transition to a third context enables the computer system to provide other types of input while operating in other contexts, thereby providing improved visual feedback to the user and performing operations without further user input when a set of conditions are met.
[0207] In some implementations, the output includes a representation of a second movement that performs a movement different from that of the first user (e.g., 610) (e.g., as described above regarding...). Figures 6A to 6E (As described). The output includes a representation of a second movement that performs a movement different from that of the first user, enabling the computer to replicate the different movements for the user while the computer system operates in other contexts, thereby providing the user with improved visual feedback and performing actions when a set of conditions are met without requiring further user input.
[0208] In some implementations, in response to detecting movement of a first user (e.g., 610) relative to a first context, the computer system performs an operation different from the representation of the first user's movement (e.g., as described above regarding...). Figures 6A to 6E (As described).
[0209] In some implementations, actions that differ from the representation of the first user's movement (e.g., 610) are not performed based on the detected (e.g., previously detected and / or currently detected) movement of the first user (e.g., as described above regarding...). Figures 6A to 6E (As described).
[0210] In some implementations, after representing the movement of a first user (e.g., 610), the computer system detects a movement of a third user relative to a fourth context that is different from the first user's movement. In some implementations, the third user's movement is the same as the first user's movement (e.g., as described above regarding...). Figures 6A to 6E (As described above). In some embodiments, in response to detecting movement of a third user relative to a fourth context, the computer system executes a representation of the third user's movement without executing a representation of the first user's movement (e.g., 610). In some embodiments, the representation of the third user's movement differs from the representation of the first user's movement (e.g., as described above regarding...). Figures 6A to 6E(As described). Detecting movement of a third user relative to a fourth context and, in response to detecting movement of the third user relative to the fourth context, executing a representation of the third user's movement without executing a representation of the first user's movement, enables the computer system to respond to multiple users in different ways, thereby providing improved visual feedback to users and performing operations when a set of conditions are met without requiring further user input.
[0211] In some implementations, after representing the movement of the first user (e.g., 610), the computer system detects a movement of a fourth user relative to a fifth context that is different from the first user's. In some implementations, the fourth user's movement is the same as the first user's movement (e.g., as described above regarding...). Figures 6A to 6E (As described above). In some implementations, in response to detecting movement of a fourth user relative to a fifth context, the computer system executes a representation of the first user's movement (e.g., as described above regarding...). Figures 6A to 6E (As described). Detecting the movement of a fourth user relative to a fifth context (the same as the movement of the first user) and, in response to detecting the movement of the fourth user relative to the fifth context, executing a representation of the first user's movement, enables the computer system to provide the same output to multiple users, thereby providing improved visual feedback to users and performing operations when a set of conditions are met without requiring further user input.
[0212] In some implementations, after representing the movement of the first user (e.g., 610), the computer system detects incoming communication from a fifth user different from the first user (e.g., from a second computer system different from the first user) (e.g., a telephone call, text message, and / or video call) (e.g., as described above regarding...). Figures 6A to 6E (As described above). In some implementations, in response to detecting (and / or receiving) incoming communication, the computer system performs a representation of the movement of a fifth user (e.g., as described above regarding...). Figures 6A to 6E (As described) (e.g., simulating the behavior of an incoming communication caller). In some embodiments, the representation of the fifth user's movement is the same as that of the first user's movement. In some embodiments, the representation of the fifth user's movement differs from that of the first user's movement. Detecting incoming communication from the fifth user and executing the representation of the fifth user's movement in response to the detection of incoming communication enables the computer system to provide customized output for various communications, thereby providing the user with improved visual feedback and / or reducing the amount of input required to perform an operation.
[0213] In some implementations, before representing the movement of the first user (e.g., 610), the computer system detects a move towards a sixth context (e.g., as described above regarding...). Figures 6A to 6EThe transition (as described above) is a transition of the first user's movement (e.g., a first context and / or another context, such as a second context or a context different from the second context). In some implementations, in response to the detection of a sixth context, based on determining that the first user's movement (e.g., 610) has been detected more than a threshold number of times (e.g., 1 to 1000 times) before the transition to the sixth context, the computer system executes a representation of the first user's movement (e.g., as described above regarding...). Figures 6A to 6E (As described above). In some implementations, in response to the detection of a sixth context, based on the determination that the number of moves by the first user (e.g., 610) detected before transitioning to the sixth context is less than a threshold number, the computer system abandons the execution of the representation of the first user's moves (e.g., as described above regarding...). Figures 6A to 6E (As described). The representation of whether or not to execute the first user's movement based on the satisfaction of specified conditions enables the computer system to replicate the user's movement only after multiple detections of movement execution, thereby providing the user with improved visual feedback and performing operations without further user input when a set of conditions are met.
[0214] It should be noted that the above text regarding method 700 (for example, Figure 7 The details of the process described herein also apply in a similar manner to the methods described below / above. For example, method 900 may optionally include one or more characteristics of the various methods described above with reference to method 700. For example, a computer system may perform a representation of movement using the techniques described with respect to method 900 in response to detecting a gesture and a representation of movement induced by the techniques described with respect to method 700. For the sake of brevity, these details will not be repeated below.
[0215] Figures 8A to 8E Exemplary user interfaces for configuring operations to be performed based on learned characteristics, according to some implementation schemes, are illustrated. The user interfaces in these figures are used to illustrate the processes described below, including... Figure 9 The process in.
[0216] Figures 8A to 8EThe computer system 800 is illustrated as a tablet computer displaying different user interfaces. It should be understood that the computer system 800 can be other types of computer systems, such as smartphones, smartwatches, laptops, public facilities, smart speakers, accessories, personal gaming systems, desktop computers, fitness trackers, and / or head-mounted display (HMD) devices. In some embodiments, the computer system 800 includes one or more sensors (e.g., one or more cameras, one or more LiDAR detectors, one or more motion sensors, one or more infrared sensors, and / or one or more microphones) and / or communicates with them. In some embodiments, the computer system 800 includes one or more output devices (e.g., displays, projectors, touch-sensitive displays, and / or speakers) and / or communicates with them. In some embodiments, the computer system 800 includes one or more moving components (e.g., actuators, movable bases, rotatable components, and / or rotatable bases) and / or communicates with them. In some embodiments, the computer system 800 includes one or more components and / or features described above with respect to electronic devices 100, 200, and / or 600.
[0217] Figures 8A to 8E This illustrates a scenario where a computer system 800 learns air gestures from a user and associates them with system responses and / or inputs. The computer system 800 detects the user's air gestures and inputs via one or more sensors, and in response, performs an operation. In this scenario, the computer system 800 configures the operation performed by the input to be executed when the computer system 800 detects an air gesture but not the input.
[0218] like Figure 8A As illustrated, computer system 800 displays a text user interface 802, including displaying first text 804 within the center of the text user interface 802 and occupying a large portion of the text user interface. In some embodiments, computer system 800 displays the first text 804 in some other size and / or at some other location within the text user interface 802. For example... Figure 8A As illustrated, the first text 804 corresponds to a request to transmit text (e.g., "Should I transmit text?"). In Figure 8AIn this example, computer system 800 detects a first gesture from user 806. In this example, the first gesture is an air gesture 805a, which is a fist. In some embodiments, air gesture 805a is some other air gesture (e.g., palm up, palm down, thumbs up, clapping, flicking, pointing, pinching, waving, and / or the OK gesture). In some embodiments, the first gesture is some other form of input (e.g., gaze input and / or body movement). In this example, air gesture 805a is not a default gesture (e.g., not pre-configured to correspond to an action). In some embodiments, air gesture 805a is a default gesture (discussed in more detail below). Figure 8A In the process, it was determined that the computer system 800 did not recognize the air gesture 805a as input corresponding to the operation.
[0219] like Figure 8B As illustrated, in response to determining that the computer system 800 does not recognize the air gesture 805a as input corresponding to an operation, the computer system 800 continues to display the first text 804. In some embodiments, in response to determining that the computer system 800 does not recognize the air gesture 805a as input corresponding to an operation, the computer system 800 briefly stops displaying the first text 804 and then displays the first text 804 again. In some embodiments, in response to determining that the computer system 800 does not recognize the air gesture 805a as input corresponding to an operation, the computer system 800 changes the display of the first text 804 to enhance its visibility.
[0220] exist Figure 8B In this example, computer system 800 detects a first input corresponding to the computer system 800 performing a first operation. In this example, the first input is audio input 805b corresponding to an affirmative response (e.g., "yes") to first text 804. Figure 8B In this process, computer system 800 continues to detect air gesture 805a (e.g., a first gesture). In some embodiments, computer system 800 detects the first input before detecting the first gesture. In some embodiments, computer system 800 detects the first input after detecting the first gesture. In some embodiments, computer system 800 detects the first input and the first gesture simultaneously.
[0221] like Figure 8C As illustrated, in response to the detection of audio input 805b, computer system 800 performs a first operation. In this example, computer system 800 performing the first operation includes computer system 800 stopping the display of first text 804 and displaying second text 808, which corresponds to an acknowledgment that text has been transmitted. In some embodiments, computer system 800 performing the first operation includes computer system 800 transmitting text corresponding to the question in first text 804. Figure 8C In response to detecting an air gesture 805a along with audio input 805b, the computer system 800 configures the air gesture 805a (e.g., a first gesture) to correspond to the computer system 800 performing a first operation when no audio input 805b (e.g., the first input) is detected. Figure 8C In this process, computer system 800 no longer detects air gestures 805a from user 806.
[0222] In some implementations, computer system 800 requires detecting a first gesture along with a first input at least twice (e.g., a threshold number of times) before configuring the first gesture to correspond to performing a first operation by computer system 800 (e.g., the same operation as the first input corresponds to the operation performed by computer system 800) without detecting the first input. In some implementations, computer system 800 requires detecting different gestures along with their corresponding inputs at different numbers of times before configuring gestures to be used in the absence of corresponding inputs. For example, computer system 800 may require detecting common gestures (e.g., clapping, pinching, swiping, and / or raising a thumb) along with their corresponding gestures more frequently than uncommon gestures (e.g., moving the back of the hand forward, raising the little finger, and / or scooping).
[0223] In some embodiments, the first operation includes launching a new application (e.g., computer system 800 displays an application). In some embodiments, the first operation includes closing an application (e.g., computer system 800 stops displaying an application). In some embodiments, the first operation is a media control (e.g., play, pause, skip forward, skip backward, play next, and / or mark as favorites). For example, computer system 800 performing the first operation may include computer system 800 pausing media currently being output by computer system 800. In some embodiments, the first operation is a communication control (e.g., starting, answering, and / or ending a telephone and / or video call). For example, computer system 800 performing the first operation may include computer system 800 starting a video call.
[0224] In some embodiments, computer system 800 recognizes a first gesture as input corresponding to a second operation (e.g., not the first operation). In some embodiments, in response to detecting the first gesture but not detecting audio input 805b, computer system 800 performs the second operation. In some embodiments, air gesture 805a is a default gesture. For example, the first gesture could be a default gesture corresponding to "no," and in response to detecting the first gesture, computer system 800 could stop displaying text 804 and display the statement "Text not transmitted." In some embodiments, the second gesture is a default gesture corresponding to computer system 800 performing the first operation. For example, in response to detecting a thumbs-up air gesture (e.g., the second gesture), computer system 800 can perform the first operation. In some embodiments, in response to detecting air gesture 805a along with audio input 805b, when computer system 800 recognizes air gesture 805a as corresponding to the second operation and audio input 805b as corresponding to the first operation, computer system 800 performs the first operation. In some implementations, after performing a first operation in response to detecting audio input 805b, when computer system 800 detects air gesture 805a and identifies it as corresponding to a second operation, and identifies audio input 805 as corresponding to the first operation, computer system 800 configures air gesture 805a to correspond to the first operation, and the scenario is as follows: Figure 8D The description continues.
[0225] like Figure 8D As illustrated, after displaying the second text 808, the computer system 800 displays a text user interface 802 that includes the first text 804. Figure 8D In this example, computer system 800 detects an air gesture 805d from user 806. In this example, air gesture 805d is the same as air gesture 805a. Figure 8D In the process, the computer system 800 recognizes the air gesture 805d as input corresponding to the first operation.
[0226] like Figure 8E As illustrated, in response to the computer system 800 detecting an air gesture 805d and recognizing the air gesture 805d as input corresponding to a first operation, the computer system 800 performs a first operation including stopping the display of first text 804 and displaying second text 808. In some embodiments, the computer system 800 performing the first operation includes the computer system 800 transmitting text corresponding to a question in the first text 804.
[0227] In some implementations, if the computer system 800 operates in different modalities, detecting an air gesture 805d will not cause the computer system 800 to perform an operation. In some implementations, in response to detecting an air gesture 805d when the computer system 800 is displaying an online shopping interface, the computer system 800 does nothing. In some implementations, in response to detecting an air gesture 805d while in some modalities, the computer system 800 performs a third operation. For example, if the computer system 800 is displaying an e-book application, in response to detecting an air gesture 805d, the computer system 800 may add a bookmark to the currently displayed page. In some implementations, the computer system 800 detecting an air gesture 805d in some modalities causes the computer system 800 to perform an operation similar to the first operation. For example, if the computer system 800 is displaying text corresponding to a request to send an email while displaying an email user interface (e.g., similar to displaying text corresponding to a request to send text, such as...),... Figure 8D (As illustrated), in response to the detection of an air gesture 805d, the computer system 800 stops displaying the text corresponding to the request and displays the text corresponding to the confirmation that an email has been sent.
[0228] In some implementations, in response to the detection of a third gesture along with audio input 805b (e.g., an input corresponding to the computer system 800 performing a first operation), the computer system 800 performs the first operation and configures the third gesture to correspond to the computer system 800 performing the first operation without detecting audio input 805b. For example, if the computer system 800 detects an air gesture input with the palm facing forward and all fingers spread, and also detects audio input 805b, the computer system 800 may perform the first operation and configure the air gesture input with the palm facing forward and all fingers spread to correspond to the computer system 800 performing the first operation. In some implementations, after configuring the third gesture to correspond to the computer system 800 performing the first operation, in response to the detection of an air gesture 805d, the computer system 800 performs the first operation, indicating that the air gesture 805d is still configured to correspond to the first operation. In some implementations, more than one gesture may be configured to perform the same operation. In some implementations, after a third gesture is configured to correspond to a first operation, the computer system 800 does not perform the first operation in response to the detection of an air gesture 805d, indicating that the air gesture 805d is no longer configured to correspond to the first operation. In some implementations, only one gesture can be configured to perform an operation at a time, and when a new gesture is configured to perform an operation, the previous gesture is no longer configured to perform the operation without initial input. For example, after an air gesture input with the palm facing forward and all fingers spread is configured to correspond to the computer system 800 performing a first operation, the computer system 800 will no longer perform the first operation because it detects the air gesture 805a and does not simultaneously detect audio input 805b.
[0229] In some implementations, computer system 800 detects a fourth gesture along with a second input corresponding to the computer system 800 performing a fourth operation, the fourth operation causing computer system 800 to configure the fourth gesture to correspond to the computer system 800 performing the fourth operation. For example, if computer system 800 detects an air gesture of two fingers swiping down and to the left along with audio input corresponding to computer system 800 opening the initial text for editing while computer system 800 is displaying text 804, then computer system 800 can open the initial text for editing and configure the air gesture of two fingers swiping down and to the left to correspond to computer system 800 opening the initial text for editing. In some implementations, computer system 800 can recognize different gestures to correspond to computer system 800 performing different operations in the same modality. For example, when computer system 800 is displaying text 804, computer system 800 can recognize a first gesture to correspond to computer system 800 performing a first operation and a fourth gesture to correspond to computer system 800 performing a fourth operation.
[0230] In some implementations, the computer system 800 will not allow gestures to be configured to correspond to actions performed by the computer system 800 if no corresponding input is detected. For example, when a prompt to change security settings is displayed, the computer system 800 may ignore the detected gesture along with the input to change security settings. In some implementations, the computer system 800 will not recognize gestures that have been configured to correspond to actions performed by the computer system 800; instead, the computer system 800 will require detection of the initial corresponding input. For example, the computer system 800 may ignore gestures corresponding to deleting files when an entire drive is selected; instead, the computer system 800 may require audio input corresponding to deleting files to perform the action.
[0231] In this example, the computer system 800 performing the first operation does not include movement (e.g., moving a UI element via a display and / or moving a portion of the computer system 800). In some embodiments, the computer system 800 performing the first operation includes movement (e.g., moving a UI element via a display and / or moving a portion of the computer system 800). For example, the computer system 800 performing the first operation may include the computer system 800 moving a volume slider on a music user interface. In this example, air gesture 805d (e.g., 805d and 805a) does not include movement. In some embodiments, air gesture 805d (e.g., 805d and 805a) includes movement. In some embodiments, if air gesture 805d includes movement and the first operation does not include movement, the amount of movement detected by the computer system 800 in air gesture 805d does not affect how the computer system 800 performs the first operation. In some embodiments, if both air gesture 805d and the first operation include movement, the amount of movement detected by the computer system 800 in air gesture 805d affects how the computer system 800 performs the first operation. For example, if the first operation is scrolling up a document within the user interface and the air gesture 805d includes raising the index finger, then the computer system 800 detects that the degree and speed at which the index finger is raised affects the degree and speed at which the computer system 800 scrolls up the document within the user interface.
[0232] Figure 9 This is a flowchart illustrating a method for configuring operations to be performed using a computer system based on learned characteristics, according to some implementation schemes. Method 900 is executed at a computer system (e.g., 100, 200, 600). Some operations in method 900 may be combined, some operations may be rearranged, and some operations may be omitted.
[0233] As described below, Method 900 provides an intuitive way to configure operations to be performed based on learned characteristics. This method reduces the cognitive burden on users when configuring operations based on learned characteristics, thereby creating a more efficient human-computer interface. For battery-powered computing devices, enabling users to configure operations based on learned characteristics more quickly and efficiently saves power and increases the time between battery charging cycles.
[0234] In some embodiments, method 900 is performed at a computer system (e.g., 600) that communicates with one or more input devices (e.g., touch-sensitive displays, rotatable input mechanisms, cameras (e.g., telephoto, wide-angle, and / or ultra-wide-angle cameras) and / or sensors (e.g., gyroscopes and / or heart rate sensors)). In some embodiments, the computer system communicates with display components (e.g., displays, projectors, and / or touch-sensitive displays) and / or one or more output devices (e.g., speakers, haptic output devices, displays, projectors, and / or touch-sensitive displays). In some embodiments, the computer system is a watch, phone, tablet, fitness tracker, processor, head-mounted display (HMD) device, public utility, media device, speaker, television, and / or personal computing device. In some embodiments, the computer system communicates with motion components (e.g., actuators, motors, electro-arms, lifts, and / or levers). In some embodiments, the computer system communicates with one or more cameras (e.g., telephoto, wide-angle, and / or ultra-wide-angle cameras).
[0235] The computer system detects (902) a first gesture (e.g., 805A) (e.g., body posture, user movement, thumbs up, thumbs down, hand gesture, and / or karate chop) via one or more input devices and combines (e.g., simultaneously, immediately before, and / or immediately after) a first input (e.g., 805B) (e.g., verbal input, air gesture, touch input (e.g., tap input, swipe input, and / or long press input) and / or gaze input) (e.g., as described above regarding...). Figures 8A to 8B (As described).
[0236] In response to (904) the detection of a first input (e.g., 805B) and the detection of a first gesture (e.g., 805A), the computer system performs (906) a first operation (e.g., corresponding to the first input) (e.g., as described above regarding...). Figures 8A to 8E (As described). In some embodiments, the first input corresponds to a request to perform a corresponding operation. In some embodiments, the first operation is an action to be performed by the computer system (e.g., starting an application, displaying an application, and / or stopping the display of an application) and / or causing that action to be performed by an external computer system.
[0237] In response to (904) the detection of a first gesture along with the detection of a first input, the computer system configures (908) a first operation to be performed in the absence of the first input (e.g., 805B) (and / or in response to the detection of user movement) (e.g., as described above regarding...). Figures 8B to 8C (As described).
[0238] After performing the first operation and after configuring the first operation to be performed without detecting the first input (e.g., 805B), the computer system detects (910) the first pose (e.g., 805A) via one or more input devices (e.g., another instance of the first pose and / or a second pose independent of the first pose (e.g., detected at different times and / or in different contexts) (e.g., after the computer system has detected a pose that is different from and / or not of the same type as the first pose) without detecting the first input (e.g., 805B) (e.g., as described above regarding...). Figures 8D to 8E (As described).
[0239] In response to the detection of a first gesture (e.g., 805A) but the absence of a first input (e.g., 805B), the computer system performs (912) a first operation (e.g., as described above regarding...). Figures 8D to 8E (As described). After configuring the first operation to be performed in the absence of a first input, the first operation is performed in response to a detected gesture but no first input, enabling the computer system to automatically configure the operation to be performed in the absence of input, and allowing the computer system to learn the user's gesture and associate the gesture with the system response based on the gesture and the input associated with the system response, thereby reducing the amount of input required to perform the operation and allowing the computer system to automatically associate the user's personal gesture with the system response, thereby providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without further user input, and providing improved feedback to the user.
[0240] In some implementations, before configuring the first operation to be performed if the first input is not detected (e.g., 805B), the computer system detects a third gesture (e.g., the same as or different from the first gesture) via one or more input devices along with detecting the first input (e.g., as described above). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a third gesture along with the detection of a first input (e.g., 805B) and based on the determination that the first operation is not configured to be performed in the absence of the first input, the computer system performs a second operation (e.g., as described above regarding...). Figures 8A to 8E(As described). In some embodiments, the second operation differs from the first operation. In some embodiments, the second operation is the same as the first operation. In some embodiments, the computer system was not previously configured to perform the first operation in the absence of detected input before configuring the first operation to be performed in the absence of detected input. In some embodiments, the computer system performs the second operation because the second operation was not configured to be performed in the absence of detected first input. In some embodiments, in response to the detection of a third gesture along with the detection of the first input and based on the determination that the first operation is configured to be performed in the absence of detected first input, the computer system performs the first operation but not the second operation. Performing the second operation in response to the detection of a third gesture along with the detection of the first input and based on the determination that the first operation is not configured to be performed in the absence of detected first input enables the computer system to perform additional operations prompted by the user, thereby providing additional control options without cluttering the user interface with additional displayed controls, performing operations without requiring further user input, and providing improved feedback to the user.
[0241] In some implementations, before configuring the first operation to be performed if the first input is not detected (e.g., 805B), the computer system detects a fourth gesture (e.g., the same as or different from the third gesture) via one or more input devices along with detecting the first input (e.g., as described above). Figures 8A to 8E (As described above). In some embodiments, in response to the detection of a fourth gesture along with the detection of a first input (e.g., 805B) and based on the determination that the first operation is not configured to be performed in the absence of the first input, the computer system abandons the execution of the first operation (and in some embodiments, abandons the execution of additional and / or different operations) (e.g., as described above regarding...). Figures 8A to 8E (As described). In some implementations, in response to the detection of a fourth gesture along with the detection of a first input and based on the determination that a first operation is configured to be performed in the absence of a first input, the computer system performs the first operation. In response to the detection of a fourth gesture along with the detection of input and based on the determination that the first operation is not configured to be performed in the absence of a first input, the first operation is not performed, enabling the computer system to automatically perform the operation without detecting input when the computer system is configured to detect a gesture but there is no input to trigger the operation. This provides additional control options without cluttering the user interface with additional displayed controls, performs the operation without requiring further user input, and provides improved feedback to the user.
[0242] In some implementations, the computer system detects the first pose without detecting the first input (e.g., 805A) before detecting the first pose (e.g., 805B) along with detecting the first input (e.g., 805B). Figures 8A to 8E (As described above). In some embodiments, in response to the detection of a first gesture (e.g., 805A) but the absence of a first input (e.g., 805B), the computer system abandons the execution of the first operation (and in some embodiments, abandons the execution of additional and / or different operations) (e.g., as described above regarding...). Figures 8A to 8E (As described). In some implementations, detecting a first gesture does not cause the system to perform an operation (e.g., the first operation and / or additional and / or different operations) even when the first input is not detected. Not performing the first operation in response to detecting a first gesture but not detecting the first input allows the computer system to do so only if previously configured to perform an operation even when no input is detected, thereby providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0243] In some implementations, the computer system detects the first gesture but not the first input via one or more input devices before detecting the first gesture (e.g., 805A) along with the first input (e.g., 805B) (e.g., as described above). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a first gesture (e.g., 805A) but the absence of a first input (e.g., 805B), the computer system performs a third operation different from the first operation (e.g., as described above regarding...). Figures 8A to 8E (As described). In some implementations, detecting a first gesture causes the computer system to perform an operation even when the first input is not detected. Responding to the detection of the first gesture but the absence of the first input, performing a third operation different from the first operation enables the computer system to perform different operations when the computer system is not configured to automatically perform different operations. This provides additional control options without cluttering the user interface with additional displayed controls, performs the operation without requiring further user input, and provides improved feedback to the user.
[0244] In some implementations, before configuring the first operation to be performed if the first input (e.g., 805B) is not a previously stored type of pose (e.g., not the default pose, not a pose recorded by the system, and / or not a pose registered by the system) (e.g., as described above regarding...). Figures 8A to 8E(As described). In some embodiments, the first gesture is a second type of gesture (e.g., a customized type, a user (e.g., a subject, person, object, and / or animal), a specific gesture, and / or a user-acquired characteristic) that differs from the first type of gesture. After configuring the first operation to be performed in the absence of a first input, the first operation is performed in response to the detection of a gesture but the absence of a first input, and the first gesture is not a first type of gesture that was previously stored before configuring the first operation to be performed in the absence of a first input. This enables the computer system to automatically configure the operation to be performed in response to a gesture that has not been previously registered, without prompting the user to perform the operation, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0245] In some implementations, prior to detecting the first gesture (e.g., 805A) along with detecting the first input (e.g., 805B), the computer system detects a fifth gesture (e.g., a gesture different from the first gesture, e.g., one with a different type and / or characteristic) via one or more input devices. This fifth gesture may be independent of the first gesture and / or a different type of gesture (e.g., both gestures involve a fist and a thumb (e.g., the first gesture is a raised thumb and / or the seventh gesture is a thumbs-down), both gestures involve waving (e.g., the first gesture is a wave and / or the seventh gesture is a wave from a different position than the first gesture), and / or both gestures involve head movement (e.g., the first gesture is a nod (e.g., affirmation), and the seventh gesture is a head movement upwards)). The fifth gesture is a previously stored type of gesture (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to detecting a fifth pose of a previously stored type pose, the computer system performs a first operation (e.g., as described above regarding...). Figures 8A to 8E (As described). A first operation is performed in response to the detection of a fifth gesture of a previously stored first type of gesture, enabling the computer system to perform operations in response to input previously stored on the computer system, thereby providing additional control options without cluttering the user interface with additional displayed controls, performing operations without requiring further user input, and providing improved feedback to the user.
[0246] In some embodiments, the computer system communicates with a first display component. In some embodiments, performing the first operation includes displaying, via the first display component, a representation (e.g., graphical representation, image, user interface element, button, and / or function representation) of content (e.g., application content, media content, and / or symbolic content) that was not previously displayed on the user interface (e.g., application content, media content, and / or symbolic content). Figures 8A to 8E(As described). In some embodiments, performing the first operation includes moving content, changing content, removing content, and / or emphasizing and / or de-emphasizing content (e.g., facial representations, application content, media content, and / or symbolic content). In some embodiments, performing the first operation includes launching an application. In some embodiments, content is displayed on a user interface (e.g., a lock screen, a home screen, a user interface displayed when the computer system is locked (e.g., a state requiring a password and / or other information before the computer system can be unlocked and / or a state that is more secure, less functional, and / or includes less information than another operational state of the computer system), an unlocked screen user interface, and / or a user interface displayed via a first display component when the computer system is unlocked and / or unlocked). Performing an operation including displaying a representation of content not previously displayed on the user interface enables the computer system to display additional information without prompting the user to perform the operation, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing the user with improved visual feedback.
[0247] In some embodiments, the computer system (e.g., 600) communicates with a second display component. In some embodiments, before the detection of a first gesture (e.g., 805A) along with the detection of a first input (e.g., 805B) (and in some embodiments, during and / or after the detection of the first gesture), the computer system displays corresponding content (e.g., one or more user interface objects (e.g., user interface elements, representations of software applications, avatars, system avatars, menus, and / or buttons)) (e.g., application content, media content, and / or symbol content) via the second display component. In some embodiments, performing the first operation includes stopping the display of the corresponding content via the second display component (e.g., as described above regarding...). Figures 8A to 8E (As described). In some implementations, the corresponding content is initially displayed on a second user interface (e.g., a lock screen, home screen, and / or a user interface displayed when the computer system is in a locked state (e.g., a state where a password and / or other information is required before the computer system can be transitioned to an unlocked state, and / or a state that is more secure, less functional, and / or includes less information than another operational state of the computer system)). Displaying the corresponding content before detecting the first gesture along with detecting the first input and stopping the display of the corresponding content allows the computer system to automatically configure the different content to be displayed, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing the user with improved visual feedback.
[0248] In some implementations, after performing the first operation (e.g., in response to detecting a second gesture but not detecting the first input) (and in some implementations, together with or without detecting the first input), the computer system detects the first gesture via one or more input devices (e.g., 805A) (and in some implementations, together with or without detecting the first input) (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some embodiments, in response to the detection of a first gesture (e.g., 805A) (and in some embodiments, together with the detection or non-detection of the first input), the computer system performs a fourth operation different from the first operation (e.g., as described above regarding...). Figures 8A to 8E (As described). Responding to the detection of a first gesture to perform a fourth operation enables the computer system to perform different operations based on the same prompt from the user, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0249] In some implementations, the first operation is a first-type operation (e.g., displaying, outputting, and / or playing back media), and the fourth operation is a first-type operation (e.g., as described above regarding...). Figures 8A to 8E (As described). In some embodiments, the first type of operation is to output content in a certain way. In some embodiments, the first operation is to output a first medium, and the fourth operation is to output a second medium different from the first medium. The fourth operation, performed in response to detecting a first gesture, enables the computer system to perform a different operation of the same type based on the same gesture from the user, thereby providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0250] In some implementations, the first and fourth operations are the same operation (e.g., as described above regarding...). Figures 8A to 8E (As described). In some implementations, the first operation is to output a first medium, and the fourth operation is to output the first medium. The fourth operation, performed in response to detecting a first gesture, enables the computer system to perform a different operation of the same type based on the same gesture from the user, thereby providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0251] In some implementations, after (and / or before) performing the first operation (e.g., in response to detecting the first gesture but not the first input, or in response to detecting the first gesture along with detecting the first input), the computer system detects a sixth gesture different from the first gesture (e.g., body posture, user movement, thumbs up, thumbs down, waving, and / or karate chop) via one or more input devices (e.g., independent of the first gesture and / or a different type of gesture (e.g., both gestures involve a fist and a thumb (e.g., the first gesture is a thumbs up, the seventh gesture is a thumbs down), both gestures involve waving (e.g., the first gesture is a waving gesture, the seventh gesture is a waving gesture in a different position than the first gesture), and / or both gestures involve head movement (e.g., the first gesture is a nod, the seventh gesture is a head movement upwards)) (e.g., 805A) along with detecting a second input (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some embodiments, the second input is different from the first input. In some embodiments, the first input is the same as the second input. In some embodiments, in response to the detection of a sixth gesture in conjunction with the detection of the second input, based on determining that the sixth gesture has been detected more than a predetermined number of times (e.g., 1 to 50 times) in conjunction with the detection of the second input, the computer system performs a fifth operation (e.g., corresponding to the second input) (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some embodiments, the fifth operation is an action to be performed by the computer system (e.g., starting an application, displaying an application, and / or stopping the display of an application) and / or causing that action to be performed by an external computer system. In some embodiments, in response to the detection of a sixth gesture along with the detection of a second input, based on determining that the sixth gesture has been detected more than a predetermined number of times along with the detection of the second input, the computer system configures the fifth operation to be performed even when the second input is not detected (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a sixth gesture along with the detection of a second input, based on determining that the sixth gesture, along with the detection of the second input, has not exceeded a predetermined number of times, the computer system performs a fifth operation (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a sixth gesture along with the detection of a second input, and based on the determination that the sixth gesture has not been detected more than a predetermined number of times along with the detection of the second input, the computer system abandons configuring the fifth operation to be performed in the absence of a detected second input (e.g., as described above regarding...). Figures 8A to 8E (As described).
[0252] In some implementations, before configuring the first operation to be performed if the first input is not detected (e.g., 805B), the computer system detects a seventh posture (e.g., independent of the first posture and / or a different type of posture (e.g., both postures involve a fist and a thumb (e.g., the first posture is a raised thumb, the seventh posture is a thumbs-down thumb), both postures involve waving (e.g., the first posture is a wave, the seventh posture is a wave at a different position than the first posture), and / or both postures involve head movement (e.g., the first posture is a nod, the seventh posture is a head-up movement)) via one or more input devices (e.g., along with the detection and / or non-detection of the first input) (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a seventh gesture (e.g., together with the detection of the first input and / or the non-detection of the first input), the computer system performs a first operation (e.g., as described above regarding...). Figures 8A to 8E (As described). Responding to the detection of a seventh posture to perform a first operation enables the computer system to automatically perform the same operation for multiple different postures without requiring input, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0253] In some implementations, after the first operation is configured to be performed if the first input is not detected (e.g., 805B) (and in some implementations, when the first operation is configured to be performed if the first input is not detected), the computer system detects a seventh gesture via one or more input devices (e.g., together with the detection and / or non-detection of the first input) (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a seventh gesture but the absence of a first input (e.g., 805B), the computer system abandons the execution of the first operation (e.g., as described above regarding...). Figures 8A to 8E(As described). In some embodiments, the first operation is configured to be performed if the first input is not detected, such that the seventh gesture is not configured to cause the computer system to perform the first operation if the first input is detected. In some embodiments, the first operation is configured to be performed if the first input is not detected, such that the seventh gesture is not configured to cause the computer system to perform the first operation if the first input is not detected. In some embodiments, the first operation is configured to be performed if the first input is not detected, such that the computer system does not perform the first operation when the computer system detects the seventh gesture. Not performing the first operation in response to detecting the seventh gesture but not detecting the first input prevents the computer system from automatically performing the operation if no input is detected, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0254] In some implementations, after performing the first operation, the computer system detects an eighth gesture (e.g., independent of the first gesture and / or a different type of gesture, such as both gestures involving a fist and a thumb (e.g., the first gesture is a raised thumb, the seventh gesture is a thumbs-down thumb), both gestures involving a wave (e.g., the first gesture is a wave, the seventh gesture is a wave at a different position than the first gesture), and / or both gestures involving head movement (e.g., the first gesture is a nod, the seventh gesture is a head-up movement)) via one or more input devices, and combines this with the detection of a third input (e.g., independent of the first input and / or a different type of input, such as the first input is pressing a button and the second input is pressing a different button; and / or the first input is tapping and the second input is tapping a different position)) (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some embodiments, in response to the detection of an eighth gesture along with a third input, the computer system performs a sixth operation different from the first operation (e.g., independent of the first operation and / or a different type of operation (e.g., the first operation is launching an application and the sixth operation is causing the computer system to launch the application separately; and / or the first operation is taking a picture and the sixth operation is displaying the home screen user interface)) (and in some embodiments, the first operation is not performed) (e.g., as described above regarding Figures 8A to 8E (As described above). In some implementations, in response to the detection of an eighth gesture along with a third input, the computer system configures a sixth operation to be performed even when the first input is not detected (e.g., 805B) (e.g., as described above regarding...). Figures 8A to 8E(As described above). In some implementations, after performing the sixth operation, the computer system detects a ninth posture via one or more input devices, wherein the ninth posture is the same as the eighth posture (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a ninth gesture but the absence of a third input, the computer system performs a sixth operation (e.g., as described above regarding...). Figures 8A to 8E (As described) (and in some embodiments, the first operation is not performed). In some embodiments, after performing the sixth operation, the computer system detects a corresponding posture different from the ninth posture via one or more input devices. In some embodiments, in response to detecting a corresponding posture but not detecting a third input, the computer system performs a corresponding operation different from the sixth operation. By configuring the sixth operation to be performed in the absence of a first input, and then performing the sixth operation in response to detecting the ninth posture but not detecting a third input, the computer system is able to configure multiple different operations for detecting different postures and perform different operations without corresponding input, thereby reducing the number of inputs required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without further user input, and providing improved feedback to the user.
[0255] In some implementations, after performing the first operation, the computer system detects a tenth gesture (e.g., independent of the first gesture and / or a different type of gesture (e.g., both gestures involve a fist and a thumb (e.g., the first gesture is a raised thumb, the seventh gesture is a thumbs-down thumb), both gestures involve waving (e.g., the first gesture is a wave, the seventh gesture is a wave at a different position than the first gesture), and / or both gestures involve head movement (e.g., the first gesture is a nod, the seventh gesture is a head-up movement)) via one or more input devices, along with detecting a fourth input (e.g., as described above regarding the first input (e.g., 805B)) that is different from the first input (e.g., 805B). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a tenth gesture along with a fourth input, the computer system performs a seventh operation different from the first operation (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of the tenth gesture along with the fourth input, the computer system configures the seventh operation to be performed even if the fourth input is not detected (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, after performing the seventh operation, the computer system detects an eleventh gesture via one or more input devices, wherein the eleventh gesture is the same as the tenth gesture (e.g., as described above regarding...). Figures 8A to 8E(As described above). In some implementations, in response to the detection of an eleventh gesture but the absence of a fourth input, the computer system abandons the execution of the seventh operation (e.g., as described above regarding...). Figures 8A to 8E (As described) (and in some embodiments, the seventh operation is abandoned in the absence of a fourth input). In some embodiments, in response to the detection of a gesture (e.g., a tenth gesture, an eleventh gesture, and / or a different gesture) but the absence of an input (e.g., a fourth input and / or a different input), the computer system abandons the execution of the seventh operation. In some embodiments, in response to the detection of a gesture but the absence of an input, the computer system abandons the execution of the seventh operation and abandons the configuration of the seventh operation to be executed in the absence of a fourth input. In some embodiments, certain operations (e.g., the seventh operation) cannot be configured to be performed in response to both a gesture and an input. In some embodiments, the computer system is not allowed to be configured to perform an operation in the absence of an input but only a gesture. After not configuring the seventh operation to be executed in the absence of a fourth input, the seventh operation is not executed in response to the detection of an eleventh gesture but the absence of a fourth input, so that the computer system cannot but automatically configure the operation to be executed in the absence of an input, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0256] In some implementations, after configuring the first operation to be performed in the absence of detected input, the computer system detects a first gesture (e.g., 805A) (and / or a second gesture and / or a gesture identical to the first gesture) via one or more input devices (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some embodiments, in response to the detection of a first pose (e.g., 805A), based on determining that a first operation has been performed without the detection of a first input (e.g., 805B) (and in some embodiments, in response to the detection of a different pose identical to the first pose) more than a threshold number of times (e.g., 1 to 50 times), the computer system abandons the execution of the first operation (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a first gesture, based on the determination that the number of first operations performed without the first input (e.g., 805B) is less than a threshold number (e.g., 1 to 50 times), the computer system performs a first operation (e.g., as described above regarding...). Figures 8A to 8E(As described). By automatically performing the first operation based on determining that the first operation has been performed less than a threshold number of times without detecting the first input, the computer system can automatically perform the operation before the input needs to be detected and the gesture to perform the operation is required, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0257] In some implementations, prior to detecting the first gesture (e.g., 805A), the first operation has been tracked (e.g., recorded, identified, marked, and / or stated) as being performed a first number (e.g., 1 to 10 times) without input being detected (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, based on the determination that a first gesture has been detected (e.g., 805A) along with the detection of a first input (e.g., 805B), the first operation has been performed a second number of times (e.g., 0 to 1 times) less than the first number when no input has been detected (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, the second number is less than the first number. In some implementations, the second number is a reset value (e.g., 0 to 1). In some implementations, based on the determination that a first gesture was detected (e.g., 805A) along with the detection of a first input (e.g., 805B), the first operation has been performed a third number (e.g., 1 to 10 times) greater than the first number even when no input was detected (e.g., as described above regarding...). Figures 8A to 8E (As described). In some implementations, the third count is one more than the first count. In some implementations, the third count is more than one more than the first count. In some implementations, the first count is refreshed after a gesture along with the first input is detected. The first operation is automatically performed based on determining that the first operation has been performed less than a threshold number of times without detecting the first input, and the threshold number varies based on the detected gesture. This allows the computer system to refresh the counter to deconfigure the gesture even without input to perform the operation, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without further user input, and providing improved feedback to the user.
[0258] In some implementations, after configuring the first operation to be performed in the absence of detected input, the computer system detects a first gesture (e.g., 805A) (and / or a second gesture and / or a gesture identical to the first gesture) via one or more input devices (e.g., as described above regarding...). Figures 8A to 8E(As described above). In some implementations, in response to the detection of a first gesture (e.g., 805A), based on determining that a threshold time amount has elapsed (e.g., 1 second to 1000 seconds) (e.g., because the first gesture and the first input were detected together and / or together with each other at a previous time that exceeds the threshold time amount from the current time (e.g., the time after the first operation is configured to be performed without input being detected) (e.g., an inactivity threshold amount), the computer system abandons the execution of the first operation (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to detecting a first pose, the computer system performs a first operation based on determining that a threshold amount of time has not yet elapsed (e.g., as described above regarding...). Figures 8A to 8E (As described). The computer system can automatically perform operations based on whether a threshold time has elapsed, without requiring user prompts based on the threshold time. This reduces the amount of input required to perform the operation, provides additional control options without cluttering the user interface with additional displayed controls, executes operations without further user input, and provides improved feedback to the user.
[0259] In some implementations, the threshold time quantity is the threshold time quantity of inactivity (e.g., as mentioned above regarding...). Figures 8A to 8E (as described) (e.g., 1 second to 1000 seconds) (e.g., no input has been detected, no interaction with the computer system has occurred, and / or the computer system has not been used to perform an operation in response to the detection of input and / or gesture). Performing or not performing a first operation based on whether a threshold time has elapsed enables the computer system to automatically perform the operation without prompting the user before the inactivity time has elapsed, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0260] In some implementations, the threshold time is the amount of time since the first input (e.g., 805B) was detected (e.g., 1 second to 1000 seconds) (e.g., as described above regarding...). Figures 8A to 8E (As described). The computer system can automatically perform an operation based on whether a threshold time has elapsed to determine whether to perform the first operation. This allows the system to respond to an action detected only when a gesture is detected, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0261] In some implementations, the threshold time is the amount of time (e.g., 1 second to 1000 seconds) since the detection of the first pose (e.g., 805A) (and / or a pose of the same type as the first pose) (e.g., as described above regarding...). Figures 8A to 8E (as described) (e.g., regardless of whether input is received) (e.g., no input is detected). Performing or not performing a first operation based on whether a threshold time has elapsed enables the computer system to automatically perform an operation in response to detecting only a gesture as recently detected, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0262] In some implementations, the threshold time is the amount of time (e.g., 1 second to 1000 seconds) since the first pose (e.g., 805A) was detected along with the first input (e.g., 805b) (e.g., as described above regarding...). Figures 8A to 8E (As described) (e.g., at a previous time exceeding a threshold time (e.g., the time after the first gesture is detected following the first operation being configured to be performed without the first input being detected), the first gesture and / or the first input have been detected together and / or together with each other). Executing or not executing the first operation based on whether the threshold time has elapsed enables the computer system to automatically execute the operation in response to detecting only the gesture when both the gesture and input have been recently detected, thereby reducing the amount of input required to execute the operation, providing additional control options without cluttering the user interface with additional displayed controls, executing the operation without requiring further user input, and providing improved feedback to the user.
[0263] In some implementations, after performing the first operation and after configuring the first operation to be performed without detecting the first input (e.g., 805B), the computer system detects a twelfth pose different from the first pose (e.g., 805A) via one or more input devices without detecting the first input (e.g., as described above regarding...). Figures 8A to 8E (As described above). In some implementations, in response to the detection of a twelfth gesture but the absence of a first input (e.g., 805B), the computer system performs a first operation (e.g., as described above regarding...). Figures 8A to 8E(As described). In some embodiments, the twelfth posture is similar to but different from the first posture (e.g., head shaking from left to right versus head shaking from right to left and / or head nodding from bottom to top versus head nodding from top to bottom). Responding to the detection of the twelfth posture but the absence of the first input, and after performing the first operation and after configuring the first operation to be performed in the absence of the first input, the computer system is able to detect different postures (e.g., similar postures) and perform the operation in the absence of input, thereby reducing the amount of input required to perform the operation, providing additional control options without cluttering the user interface with additional displayed controls, performing the operation without requiring further user input, and providing improved feedback to the user.
[0264] It should be noted that the above text regarding method 900 (for example, Figure 9 The details of the process described herein also apply in a similar manner to the methods described below / above. For example, method 700 may optionally include one or more characteristics of the various methods described above with reference to method 900. For example, a computer system may perform a representation of movement using the techniques described with respect to method 900 in response to detecting a gesture and a representation of movement induced by the techniques described with respect to method 700. For the sake of brevity, these details will not be repeated below.
[0265] Figures 10A to 10C An exemplary user interface for automatically outputting content based on context and / or input characteristics, according to some implementation schemes, is illustrated. Figures 10A to 10C The user interface in the document is used to illustrate the processes described below, including Figure 11 and Figure 12 The process in.
[0266] Figures 10A to 10CThe computer system 1000 is illustrated as a tablet computer displaying different user interfaces. It should be understood that the computer system 1000 can be other types of computer systems, such as smartphones, smartwatches, laptops, public facilities, smart speakers, accessories, personal gaming systems, desktop computers, fitness trackers, and / or head-mounted display (HMD) devices. In some embodiments, the computer system 1000 includes one or more sensors (e.g., cameras, LiDAR detectors, motion sensors, infrared sensors, and / or microphones) and / or communicates with them. In some embodiments, the computer system 1000 includes one or more output devices (e.g., displays, projectors, touch-sensitive displays, and / or speakers) and / or communicates with them. In some embodiments, the computer system 1000 includes one or more moving components (e.g., actuators, movable bases, rotatable components, and / or rotatable bases) and / or communicates with them. In some embodiments, the computer system 1000 includes one or more components and / or features described above with respect to computer system 100 and / or electronic device 200.
[0267] Figures 10A to 10C This illustrates a scenario where a computer system 1000 detects input and adjusts its content output in response to the input based on the detected contextual characteristics. Figures 10A to 10C In this context, computer system 1000 detects contextual characteristics via one or more sensors connected to and / or communicating with computer system 1000. Figures 10A to 10C The illustrated scenario covers an example in which a computer system 1000, in response to input, outputs content having characteristics corresponding to a first set of one or more learned characteristics when operating in a first context, and outputs content having characteristics corresponding to a second set of one or more learned characteristics when operating in a second context. In some embodiments, the second context and the corresponding second set of one or more learned characteristics are different from the first context and the corresponding first set of one or more learned characteristics. In some embodiments, the second context and the corresponding second set of one or more learned characteristics are the same as the first context and the corresponding first set of one or more learned characteristics.
[0268] In some implementations, contextual characteristics include the location in which the computer system 1000 is currently operating (e.g., home, car, theater, museum, park, and / or workplace). For example, the computer system 1000 may respond differently to the same input when located in a park compared to when located in a car (where the user may not be able to give the computer system 1000 their full attention). In some implementations, contextual characteristics include environmental characteristics of the location where the computer system 1000 is currently located (e.g., volume, light level, and / or the number of users in the environment). For example, the computer system 1000 may respond differently to the same input in an environment with high volume (e.g., an environment where the user is considered to have difficulty hearing the audio output by the computer system 1000 (e.g., an environment with a large amount of audio competing with the audio output by the computer system 1000)) compared to an environment with low volume (e.g., an environment where the user is considered to have easy access to the audio output by the computer system 1000 (e.g., an environment with a small amount of audio competing with the audio output by the computer system 1000)). In some implementations, contextual characteristics include user characteristics. In some implementations, user characteristics include facial expressions (e.g., smiling, crying, frowning, laughing, and / or lack of emotion), user clothing (e.g., earmuffs, sunglasses, goggles, face mask, and / or gloves), user activities (e.g., phone calls, video calls, running, driving, talking to another user, cooking, and / or relaxing), calendar events, and / or audio characteristics of the user's audio output (e.g., pitch, volume, rhythm, and / or cadence). For example, computer system 1000 may respond differently to the same input when it detects that a user is running versus when it detects that a user is preparing to go out. In some implementations, user characteristics include general characteristics applicable to a number of users. In some implementations, user characteristics include specific characteristics applicable to a particular user. For example, in response to the detection of input, computer system 1000 may output a different response when it detects a smile from a first user versus when it detects a smile from a second user.
[0269] In some implementations, contextual characteristics include characteristics of the detected input. For example, computer system 1000 may respond differently to the same tap input depending on whether the tap input is detected as a hard tap or a soft tap. Examples described below include verbal input and touch input. In some implementations, input includes air gestures. In some implementations, input includes gaze input. In some implementations, input includes user movement in the environment. One example described below includes input corresponding to a request and / or question (e.g., a user requests information about the outside temperature (e.g., “What’s the outside temperature?”)). In some implementations, if the input corresponds to a request and / or question, computer system 1000 outputs content corresponding to the request and / or question. In some implementations, the input corresponds to a declarative statement. For example, a user can say, “It looks like it’s going to rain today.” In some implementations, if the input corresponds to a declarative statement, computer system 1000 outputs content corresponding to the declarative statement. For example, in response to detecting the audio input “It looks like it’s going to rain today,” computer system 1000 may output the audio “There’s a 60 percent chance it will rain today.” In some implementations, the input corresponds to a command. For example, a user can say, “Tell me the weather today.” In some implementations, if the input corresponds to a command, the computer system 1000 outputs the content corresponding to the command.
[0270] like Figure 10A As illustrated, computer system 1000 displays a main screen user interface 1002. In this example, the main screen user interface 1002 includes a weather application control 1004, a text application control 1006a, an email application control 1006b, a telephone application control 1006c, and a calendar application control 1006d. Figure 10A As illustrated, text application control 1006a, email application control 1006b, telephone application control 1006c, and calendar application control 1006d are displayed along the bottom edge of the main screen user interface 1002. Similarly... Figure 10A As illustrated, computer system 1000 displays weather application controls 1004 in the upper left position within the main screen user interface 1002. In some embodiments, the main screen user interface 1002 includes one or more other controls and / or indicators. In some embodiments, in Figure 10A In the middle, computer system 1000 detects a tap input 1005a1. In some implementations, in Figure 10A In the middle, computer system 1000 detects audio input 1005a2. In some implementations, in Figure 10A In this context, computer system 1000 determines that it is operating in a first context. In some implementations, in... Figure 10A In this context, computer system 1000 determines that computer system 1000 is operating in a second context.
[0271] Figure 10B An example is illustrated of a computer system 1000 in a first context, which corresponds to a first set of one or more contextual characteristics detected and learned by the computer system 1000. In this example, the first context corresponds to the context in which the computer system 1000 outputs information in a complete format (e.g., a complete sentence and / or multiple messages). In some embodiments, the first set of one or more contextual characteristics includes the computer system 1000 being located in a location frequently visited by the user (e.g., home, office, park, cafe, and / or friend's house). For example, the computer system 1000 and the user may be located in the user's home, a location where the user is more likely to feel comfortable in the environment. In some embodiments, the first set of one or more contextual characteristics includes environmental characteristics that facilitate the user's reception of the content output by the computer system 1000 (e.g., proximity to other users, maximum volume level, maximum light level, and / or minimum light level). For example, the computer system 1000 and the user may be located in a quiet room with a light level comfortable for reading. In some embodiments, the first set of one or more contextual characteristics includes characteristics of the user. For example, the user may remain calm and not engage in any activity that would require additional attention to be taken away from the computer system 1000.
[0272] In some embodiments, the first set of one or more contextual characteristics includes audio input 1005a2 corresponding to the first set of one or more audio characteristics (e.g., pitch, volume, rhythm, and / or cadence). For example, the first set of one or more audio characteristics may include audio input 1005a2 having characteristics of a speaking user's normal rhythm and volume. In some embodiments, the first set of one or more contextual characteristics includes tap input 1005a1 having characteristics of a normal press compared to other presses by the user. In some embodiments, tap input 1005a1 is gaze input, and the first set of one or more contextual characteristics includes gaze input for sustained gaze. In some embodiments, tap input 1005a1 is an air gesture, and the first set of one or more contextual characteristics includes air gestures for controlled air gestures.
[0273] like Figure 10B As illustrated, in response to detecting a tap input 1005a1 and / or an audio input 1005a2, the computer system 1000 stops displaying the main screen user interface 1002 and displays the weather application user interface 1012. Figure 10BAs illustrated, in response to determining that the computer system 1000 is operating in a first context and detecting a tap input 1005a1 and / or an audio input 1005a2, the computer system 1000 displays a weather indication 1014 corresponding to one or more learned context features of the first group. Figure 10B As illustrated, in response to the detection of a tap input 1005a1 and / or an audio input 1005a2, the computer system 1000 displays a weather indicator 1014 at a central location within the weather application user interface 1012. In this example, the display of the weather indicator 1014 by the computer system 1000 corresponding to one or more learned contextual characteristics of a first group includes displaying a weather indicator 1014 with multiple lines of information (e.g., current location, current outside temperature, current weather conditions, and / or the predicted high and / or low temperatures for the day). In some embodiments, the computer system 1000 displays the weather indicator 1014 and simultaneously outputs an audio response 1016.
[0274] like Figure 10B As illustrated, in response to determining that the computer system 1000 is operating in a first context and detecting a tap input 1005a1 and / or audio input 1005a2, the computer system 1000 outputs an audio response 1016 having a second set of one or more audio characteristics corresponding to a first set of one or more learned contextual characteristics. In this example, the computer system 1000 outputting an audio response 1016 having a second set of one or more audio characteristics includes outputting a sentence indicating the temperature (e.g., “It’s currently 75 degrees outside”). In some embodiments, the computer system 1000 outputting audio with the second set of audio characteristics corresponds to the computer system 1000 outputting audio with default audio characteristics. In some embodiments, the default audio characteristics are audio characteristics preset (e.g., previously configured) by the publisher of the weather application user interface 1012 (e.g., or otherwise defaulted to, rather than based on user input and / or what is being output). In some embodiments, the default audio characteristics are audio characteristics preset by the user. In some embodiments, the computer system 1000 outputting audio with the second set of audio characteristics corresponds to audio characteristics that the computer system 1000 matches the user's speech. For example, in response to determining that the computer system 1000 is operating in a first context and detecting a tap input 1005a1, if the computer system 1000 detects that the user is speaking with a uniform rhythm, relaxed tempo, and conversational volume, the computer system 1000 may output an audio response 1016 with the same and / or similar uniform rhythm, relaxed tempo, and conversational volume. In some embodiments, the audio characteristics of the user's speech are sufficiently similar to the audio characteristics of a default audio characteristic, such that the computer system 1000 outputs audio with the default audio characteristics when matching the audio characteristics of the user's speech.
[0275] In some implementations, the computer system 1000 matches the user's level of detail. For example, if the user provides concise input such as "What's the weather like today?", the computer system 1000 might infer a preference for low detail and respond with a concise output such as "Sunny, 72 degrees." Conversely, if the user asks, "Can you give me today's weather forecast? Including the temperature range, chance of precipitation, and wind conditions?", the computer system 1000 might provide a more detailed and / or more comprehensive weather report to match the user's level of detail and / or inferred expected engagement. For example, the computer system 1000 might respond, "The weather looks very nice today. We expect the temperature to be between 68 and 75 degrees. There is a small chance of rain in the afternoon, so you might want to bring an umbrella just in case." The wind is from the northwest, not very strong, approximately 5 to 10 mph. In some implementations, the user's level of detail corresponds to the number of words and / or phrases used to convey content (e.g., requests, messages, statements, and / or commands). In such implementations, the content may have an average and / or minimum level of detail, such that matching the user's level of detail may include increasing or decreasing the level of detail of the response based on whether the user's level of detail is below or above the average and / or minimum level of detail and / or whether the user's level of detail is below or above the average and / or minimum level of detail. In some implementations, changing the level of detail of the content does not change the content's... The level of detail should not result in certain parts of the content not being conveyed. In some implementations, increasing the level of detail involves using more words to convey the same content. In some implementations, decreasing the level of detail involves using fewer words to convey the same content. It should be recognized that the level of detail may only consider a single most recent input from the user and / or a predefined amount of recent input from the user (e.g., including multiple inputs, such as by averaging the level of detail of those inputs and / or taking the minimum, maximum, and / or mode of those inputs). It should also be recognized that matching detail may be performed as a complement to or alternative to the matching audio features described above.
[0276] In some embodiments, in response to detecting audio input 1005a2 corresponding to one or more audio characteristics of a first group, computer system 1000 outputs audio with one or more audio characteristics of a third group. In some embodiments, the third group of one or more audio characteristics is the same as the first group of one or more audio characteristics. For example, in response to detecting audio input 1005a2 with musical rhythm and pitch-changing characteristics, computer system 1000 may output an audio response 1016 with musical rhythm and pitch-changing characteristics, thereby reflecting the audio characteristics of the user's speech and extending to reflecting the user's emotions. In some embodiments, the first group of one or more audio characteristics is sufficiently similar to a default group of audio characteristics such that the third group of one or more audio characteristics is the same as the default audio characteristics. In some embodiments, the third group of one or more audio characteristics differs from the first group of one or more audio characteristics. For example, in response to detecting audio input 1005a2 with high pitch and uneven rhythm, computer system 1000 may output an audio response 1016 with lower pitch and even rhythm, creating a balanced contrast with the audio characteristics of audio input 1005a2.
[0277] Figure 10CAn example is illustrated of a computer system 1000 in a second context, which corresponds to a second set of one or more contextual characteristics detected and learned by the computer system 1000. In this example, the second context corresponds to the context in which the computer system 1000 outputs information in a shortened format (e.g., incomplete sentences and / or fewer messages). In some embodiments, the second set of one or more contextual characteristics includes the computer system 1000 being located in a location not frequently visited by the user (e.g., a hotel, amusement park, resort, and / or airport). For example, the computer system 1000 and the user may be located in an airport, a location where the user is more likely to be distracted and / or inattentive, and where a more concise response from the computer system 1000 may be easier for the user to understand. In some embodiments, the second set of one or more contextual characteristics includes the computer system 1000 being located in a location where lower volume is generally considered more appropriate (e.g., a library, religious site, museum, and / or theater). For example, the computer system 1000 and the user may be located in a library, where a lower volume and shorter response from the computer system 1000 is beneficial for not disturbing nearby users. In some implementations, the second set of one or more contextual characteristics includes environmental characteristics that are detrimental to a use...
Claims
1. A method, the method comprising: At a computer system that communicates with one or more input devices and one or more output devices: Input from a first user is detected via the one or more input devices, wherein the input includes a set of identification information for one or more users; After the input is detected, a second user different from the first user is detected via the one or more input devices; as well as In response to the detection of the second user: Based on the identification information that determines the second user corresponds to the group of one or more users, a first confirmation pointing to the second user is output via the one or more output devices; as well as Based on the determination that the second user does not correspond to the identification information corresponding to the group of one or more users, a second confirmation pointing to the second user is output via the one or more output devices, wherein the second confirmation is different from the first confirmation.
2. The method of claim 1, wherein the identification information corresponds to a set of one or more characteristics corresponding to the set of one or more users, wherein the second user corresponds to the identification information of the set of one or more users if the second user matches the set of one or more characteristics, and wherein the second user does not correspond to the identification information of the set of one or more users if the second user does not match the set of one or more characteristics.
3. The method of claim 1, wherein the set of one or more characteristics comprises a set of one or more visually identifiable features.
4. The method of claim 1, wherein the identification information includes a name corresponding to the group of one or more users, wherein the identification information of the second user corresponding to the name is determined to be true, and wherein the identification information of the second user corresponding to the group of one or more users is determined to be false.
5. The method of claim 4, wherein detecting the second user comprises detecting input including the name corresponding to the group of one or more users via the one or more input devices.
6. The method of claim 1, wherein the input from the first user is detected but the second user is not detected.
7. The method of claim 1, wherein the input is an audible input.
8. The method of claim 1, wherein the input from the first user includes a first indication of a first operation to be performed when one or more of the group of one or more users are detected.
9. The method of claim 1, wherein the input is a first input, and the method further comprises: A second input from the first user is detected via the one or more input devices, the second input including a second indication of a second operation to be performed when one or more of the group of one or more users are detected, wherein the second input is different from the first input.
10. The method of claim 9, wherein the second instruction for the second operation includes movement of the first user.
11. The method of claim 10, wherein the first confirmation is an indication of the movement.
12. The method of claim 1, wherein the input includes nonverbal input.
13. The method of claim 1, wherein the first confirmation includes a first audio characteristic, and wherein the second confirmation includes a second audio characteristic different from the first audio characteristic.
14. The method of claim 13, wherein the first audio characteristic includes outputting audio in a first language, and the second audio characteristic includes outputting audio in a second language different from the first language.
15. The method of claim 1, wherein the first confirmation includes a first movement, and wherein the second confirmation includes a second movement different from the first movement.
16. The method of claim 1, wherein the first confirmation includes a first visual characteristic, and wherein the second confirmation includes a second visual characteristic different from the first visual characteristic.
17. The method of claim 1, wherein the group of one or more users is a first group of one or more users, the method further comprising: Detect a third input from a third user via the one or more input devices, wherein the third input includes identification information of a second group of one or more users that is different from the first group of one or more users; as well as After the third input is detected, a fourth user is detected via the one or more input devices; as well as In response to the detection of the fourth user: Based on the identification information that determines the fourth user corresponds to one or more users in the second group, a third confirmation is output to the fourth user via the one or more output devices; as well as Based on the determination that the fourth user does not correspond to the identification information corresponding to one or more users in the second group, a fourth confirmation is output to the fourth user via the one or more output devices, wherein the fourth confirmation is different from the third confirmation.
18. The method according to claim 1, further comprising: The first user is detected via the one or more input devices; as well as In response to detecting the first user, a fifth confirmation directed to the first user is output via the one or more output devices, wherein the fifth confirmation is different from the first confirmation and the second confirmation.
19. A computer system that communicates with one or more input devices and one or more output devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 18.
20. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more input devices and one or more output devices, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 18.